BHI 3.1·Bulletin
Validation: preliminary
Today's bulletin →
Highest B
§ 06 · VALIDATION REPORT · V3.1

Validation

A preliminary, concurrent association between B-index scores and documented annual retention for 14 platforms — reported with the defects of the cohort stated beside every figure, and without a claim that it validates the formula.

Spearman ρ · as published
0.79
all n = 14 · p = 0.0009 (two-sided, t approximation)
Spearman ρ · measured rates only
0.74
n = 9 without asserted, non-retention and mismatched rows · p = 0.0237; n = 10 without the asserted rows: ρ = 0.74, p = 0.0153
ICC(3,k)
Pending
Pre-declared · recruitment deferred by decision · target > 0.75
Status
Concurrent
Single evaluator with anchored rubrics · no predictive test has been run

BHI score vs. annual retention

B-index vs. the retention figure recorded for each of the 14 platforms of paper §7.4 (SEC filings, earnings calls, CIRP, Antenna, industry reports). Hollow points are rows whose value is asserted rather than measured, is not a retention rate, or belongs to an object this index does not score. Dashed line is the OLS fit on all 14.

05101560%70%80%90%100%B-INDEX (V3.1)ANNUAL RETENTION RATE (%)Synopsys/Cadence (EDA Duopoly) · B=13.50, retention=99% (asserted, not a measured rate)TSMC · B=12.28, retention=99% (estimated rate (medium source quality))Visa · B=8.82, retention=99.9% (asserted, not a measured rate)Amazon · B=8.40, retention=95% (retention of an object this index does not score)NVIDIA (CUDA) · B=8.29, retention=98% (not a retention metric)Bloomberg Terminal · B=7.21, retention=96% (estimated rate (medium source quality))SWIFT · B=6.78, retention=99.9% (asserted, not a measured rate)Nubank · B=5.21, retention=94% (documented annual rate)Salesforce · B=4.76, retention=92% (documented annual rate)Apple · B=4.60, retention=89% (documented annual rate)Spotify · B=2.31, retention=96.5% (documented annual rate)Netflix · B=1.00, retention=78% (documented annual rate)Zoom · B=0.90, retention=92% (documented annual rate)Disney+ · B=0.71, retention=61% (documented annual rate)ρ = 0.79 (p = 0.0009) · n = 14 · hollow = asserted / mismatched

What this cohort shows, and what it cannot

Retention is positively associated with B in this cohort. The association is concurrent — the retention periods precede or coincide with the scores — so it is not a prediction, and no prospective test has been run. A monotone “retention floor” does not hold: in 16 of 88 comparable platform pairs the higher-B platform has the lower retention (Spotify, B = 2.31, retains 96.5% against Apple’s 89% at B = 4.60).

The cohort does not distinguish the full formula from simpler forms. On the same 14 objects and the same outcome, the escape half of the model alone gives ρ = 0.87 and the single parameter portability gives ρ = 0.87, against ρ = 0.79 for B; every paired difference between these correlations includes zero at this sample size. No incremental validity of the formula’s complexity, and no predictive validity, has been demonstrated. This cohort is a preliminary external-association check, not evidence for the formula.

Limitations of this cohort, corrected 13 September 2026:
Small sample (n = 14) · Concurrent, not predictive · Author-selected · One retention figure belongs to an object this index does not score (Amazon Prime’s renewal rate against the Amazon ecosystem row) · One value is not a retention rate (NVIDIA: AI-framework share) · Three values are characterisations, not measured rates (EDA, Visa, SWIFT) · Heterogeneous metric definitions and source quality · Single evaluator · The association is carried by the contrast between infrastructure rows near 99% and consumer rows; within either group it is weak · Correlation does not establish causation
Scoring-noise and aggregation-form experiments: robustness note

Reference cohort · n = 14

The 14 platforms of paper §7.4, with the basis of each retention value. The paper’s figure is reproducible from this table; the “measured rates only” figure above excludes the rows marked asserted, non-retention or mismatched.

PlatformSectorBHI · V3.1Retention (%)SourceBasis
Synopsys/Cadence (EDA Duopoly)infra13.5099.0%Industry reportsasserted, not a measured rate — Period recorded as "Structural"; a characterisation of the market, not a measured annual rate.
TSMCtech12.2899.0%TSMC annual reportestimated rate (medium source quality)
Visafintech8.8299.9%Visa 10-K filingasserted, not a measured rate — "Zero voluntary network departures" coded as 99.9%.
Amazontech8.4095.0%CIRP consumer surveyretention of an object this index does not score — The figure is Amazon Prime’s annual renewal rate; the scored row is the Amazon ecosystem, and no Prime row exists.
NVIDIA (CUDA)tech8.2998.0%CUDA developer surveynot a retention metric — The underlying metric is AI-framework share (~92%), stored as 98; it is not a retention rate.
Bloomberg Terminalbanks7.2196.0%Bloomberg LP reportsestimated rate (medium source quality) — "<5% annual churn estimated".
SWIFTbanks6.7899.9%SWIFT annual reviewasserted, not a measured rate — "Zero voluntary exits since 1973" coded as 99.9%.
Nubankfintech5.2194.0%Nubank earningsdocumented annual rate
Salesforcesaas4.7692.0%Salesforce 10-Kdocumented annual rate
Appletech4.6089.0%Antenna churn datadocumented annual rate
Spotifygaming2.3196.5%Spotify earningsdocumented annual rate
Netflixgaming1.0078.0%Antenna churn datadocumented annual rate
Zoomsaas0.9092.0%Zoom earningsdocumented annual rate
Disney+gaming0.7161.0%Antenna churn datadocumented annual rate

Validation roadmap

Stated as current practice
Scoring protocol: “Every parameter score must be backed by a citation.” “Each score includes a prose justification.” “All scores, rationales, and computed B values are published.” (/methodology §07)
IN PROGRESS · assessed 2026-09-19 As of 19 September 2026: structured per-parameter assessment and provenance records exist for a subset of the universe, 16 of the 184 platforms (176 current rows with a prose rationale; citations on 35 cells across 12 platforms); this work is incomplete and is not yet served as release-level evidence. The public history API returns a change reason for the 14 platforms re-scored in Q4-2026 and no citation, and of the 184 platforms one carries a stored platform-level rationale. Established 13 September 2026: platform-level rationales (one paragraph each, 13–46 words, no per-parameter citation) were written for the 100 platforms of the Q1-2026 release and published until the 4 June 2026 migration to the database, which carried none of them over; the 84 platforms added in Q2-2026 never had one. Whether the 100 are restored is a pending decision. Score changes are logged with a reason; the standing levels are not sourced. The protocol is met in full only in the determinations prepared for regulators, which record score, evidence tier, searches, sources and provenance for each of eleven parameters. record
Q2 26
Expanded reference set. Lift cohort from 14 to 30+ platforms with documented switching cost or retention data.
NOT EXECUTED · assessed 2026-09-07 The reference cohort remains n = 14. No expansion was carried out.
Q3 26
Independent evaluators. Recruit 3-5 independent scorers. Compute ICC and Krippendorff’s alpha on 15-20 platforms.
PREPARED, NOT STARTED · assessed 2026-09-07 The cohort was drawn by lot from the pinned release and frozen at n = 22 on 22 August 2026, larger than the 15–20 promised; the statistical plan was frozen on 6 September 2026 before any evaluator was approached, with its digest published. No evaluator has been approached and no rating exists. Recruitment was deferred by decision on 8 September 2026: the study stays pre-declared and the cohort frozen, and it is not being staffed. record
Q3 26
Public scoring rubric audit. Anchor definitions opened for 30-day public comment.
NOT EXECUTED · assessed 2026-09-07 No comment period has been opened, no channel for comment has been published, and the anchor definitions stand as first written. The rubrics are public and can be objected to at any time, but that is not the audit that was promised.
Q4 26
Pre-registered V4 study. OSF-registered protocol filed before scoring begins.
IN PROGRESS · assessed 2026-09-07 A pre-declared protocol exists, frozen by digest before execution — estimators, scale, interval procedure and thresholds, each with a stated consequence — with its digest published on 6 September 2026. It is not filed on the Open Science Framework, which is what was promised; the filing has not been done. record
Q1 27
External replication package. Full scoring spreadsheets and code released under CC BY 4.0.
IN PROGRESS · assessed 2026-09-07 The analysis script is written, tested against known answers and frozen by digest, and a rater pack is generated from the published rubrics. The scoring spreadsheets are not released, and the package as promised is not complete. record
The rows above are the commitments as first published, unedited. The status beside each one is held in a single register and assessed against the record on the date shown, so that a promise and its outcome cannot drift apart in two separately maintained lists — which is what happened to the note that stood here before, and why it is gone.

Open challenges

  1. Single evaluator. All 184 platforms were scored by a single evaluator. This is the most critical methodological limitation. The reliability study is pre-declared, its protocol and cohort frozen by digest before execution, but recruitment is deferred — so this limitation stands until independent evaluators score that cohort.
  2. Two parameters have no recorded reference subject. Closeness and human fallback describe a state of the user rather than a property of the platform, and the rubric does not say which user. Coinbase’s closeness of 7 — “first interface opened each morning” — is reachable for an active trader and unreachable for the median registered account; because closeness multiplies capture directly, those two defensible readings put its B at 0.78 and 1.83 — values that fall in two different zones, Transition Zone and Event Horizon. No score is withdrawn and none is refuted: what is established is that these two parameters, as published, cannot be tested against evidence however good the evidence is. An attempt to evidence Coinbase’s closeness from audited SEC filings reached that conclusion on 8 September 2026, and the cloud-provider study had reached it independently on 22 August and recorded both parameters as not scorable. record
  3. Related rows are not independent measurements. Where one platform sits inside another’s corporate group, the two agree on 4.4 of the eleven parameters on average, against 2.7 for pairs merely in the same sector and 2.3 for pairs at random; in ten thousand random draws of twenty same-sector pairs, eleven reached that level. The effect is strongest in network effects (65% agreement against a 25% baseline) and portability, and it is weakest in closeness — refuting the hypothesis that prompted the test, which was pre-declared and hashed before any of it was computed. Two readings fit: related platforms genuinely resemble each other, or they were not scored independently. The record cannot choose, because the served dataset carries no recorded reasoning for 183 of 184 rows (platform-level rationales written for the 100 platforms of the Q1-2026 release were not carried into the database at the June 2026 migration; whether they are restored is a pending decision). One fact cuts against the second reading and is reported because it does: in thirteen cases a member scores above its parent on a capture parameter. No score is withdrawn; what follows is that ranking a platform against its own parent within one distribution has no basis in the record. record
  4. Sector heterogeneity. One scale is applied across every sector, and outside B = 1 no part of it is calibrated. B = 1 is definitional: the point at which capture equals escape. The four zone boundaries around it have no published derivation, and the zone names are interpretive labels rather than measured categories, so a comparison of rank or of relative position is better supported by this record than a reading of the absolute level. Sector-relative percentiles may be more useful than the global scale. Corrected 19 September 2026: this line read “The B-scale is calibrated globally”, in which the word calibrated asserted something this project does not hold and states the opposite of elsewhere.
  5. Time resolution. Quarterly cadence is coarse for fast-moving AI-platform dynamics. The quarterly record holds three observed vintages (Q1-2026 to Q3-2026); the Q4-2025 rows in the history are a modelled back-projection and are labelled as such.
  6. Causality. BHI is descriptive, not causal. Correlation does not establish causation.