Black Hole Index — frozen research snapshots ============================================ Public timestamp-and-hash records. The content of each snapshot remains private until its quality assurance completes. The hash fixed here on the day of freezing proves that whatever is later published is the content that existed on the stated date — not a back-edit. Snapshot S1 ----------- Name: cloud-core-review-package-20260822.tar.gz Date frozen: 2026-08-22 SHA-256: 6565778b84b38a9420c41fec37219303548cf6b7fb52f88237c6677b949712fe Contents (15 files; per-file SHA-256 inside the package manifest): the BHI regulatory protocol; the provider-neutral Cloud Core measurement boundary; the evidence rules; the comparability matrix (48/48 cells confirmed against first-party documentation); four provider evidence ledgers with the determination complete (8 parameters scored, 3 recorded as not scorable, for all four providers); the construct-validity audit of all 32 scored rows; the comparative structural profile; a pre-registered CMA-remedy prediction register; the open research issues; the integrated-stack boundary note; and the two enforcement scripts. Statement: frozen before any external review of the corpus and before the outcome of the CMA Board's six-month cloud progress review. External timestamps (added 2026-08-22, same day) ------------------------------------------------ The SHA-256 above was independently timestamped via RFC 3161 (hash only; no content transmitted): FreeTSA (freetsa.org) genTime 2026-08-22 17:29:46 UTC — token verified: OK DigiCert (timestamp.digicert.com) genTime 2026-08-22 17:30:00 UTC — imprint matches Third parties therefore attest that this exact digest existed on 22 August 2026. The original S1 block above is unchanged since first publication. Snapshot S2 ----------- Name: integrated-stack-package-20260829.tar.gz Date frozen: 2026-08-29 SHA-256: 6f5798718c763a2c718eb65b0c1bccdfdc579d118a4eb584ceebbd04e102ebaf Contents (793 files; per-file SHA-256 inside the package manifest): the Integrated Stack measurement boundary (v1.0.8, with every superseded version archived and hashed), the rules delta (v1.7), the evidence journal (frozen v1.0 2026-08-29 08:00 UTC, revised 1.0.1/1.0.2 with status and declaration lines only), the structural profile (v1.0: eight parameter values, no composite), the standalone-anchor matrix, the comparability register, the financial-exit specification and register (v1.3, generated by script), the evidence ledgers with fixed question sets, search-opening stamps and retained page captures, and the two enforcement scripts. Statement: frozen after twenty-one internal audit passes and before any external review of the corpus; the profile records every value withdrawn or lowered on review and every reading registered after scoring, with the principal investigator's confirmation status stated. External timestamps (added 2026-08-29, same day) ------------------------------------------------ The SHA-256 above was independently timestamped via RFC 3161 (hash only; no content transmitted): FreeTSA (freetsa.org) genTime 2026-08-29 17:41:50 UTC — token verified: OK DigiCert (timestamp.digicert.com) genTime 2026-08-29 17:41:51 UTC — token verified: OK Snapshot S3 ----------- Name: jftc-package-20260829.tar.gz Date frozen: 2026-08-29 SHA-256: ae8ca1aacf27a9276273b2240076bd3df3b8c6735add47d05d9aaf167f92aa9e Contents (116 files; per-file SHA-256 inside the package manifest): the JFTC Mobile Core measurement boundary (v1.0.1, frozen 2026-08-26 before any evidence row), the pre-registered obligation register (v1.0.1: observables, assumptions, predictions and falsifiers, frozen 2026-08-26 ≈17:55 UTC), the evidence ledger (frozen v1.0 2026-08-29 13:01 UTC: S-pre and S-post for the Apple and Google objects, predictions graded), the hashed primary sources (four operator compliance reports, the JFTC designation record, the Act, Cabinet Order and Rules from e-Gov, Apple's MSCA data-use disclosure), the fixed question set with its search-opening stamp, the capture log, the retained captures and byte-verified quotation sets, and the enforcement script. Statement: the register was frozen before the evidence was gathered; the ledger grades the register's predictions on dated first-party policy text and prints every assumption that did not hold. External timestamps (added 2026-08-29, same day) ------------------------------------------------ The SHA-256 above was independently timestamped via RFC 3161 (hash only; no content transmitted): FreeTSA (freetsa.org) genTime 2026-08-29 17:41:51 UTC — token verified: OK DigiCert (timestamp.digicert.com) genTime 2026-08-29 17:41:51 UTC — token verified: OK Snapshot S4 ----------- Name: integrated-stack-package-20260831.tar.gz Date frozen: 2026-08-31 SHA-256: d7a3385e0b0bfd466a84c6516161c995e24e537a56eb1c1d4e580e9a24428f0a Contents (802 files; per-file SHA-256 inside the package manifest): the Integrated Stack measurement boundary (v1.0.9, with every superseded version archived and hashed), the rules delta (v1.7), the evidence journal (v1.0.2), the structural profile (v1.0.7, revisions 1.0.1-1.0.7 dated and archived), the standalone-anchor matrix, the comparability register, the financial-exit specification and register (v1.3, generated by script), the evidence ledgers with fixed question sets, search-opening stamps and retained page captures, the findings note for the principal investigator, and the two enforcement scripts. Statement: successor state of Snapshot S2, with every subsequent revision of the boundary (to 1.0.9) and the profile (to 1.0.7) dated, archived and hashed inside the package; disclosure text only - no scored value changed after the evidence journal froze. Frozen after the derived information provision to the UK CMA completed its hostile audit passes and before that document's despatch or any external review of the corpus. External timestamps (added 2026-08-31, same day) ------------------------------------------------ The SHA-256 above was independently timestamped via RFC 3161 (hash only; no content transmitted): FreeTSA (freetsa.org) genTime 2026-08-31 07:47:42 UTC — token verified: OK DigiCert (timestamp.digicert.com) genTime 2026-08-31 07:47:44 UTC — token verified: OK Correction to the S4 contents line (2026-08-31, same day) --------------------------------------------------------- The phrase "with every superseded version archived and hashed" in the S4 contents above overstates: boundary bytes before revision 1.0.3 were not retained (declared in the boundary's revision 1.0.4; 1.0.3 is the first archived state), and the comparability register's superseded 0.1/0.2/0.3 bytes are likewise not archived. Read: "with every retained superseded version archived and hashed"; the declared exceptions are enumerated in the despatched document (§1.3, §7.2). The same qualification applies to the identical phrase in the S2 contents line. The archive digests are unaffected; no snapshot block above is edited. Clarification to the correction above (2026-08-31, same day) ------------------------------------------------------------ The correction's phrase "the despatched document" named a despatch that had not yet occurred when the correction was appended. Read instead: "the BHI information provision to the CMA Microsoft business-software SMS investigation, dated 31 August 2026 (its §1.3 and §7.2), available to the case team on request". Nothing above is edited. Status note (2026-08-31 08:19 UTC) ---------------------------------- As at this timestamp, the information provision named above — "Information provided ahead of the proposed decision: a structural profile of Microsoft's integrated business-software stack" (31 August 2026) — has not been despatched. Hostile audit passes on that document continued after the S4 freeze; they changed the outward document, its preparation logs and this file (the two dated appends above and this note). No byte of any snapshot archive is affected; the digests above stand. Status note (2026-09-01 19:47 UTC) ---------------------------------- The information provision named above was despatched to the CMA case team by the principal investigator on 31 August 2026 (his mailbox record shows 08:28 UTC, nine minutes after the preceding note was appended; the case team acknowledged receipt and circulation on 1 September 2026). The 31 August status note was true when written and is superseded by this one. No snapshot archive byte is affected; the digests above stand. Status note (2026-09-01 22:49 UTC) — comparability register ----------------------------------------------------------- The correction of 31 August above declared that the Integrated Stack comparability register's superseded 0.1/0.2/0.3 bytes are not archived. The information provision despatched to the CMA on 31 August (its §1.3) described that exception as visible in the register's own revision history; at despatch the history carried no such statement. On 1 September the register was revised (0.3a to 0.3b, superseded bytes archived and hashed) to carry the declaration itself, with the sequence disclosed in the revision row. No cell, link, test or score changed; the S2 and S4 manifests continue to fix the 0.3a bytes. Clarification to the comparability note above (2026-09-01 23:00 UTC) -------------------------------------------------------------------- The despatched provision's §1.3 wording is, verbatim, "visible only in the comparability register's own revision history". The register's revision 0.3b of 1 September supplies the declaration that description presupposed; the word "only" was inaccurate at despatch and remains so, this public record having carried the declaration first (31 August). The superseded register states are archived and hashed (0.3a: 09ac25c6...; the interim note state: 24a07965...) and the 0.3b bytes are pinned in the enforcement script. Addendum (2026-09-01 23:12 UTC) ------------------------------- Three precisions to the two notes above. (1) §1.3's same parenthetical also states that the comparability register carries no SHA pin in the enforcement script; that was true at despatch and is no longer true — the 0.3b bytes are pinned since 1 September, so the statement describes the position at despatch only. (2) The two register states superseded on 1 September are archived and hashed; the 0.1/0.2/0.3 bytes remain unretained. (3) Quotations of §1.3 in these notes normalise the typographic apostrophe to ASCII. Prior-Knowledge Ledger S1 (2026-09-02 06:42 UTC) ------------------------------------------------ A post-hoc prior-knowledge ledger for snapshot S1 now exists at paper/cloud-core/PRIOR-KNOWLEDGE-LEDGER-S1.md (v1.1, 2 September 2026; file sha256 5572d7f040e8c0b6482a21f7d444283a8ba5cc6412a9204ae5016f4a19288081). It is not part of the frozen corpus and says so. Its Section 2 fixes a CLOSED list of eight incremental claims (IC1-IC8); that section's byte digest, sha256 59af82be991e68fdbf9661cfc458b1417b2058a2c3904b7a4d8742d98e0418c8, is pinned in the enforcement script and nothing will ever be added to the list. Section 4 (gradings of later external events) grows by dated entries; each such revision changes the file digest and is fixed here by a further dated append, with the Section 2 digest unchanged. The v1.0 bytes of the same day are archived and hashed in the repository. Prior-Knowledge Ledger S1 - revision v1.2 (2026-09-02 07:01 UTC) ---------------------------------------------------------- Two disclosures and a re-fixation. (1) The ledger's Section 2 (the CLOSED claim list) was reworded once on 2 September, between the first writing (v1.0, file sha256 edc6545f6781803caa6127db3e415053f6b960bf0b2ac53f6f9926219300abd4, archived) and the closure fixed in the block above; the closed digest 59af82be... published above is v1.1's and is byte-identical in v1.2. (2) v1.2, after a second hostile audit, revises only the surrounding sections: the contamination table now records that the Commission received the portability-is-not-exit proposition on 22 August 2026 itself (in a consultation contribution), plus three further dated contacts to watched bodies; provenance cells were corrected, one in the direction of MORE archived evidence than previously claimed. Corrections that touch Section 2's content are dated annotations in a new Section 5; the closed bytes are untouched. New pins: the region before Section 2 is now also digest-pinned in the enforcement script (56b4acf3e906534cb7342202f69f4abdaeb44a42ae21a706c7760bc942634d5d) and may only grow by disclosed addition; v1.1 bytes archived (sha256 5572d7f040e8c0b6482a21f7d444283a8ba5cc6412a9204ae5016f4a19288081); the v1.2 file digest is 8ec4454f3fe0c0601f21a2d4f3be4dcb63c540d1fc8fd245baa502e05acb5fd6. Prior-Knowledge Ledger S1 - annotation A4 / revision v1.3 (2026-09-02 07:29 UTC) --------------------------------------------------------------------- A re-verification requested by the principal investigator found that annotation A1 (fixed in v1.2 above) was itself one notch more absolute than the bytes it described: RN-002's opening sentence over-compresses, but the note's own body qualifies it twice, so the divergence is a headline compression corrected within the same published note, not a contradiction between the public record and the ledger. A4 records this and downgrades the proposed site correction from needed to optional. The pinned Section 1/2 region digests are unchanged; the v1.2 bytes are archived (sha256 8ec4454f3fe0c0601f21a2d4f3be4dcb63c540d1fc8fd245baa502e05acb5fd6); the v1.3 file digest is 24cde611a113393235641ec6999ed148dfc6ab48593a91b6ff89b8e6d29f1ec0. Scope note to the comparability entries of 1 September, and register freeze (2026-09-02 10:55 UTC) -------------------------------------------------------------------------------------- Two sentences above are scoped. The 23:00 Clarification's sentence "The superseded register states are archived and hashed" and the 23:12 Addendum's item (2) both speak of the two register states superseded ON 1 SEPTEMBER - the 0.3a bytes as frozen (sha256 09ac25c6...), and the later state that added the declaration in place while still carrying the 0.3a version string (sha256 24a07965...). Neither sentence extends to the revision row's own in-place byte-states of 1-2 September, which were overwritten and are not retained, as the row itself discloses event by event. Today the register was revised 0.3b to 0.3c, whose sole changes are the version string and a freeze line: no further in-place modification of the revision row is permitted, and any later change to the register is a new dated revision with its own row. The superseded 0.3b bytes are archived and hashed (sha256 17d5401614906ff839173f7e10c3e51b800c8df17a0547bde25727e971ecd788) and the 0.3c bytes are pinned in the enforcement script (sha256 cc873e2b84b3bf6187ce7d85d159161adde0eb2385eaa41fa60e97e08a10906d). Ecosystem groups: interim membership change and network view (2026-09-02 12:44 UTC) ------------------------------------------------------------------------ Two dated site facts, outside every frozen snapshot (the pinned Q3-2026 release CSV carries no group column and is unaffected). (1) Cursor (Anysphere) is tagged into the SpaceX group: the merger became effective 14 August 2026 per SpaceX's Form 8-K (accession 0001628280-26-056945, Item 2.01 - wholly owned subsidiary, $60.0B implied equity); the SpaceX group is now three scored platforms. (2) The rankings group view now carries, per group, a network map and table of selected UNSCORED entities under group operational control, plus a beneficial-ownership contour where the same ultimate owner holds entities outside corporate control; every row cites a source (SEC filings, annual reports, official pages); no aggregate group score is computed and member scores are never combined into a composite; descriptive member statistics are labelled as such. An enforcement check (scripts, audit item 43) asserts full coverage and per-row sources. Correction to the ecosystem-groups entry above (2026-09-02 14:46 UTC) ---------------------------------------------------------- The 12:44 entry described every row's source as "SEC filings, annual reports, official pages". That was over-broad: a small set of rows (about seven, e.g. Starlink's scale figures and Trust Wallet's post-spin-out ownership) cite encyclopaedic pages where no primary source was located; the page itself links each row's source so the class is visible per row. In the same revision the network was expanded (now 12 groups, 105 group entities plus 4 beneficial-ownership rows, every row and note source-linked and machine-checked), several rows were corrected against filings (Informatica restated to the announcement basis; Careem control dated 30 July 2026 per Uber's Q2-2026 10-Q; all four SpaceX Exhibit-21 entities now shown), and the map's scale is derived from the live dataset. The enforcement script now also requires sources on notes and validates in-group parent references against the scored universe. Count correction to the entry above (2026-09-02 15:19 UTC) ------------------------------------------------- The 14:46 correction itself miscounted: the network ships 105 group entities plus THREE beneficial-ownership rows (108 source-linked rows in total), not four - the enforcement script's own output (108 sourced rows) was correct throughout. Several row citations were also re-pointed in the same revision so that every quoted figure appears in the document it cites (Informatica to the May-2025 announcement; the three Deutsche Telekom operator revenues restated to the entity figures on the cited principal-subsidiaries note; Oracle's OFSS stake reduced to the filing's own words; the Salesforce subsidiary-list count restated by counting rule). The pattern that every correction pass must itself be audited is why these appends exist. Single-platform corporations, tier 1 (2026-09-02 15:37 UTC) -------------------------------------------------- The rankings group view now also carries, for twelve corporations where exactly one product is scored (Tencent, ByteDance, Sony, Samsung, Disney, Roche, Thomson Reuters, LSEG, SoftBank Group, Sea, GoTo, Alibaba), the corporate network around that platform: 72 source-linked rows drawn from the corporations' own filings (annual reports, 20-F/40-F, interim reports) plus official pages, with control nuances stated per corporation - VIE structures for the Chinese groups, the Samsung chaebol cross-holdings, family voting pools (Roche, Thomson Reuters/Woodbridge), and honest negative facts (GoTo's Tokopedia is TikTok-controlled at 75.01 percent; Ant Group is a 33 percent equity-method affiliate of Alibaba, not a subsidiary). No corporate composite is computed; the scored platform keeps its own B. The enforcement check covers these rows (platform-name resolution, per-row sources, sourced notes). Scope correction, and figure-fidelity now machine-checked (2026-09-02 16:08 UTC) -------------------------------------------------------------------------- Two corrections and one structural change. (1) The 15:19 entry claimed universally that "every quoted figure appears in the document it cites"; that was true of the four rows that entry re-pointed and false as a universal - a fourth audit found further figures (acquisition prices and dates on several rows) cited to documents that do not carry them. Those rows have been re-pointed to figure-bearing documents, given second sources, or de-figured. (2) The same audit found the row detail column was being suppressed wherever a headline figure was also present (ownership percentages disappeared from display); fixed. (3) Structural: a verification script now fetches every cited source and asserts that each numeric token shown on the page occurs in the fetched bytes (with a persistent best-fetch cache); the full set - both views' 180 rows plus notes and ownership headers - passes with zero failures as of this append. This closes the defect class that survived the previous rounds: source links were checked for liveness, never for content. Figure-fidelity gate rebuilt; scope of the 16:08 claim corrected (2026-09-02 17:15 UTC) ------------------------------------------------------------------------------- A fifth audit found the verification script shipped at 16:08 was itself unsound: it substring-matched (a "$2.1B" row passed on the string "2.1" inside a page timestamp), did not enforce HTTP status, could fall back to a stale local cache on a failed fetch, and fetched only 49 of the cited sources (it skipped any row without a numeric token). So the 16:08 statement that it "closes the defect class" was premature. The script has been rebuilt: it fetches every cited source with the HTTP code enforced (non-200 is a hard failure, no silent cache fallback), records each fetch's code and SHA-256 in a manifest, matches numbers by specificity (a four-significant-digit decimal may substring-match; short or round numbers require their currency or percent sign) rather than by bare substring, and additionally asserts that each entity's own name occurs in its source. Under the rebuilt gate, several further citation defects surfaced and were fixed (Fitbit price de-figured; Verily, Starshield, MariBank, SPX and LexisNexis Risk Solutions re-pointed to sources that name them; the Tencent minority note split so Spotify is cited to Spotify's own filing; Hulu and Foundation Medicine "100%" replaced with the sourced wording; a Thomson Reuters date and a Microsoft HoloLens date re-sourced). The full set - both views, 120 distinct sources - now passes with zero failures and is deterministic across consecutive runs. Verification gate rebuilt again, and the 2 September description corrected (2026-09-05 06:32 UTC) ----------------------------------------------------------------------------------------- A sixth independent audit demonstrated that the checking script described in the 17:15 entry of 2 September did NOT do what that entry said. Its matcher compared strings after removing all whitespace, so a figure could be "found" inside an unrelated number: the auditor pushed four fabricated ownership percentages about real companies through the script and it reported zero failures. The 17:15 description was therefore wrong in two checkable ways, and is corrected here. What the gate does now, stated so it can be tested rather than trusted. Text is normalised keeping word spacing, so numbers keep their delimiters. A figure counts as found only when it is not flanked by another digit, dot or comma. A percentage must appear either with its percent sign or in the two-decimal form ownership tables use; an integer alone never satisfies a percentage. A figure must also appear close to the entity it is claimed about, unless the document is that entity's own filing (measured by how densely its name recurs) or is short enough that proximity is meaningless. Every cited source is fetched with its HTTP status enforced, and failed fetches are recorded rather than silently replaced by an earlier copy. The script now carries its own adversarial tests - the four fabricated figures from the audit must be rejected and the corresponding true figures accepted - and those tests are part of the published audit script, which now reports 44 checks. Under the rebuilt gate several further citation defects surfaced and were fixed: three figures whose cited documents do not contain them (a stake percentage, an acquisition price and an acquisition year) were removed or restated; a percentage was corrected to the exact figure in the filing; two entities were renamed to the names their filings use; and a note that attributed an oversight role to a filing that does not describe one was restated. The full set - both views, 232 rows over 120 sources - passes with zero failures, and the audit now runs the gate's self-tests on every pass. Separately, the group view was made usable on phones (below a narrow width the map shows the scored platforms and their zones, with the group's other entities in the source-linked table below it) and is now reachable from platform pages and from the platform view, rather than only from a toggle. Correction: what the verification gate actually enforced (2026-09-05 12:00 UTC) A seventh adversarial review was run on a frozen tree, with the reviewer asked to break the gate described in the entry above rather than to confirm it. It broke it, and that description was wrong in two checkable ways. First, the entry said a figure must appear close to the entity it is claimed about, and named two exemptions. A third and larger one went unstated: a row with no entity at all - every note and every ownership header - was subject to no proximity test whatsoever, and those rows carry most of the ownership percentages a reader would rely on. About half of the figures the gate accepted were decided that way. Second, the entry said a percentage must carry its percent sign or appear in the two-decimal form ownership tables use. A one-decimal number was in fact accepted bare, which allowed a dollar amount to satisfy a claimed percentage. The review also found a defect neither entry could have described, because it was invisible in the output: a single mis-escaped brace in the pattern that extracts years compiled to match only years ending in the digit 2. Every other acquisition or closing year on these pages was never checked at all, while the summary line reported zero failures. Fifteen fabricated figures were pushed through the gate in that state, including a thousand-fold error in an acquisition price and an impossible future acquisition year. The gate has been rebuilt again. It now enforces the following, each exercised by a test that ships with it. A note or ownership header may not carry a figure unless it names the entity that figure concerns, and the figure must then sit near that name. A cited source that never names the entity cannot supply its figure, so adding a second, unrelated source no longer defeats the check. A currency amount must carry its magnitude, so a figure in millions cannot satisfy a claim in billions. A name that saturates a document - a short word recurring hundreds of times inside longer words - is not treated as a location. An entity must be named in its source by its name, not by any single word of it. Years are extracted and checked. Under the two waivers that remain, a figure must additionally sit beside language that gives it a meaning. Re-running the rebuilt gate over the same 232 rows produced eighty failures, all now resolved. Thirty-six notes and ownership headers now name the entity their figures concern. Twenty-eight rows record the name their source actually uses, rather than a compound label. Twenty-one rows and notes were changed in substance: claims restated to what the cited document says - among them a cumulative investment total the filing never states, a dismissal attributed to a complaint filed two years earlier, an ownership characterisation the source does not support, a corporate action the filing does not mention, and product names absent from the filing cited for them - and rows moved to sources that name what they describe. The rebuilt gate passes over these rows with no failures, and its self-tests now include the fabrications from both reviews, so a reader can run them rather than take this description on trust. The earlier statement that these views pass with zero failures was therefore true of the script as it stood and false as evidence about the data. This entry supersedes it. The count of rows and sources published here is unchanged; what changed is that it is now the outcome of a test that runs. Corporate networks: second batch, and hierarchy put under the same rule as figures (2026-09-05 12:53 UTC) Twelve further corporations that are represented in this index by a single scored product now carry a mapped corporate network: Broadcom, Intel, NVIDIA, Adobe, SAP, Intercontinental Exchange, S&P Global, Moody's, PayPal, Block, Booking Holdings and Spotify. With the twelve published on 2 September this layer holds 205 rows, and the two views together hold 313. The batch was researched, then audited against the cited documents before any of it was published, and the audit changed it. Ten claims were contradicted by the very document cited for them: a networking brand placed in the wrong segment; a majority-owned unit placed under a division the filing reports it outside of; a count of legal shells that was too low; a broker-dealer described as a product; a registry said to be absent from a filing that names it twice; a post-trade venture shown as part of a group that sold it in October 2025; a statement that one business was the only partly-owned one, contradicted by the same exhibit; a merged company said to have left no parent entity, which the exhibit still lists; a claim that a filing uses only an old division name, which it does not; and a note asserting that no official source confirms the status of two podcast studios, when a filing names them as wholly-owned subsidiaries. Thirty-two further descriptions were absent from the documents cited for them and were restated or re-sourced. None of that reached the site. The second change is structural. Until now the nesting drawn on these maps - which entity sits under which parent inside a group - was not checked at all, and a flat subsidiary exhibit, which gives only a name and a jurisdiction, cannot support it. Every parent link must now be one of three things: two names appearing together in a cited source; a verbatim phrase from a cited source that names the entity, checked byte-for-byte; or an explicit declaration that the grouping is this index's own classification, in which case the page says so on the row. Of the 79 parent links now published, 61 are evidenced in a cited source, three of them by quotation, and 18 are declared as this index's grouping and marked as such. Nothing is left as silent inference. Two limits are worth stating plainly. Co-location of two names is a necessary condition, not a proof of the relationship: an alphabetical subsidiary list can place two names beside each other without asserting anything, which is why several links that would have passed on proximity alone were demoted to declared groupings after being read. And the checker measures how distinctive a name is inside each document rather than relying on a list of common words, because a word that recurs every few hundred characters locates nothing, while a rare one locates as well as a full name. Correction: the hierarchy test did not hold, and two sentences above were wrong (2026-09-05 13:59 UTC) An eighth adversarial review, again on a frozen tree, pushed seventeen fabricated parent links through the check described in the entry above, three of them re-confirmed against freshly fetched documents rather than a local copy. It also showed that at least twelve of the sixty-one links then published as evidenced rested on nothing: a postal address, part of a web address, page metadata, a list of sibling products, a different company's name, and - in several cases - one word standing in for both the parent and the child at the same time. The cause was the test itself. Requiring two names to appear within eight hundred characters of each other treats an alphabetical subsidiary list, a trademark enumeration and a segment revenue table as if they asserted a relationship. They do not. Proximity has therefore been removed as a basis. A parent link is now evidenced only by a phrase that occurs verbatim in a cited document and names both entities; the phrase is checked byte-for-byte and is shown to the reader on the row it supports. Everything else is declared as this index's own grouping and marked on the page. The result changes the published picture, and it should. Of 79 parent links, 16 are carried by a quotation and 63 are declared as this index's grouping. Company filings simply do not often state which brand sits under which unit; they list entities and jurisdictions. The earlier claim that 61 links were evidenced in a cited source is withdrawn. A second sentence above must also be corrected. It said the checker measures how distinctive a name is inside each document rather than relying on a list of common words. Both are true at once: the checker measures density, and it also consults two hand-written lists of ordinary words. The sentence as written overstated the first and denied the second. Three further defects found by the same review are fixed. A figure inside a short ownership exhibit could be satisfied by a neighbouring company's number, because the window spanned about forty rows; figures are now bound to the entity's own row there, with a wider window allowed only where the surrounding words say what the figure is. A single word lifted out of a longer name was being treated as locating that entity, which let one company's segment revenue stand as another company's purchase price; only a whole name now carries that weight. And a quoted phrase was checked for the child's name but never the parent's, so a real sentence about one unit could license a claim about a different one. Withdrawal: the quoted-hierarchy tier is gone, and three sentences above were wrong (2026-09-05 15:04 UTC) A ninth adversarial review broke the rule announced in the entry above on its first attempt. Ten fabricated parent links passed the unmodified checker, every one of them built from real bytes in the documents those rows already cite. Among them: a sentence reading "our subsidiary Uber Freight acquired Transplace" used to place Uber Freight UNDER Transplace; a list of two competitors, one owned by a third company; a navigation menu; a citation headline sitting in a reference list; two consecutive rows of an alphabetical subsidiary exhibit; and the headline of the press release announcing that SAP had SOLD Qualtrics, used to place Qualtrics inside SAP. The fault is not in the wording of the rule but in what a byte test can decide. The phrase is chosen by the person entering the row. A machine can confirm that the phrase exists in the cited document and that both names occur inside it; it cannot confirm that the phrase asserts the relationship, and mentioning two names is not asserting one. The tier has therefore been removed rather than repaired. Every parent link on these pages is now declared as this index's own grouping and marked on the row: 79 of 79. The type that describes a row makes the declaration mandatory, so a link cannot be added without it. Relations that a document really does state are carried as ordinary sourced text on the row, where the figure checks apply. Three statements in the entry above are corrected. It said a quoted phrase is checked byte-for-byte: the comparison is made on normalised text - lower-cased, with tags removed and runs of spaces collapsed - which is exactly how two of the forgeries were assembled from fragments that are not adjacent on the page. It said the phrase is shown to the reader: it was carried in a hover tooltip, which a touch device never displays. And it said the checker does not rely on a list of common words while in fact consulting two such lists; one of those lists had by then stopped being consulted at all. The same review found three defects in the figure checks, all now fixed. Binding a figure to its entity's own table row was applied only to documents below a size threshold, so the five largest ownership tables fell back to a wide window: three ownership percentages were successfully moved between neighbouring companies in one subsidiary list. The binding now applies everywhere, and a figure is refused when another figure of the same kind stands between the entity's name and it - unless the row itself claims that other figure, or it is the same value repeated in a bilingual table. Counts carrying a unit - subscribers, customers, events, shares, entities - were not extracted as figures at all and so were never checked; they are now, and four claims that could not be paired to their entity in the cited document have been withdrawn or restated, including two that a bilingual or header-separated table makes mechanically unpairable. Finally, a download that timed out was being recorded as a successful fetch and answered from an older local copy; this happened four times in one day, once at the moment of publication. An incomplete transfer is no longer a fetch, and a source that yields no usable text now fails the rows that cite it. The published audit script also now runs the figure checks over every published row, not only its own self-tests: the previous arrangement reported a passing audit that had never re-checked the data. It reports 45 checks. Correction: the entry above claimed a fix that did not hold (2026-09-05 16:37 UTC) A tenth adversarial review put twenty-four fabricated figures through the unmodified checker, and among them was the very defect the previous entry announced as fixed: an ownership percentage belonging to one company in a subsidiary table was still accepted for its neighbour. That paragraph was therefore false when it was published, and this entry withdraws it. Since that file is hash-pinned and cited outward, the withdrawal matters more than the repair. What the review found, and what has been done. Five holes were in the extraction itself, each one a place where a figure displayed on these pages was never being checked at all: a magnitude the pattern did not know ("trillion") was silently dropped, so a claim in trillions could be satisfied by the same number in billions; a year at the end of a sentence was invisible because of a full stop, which defeated two of the checker's own pinned tests; a count with a qualifying word ("274 face-to-face events") was extracted only in its bare form; a share price of "$56.00" could satisfy a claimed "56%"; and a name ending in a full stop ("Carfax, Inc.") could never match its own document, so such rows fell back to a single word of their name. The attribution rule has been rebuilt around the property that actually distinguishes one company's figure from its neighbour's. A figure now counts only when it can be reached from an occurrence of the entity's own name with no other figure of the same kind standing in between — and, where the surrounding text is a table rather than prose, only when the figure follows that name and the name is written in full. Bare decimals in columns are what identify a table; prose states figures with their signs. Against the review's own swaps this now accepts every true figure and refuses every borrowed one. Three claims have been withdrawn from the pages rather than defended: two ownership percentages whose sources place the number where no rule can attach it to the company (a header block in one case, a bilingual column in the other), and a set of counts of exhibit rows that were our own tally rather than a figure printed in the document. Those counts are now stated in words, and the exhibits remain linked for anyone who wishes to count them. One further error of placement, which no checker had been asked to catch: a note about Roche's holding in Chugai was rendered inside the Thomson Reuters network. It has been moved, and the coverage checker now refuses a note whose subject belongs to another corporation while its text names nothing in the one it appears under. What the figure checker actually certifies, and what it does not (2026-09-05 18:18 UTC) An eleventh adversarial review put nineteen fabricated figures through the checker in nine classes, and it also showed that the entry above described the checker as doing more than it does. Rather than announce another repair, this entry states the boundary plainly, so that no future entry has to withdraw one. The checker certifies four things about every figure shown on these pages. The cited document was fetched in full — an incomplete transfer is not a fetch, and a source that yields no usable text fails the rows that cite it. The entity is named in that document. The figure occurs in it, in the currency and magnitude claimed: a sum in millions cannot satisfy a claim in billions, a price in one currency cannot satisfy a claim in another, and a count of things is not an amount of money. And, where the figure sits in a table, it must stand in the entity's own row: after the entity's full name, within one row's width, with no intervening text between the row's figures. That last property is attribution, and inside tables it holds. Outside tables it does not hold in the same way. In running prose the checker requires the figure to sit near a specific mention of the entity with no other figure of the same kind between them. That is a strong filter — it rejects a value taken from the next paragraph — but it is a heuristic, not proof. Three exemptions are deliberate and are named here rather than left to be discovered. Dates and counts are not subject to the interposition test at all, because a timeline crowds years the way a table crowds percentages and the rule would reject true acquisition years; they are checked for presence near the entity and read by hand. A row that states several figures about one entity may have its own other figures stand between the name and the one being checked. And in a document that is ABOUT the entity — its name recurring every few hundred characters — a figure is accepted on financial language nearby without positional proof; tightening that rejected eleven figures that are true and plainly stated, so the exemption is kept and pinned as a test that records what it lets through. The nineteen forgeries are closed. Among the defects behind them: a figure that followed the word "table" or "note" was skipped entirely and never checked; an alias supplied by the author could point a row at any anchor in the document; the table detector could be switched off by the claim under test; a claim in trillions was satisfied by the same number in billions; a share price satisfied a percentage; and a row's own figures could clear a path across a whole table. One withdrawal made in the entry above is also reversed: the audit showed the Bank Jago percentage IS pairable in its source, so it is restored rather than left withdrawn. Two further things are now visible on the page itself: where a source writes an entity's name differently from this index, that form is shown beside the row, because it is the anchor the check uses; and a note's second source is linked, which had been unreachable. Withdrawal, and an account generated from the code instead of written about it (2026-09-06 09:49 UTC) A twelfth adversarial review accepted twenty-one fabricated figures in eleven classes and showed that the entry above — the fourth attempt to describe this checker in prose — was wrong in four places. Each of the previous three descriptions was also withdrawn, each time after being tested. The pattern is the point: a description written by the same hand that wrote the code reproduces the author's intention, not the code's behaviour. This entry therefore withdraws the following sentences and replaces the method of description. Withdrawn: that a figure in a table "must stand in the entity's own row … and inside tables it holds" — the table test could never fire for a money figure at all, because the pattern meant to detect a column of amounts was written to match thousands separators that the checker itself had already removed; and a table printed with one decimal place was invisible to it, which let five neighbouring percentages be taken for one company. Withdrawn: that "a sum in millions cannot satisfy a claim in billions" — a claim in billions was satisfied by the same number in millions, and a sum in yen satisfied a claim in dollars. Withdrawn: that "dates and counts are not subject to the interposition test at all" — they were. Withdrawn: that the three named exemptions were all of them — three further paths skip a figure entirely and were not named. All of those are now fixed, and the fixes were tested against the review's own forgeries rather than described. But the fix that matters is to the method. The checker now records which rule accepted each figure, and the account below is printed by the checker itself: 98 row-scale 34 own-filing-waiver 32 table-row 29 prose-proximity+meaning 27 skipped:year-in-source-url 7 short-document 3 skipped:document-label 2 skipped:month-in-source-url 200 TOTAL accepted Read plainly: of the figures published on these pages, most are accepted because the figure sits within about a hundred characters of a specific mention of the entity with no competing figure of its kind in between; a further group because the figure stands in the entity's own row of a table, after its full name; a smaller group on proximity in prose together with language that says what the figure is; and a group inside a document that is about the entity itself, where a figure is accepted on financial language nearby without positional proof — that is the largest single exemption and it is named here, as it was before. Thirty-two figures were not checked against the body of their source at all: a year or a month that appears in the source's own web address, and three counts that follow the word "Exhibit" or "Note". Those are now counted and published rather than left to be discovered. One withdrawal from the entry above is itself reversed: an alias may no longer be any word the author chooses, because a review moved a percentage between two companies with one. An alias must now either share a word with the entity's name or be attested beside that name in a cited source. Two more withdrawals, and the account re-generated (2026-09-06 11:04 UTC) A thirteenth review accepted twenty-one fabricated figures in fifteen classes and found two statements in the entry above that do not hold. Withdrawn: "The entity is named in that document." Where a row records how a source writes a name, that form could stand in for the name entirely — so an invented company whose recorded form was "the Company" took the figures of a real one, because that phrase recurs densely in its filing. A recorded form is now admissible only if it shares a distinctive word with the entity's name or stands beside that name in the source, and the presence test is word-bounded, so a name that exists only inside a longer one no longer counts. Withdrawn: the count of figures that go unchecked. The account said thirty-two. It counted only the skips the checker knew it was making; figures the extractor never recognised at all were counted nowhere. They are now counted, and two of them turned out to be real claims that had never been verified — a founding date and a joint-venture split, both since restated to the words their sources use. The residue is published below as its own line. One further defect is worth stating because it was invisible in every previous account: the rule that accepted the largest share of figures required only that a figure lie within about a hundred characters of the entity's name — in either direction. A value printed immediately BEFORE a name belongs to the row above it, and one such figure was published as this index's own. That rule now requires the figure to follow the name, as the table rule already did, and the share of figures accepted by it fell by a third when the direction test was added. The account, printed by the checker as it stands: 70 prose-proximity+meaning 65 row-scale 37 skipped:not-extracted 35 own-filing-waiver 33 table-row 27 skipped:year-in-source-url 3 skipped:document-label 1 skipped:month-in-source-url 203 TOTAL accepted The audit now compares this table with the one the checker prints and fails if they differ, so the account cannot drift from the code again without something breaking. Three withdrawals, three defects that were weakening the checks, and three claims restated (2026-09-06 13:58 UTC) A fourteenth review found that the central new claim of the entry above is false against the code that entry describes. Withdrawn: "A recorded form is admissible only if it shares a distinctive word with the entity's name." Nothing in the checker made a word distinctive. Any word of three letters or more counted, so an invented parent called Halcyon Analytics Holdings, recorded as appearing in its source under the form "Moody's Analytics", was admitted on the shared word "analytics" and took that company's revenue. A shared word must now be four letters or more and must not be one of the ordinary words of business — analytics, services, markets, risk, cloud and the like — and the test covers every field a figure can be anchored to, not only the recorded form. Withdrawn again, in the other direction: the first repair refused ordinary business words outright, and that was also wrong. Such a word is sometimes the real name of a division — Risk at RELX, Mobility at S&P Global, Services at Apple. An ordinary word is therefore admissible as an anchor but can never buy the wide window: a figure leaning on one must stand in its own row beside it. The forgery above stays refused; the three real divisions stay published. Withdrawn: "That rule now requires the figure to follow the name." That is true of the rule that binds a figure at the scale of a table row, and false of the rule beside it, which accepts a figure on either side of the name and still does. Prose puts them either way — "announced on 15 September 2014 that it has reached an agreement to acquire Mojang" — and the entry above claimed the change for the rule that accepts the largest share of figures, which was not the rule that changed. Three defects found this round were making the checks weaker than the account of them, and each was invisible in the counts: A short document was searched for the row's name written out in full, gloss included. A press release that names Mojang eleven times was discarded whole, because the row records the company as "Mojang (Minecraft)" and that exact string appears nowhere. A rule meant to drop an anchor word that saturates a document could drop the entity's own name, if a word of the gloss happened to be tested first. The row was then left leaning on the gloss. A date appearing in a source's own web address ended the check instead of beginning it. That is the weakest evidence this instrument accepts — it shows when a document was published, not that it says what the row claims — and it was overriding the reading of the document itself. It is now a fallback, reached only when the text yields nothing. The three together were quietly withholding evidence from the checker rather than admitting bad figures. The number of published figures confirmed against the words of a cited source rose from 203 to 261, and the number accepted on nothing but a date in a web address fell from 28 to none. The stronger checks then found three published claims that their sources do not support as written. All three are restated in the words of the source: Binance. The index said the Irish holding company "holds the national operating entities — at least 24 of them per the CFTC complaint". The complaint says it "has directly or indirectly owned at least 24 corporate entities that have acted as Binance's digital asset and virtual asset service providers in a variety of jurisdictions and held Binance's non-U.S. regulatory licenses". The published row now says that. SoftBank. The index said LY Corporation is held "via A Holdings (50/50 JV with Naver)". The group's own financial report says it holds 50.0 per cent of the voting rights of A Holdings Corporation and appoints the majority of that company's board; it states no 50/50 split and does not name the other shareholder in that passage. The published row now says what the report says. Roche. The index dated the Novartis-stake repurchase to 2021. No text of the cited release carries that year next to the claim; the year came from the address of the release. It has been removed — the citation carries the date — and the voting figure it accompanied, 67.5 per cent for the founding families' pool, stands in the release word for word. Two further changes close the gap between what the checks do and what this record says they do. A table's rows are now separated by the names of the neighbouring entries as well as by legal suffixes, because a subsidiary table whose rows read "Twitch | 100" carries no legal suffix at all; giving one telecoms subsidiary the true revenue of the one below it in the same table is now refused. And the account below pins the file it describes: counts alone cannot say which data they were computed from, since one true figure exchanged for a false one leaves them unchanged. How each published figure was accepted (generated, not described): data: lib/group-contours.ts sha256:768fe48aeb9092de 87 row-scale 84 prose-proximity+meaning 57 own-filing-waiver 33 table-row 5 skipped:letter-adjacent 3 skipped:document-label 3 skipped:not-extracted 261 TOTAL accepted Eleven numbers in the published rows are not treated as claims at all, and the checker lists each one so the choice can be disputed. Three are the number of an exhibit in a filing. Three are parts of product names: Microsoft 365, HoloLens 2, Vision Fund 2. Two are the date on which this index recorded that a transaction had not yet closed — its own record, not a claim about a company. Three are the names of a reporting period, a marketplace and a product line: H1-2026, 1688 and Substance 3D. A statistical plan frozen before the data exist (2026-09-06 15:10 UTC) The single largest weakness of this instrument is that one person produced all 184 scores. The answer to that is not a better argument; it is other people scoring the same systems from the same rubric and a published account of how far they agree. Twenty-two systems were drawn by lot for that purpose on 22 August and frozen. The document that freezes them says what had to happen next: the statistical plan must be fixed before a single evaluator sees a single system. That has now been done. What is fixed: the three questions the study asks, kept separate because a favourable answer to one is habitually quoted as though it settled the others; the estimator for each, named individually; the scale on which agreement is judged; the interval procedure and its resampling unit; the thresholds, each with the consequence that follows for what this site publishes if the result falls below it; the rule that stops the study being reported at all if too much data is missing; and the commitment to publish the outcome whatever it is, including the outcome that the instrument is not reproducible. That commitment is easy to make today and would be hard to make in six months, which is the whole reason for making it today. Two choices inside the plan are worth stating here because each could have been made later, in the author's favour, and now cannot be. Agreement on the index is judged on a scale that is defined at zero, since a rater may legitimately score a system in a way that produces an index of zero and a logarithm would then have to be rescued by a constant chosen after seeing the data. And band agreement is reported by two statistics at once, not one: the conventional statistic collapses when a single category dominates, nine of the twenty-two systems sit in the top band, and choosing the friendlier of the two afterwards is exactly the manoeuvre a plan like this exists to prevent. The analysis is a script, not a description of a script. It is written in the plain standard library so that a third party can run it without installing anything, every estimator in it is tested against a case whose arithmetic can be redone on paper, and the study's precision was simulated before any evaluator was approached. That simulation settled the design and also corrected the author: the number of raters matters less here than the size of the cohort, which is frozen and cannot be enlarged, so this study will report a wide interval and will report it as a property of itself rather than explaining it away later. What this study cannot do is stated in the plan and is repeated here. It measures whether the instrument is applied consistently. It cannot show that the index measures what it claims to measure. Reliability is necessary for that and does not supply it, and the phrase "validates the index" will not be used about this study whatever it returns. The text is not published today; its digest is, which is what makes "written before the data" checkable afterwards. Nothing below can be edited without every one of these values changing. statistical plan sha256 7737200b26d3e9c257ac69f731ec362f1cddd477d6ccbbb354b8eabe93c4d97f analysis script sha256 136ace7a4490d73e426d89aa69b7e9da43c7ec70bd5fe06e3d94d3076673f145 rater pack sha256 39aa5a1d3e9425291918f2fe9469df5c8588b63c8714e3c8a91b524ca0ba85ef cohort freeze sha256 bdf63a818d0ce097badd738ca828fb5127ef58a501418c57fa846584abd92bfa dataset drawn from sha256 576c93ca3755aee652220ca507c5c0ca93db7db12441f31d8ad58630ce853265 The last of these is the pinned Q3-2026 release, already public at blackholeindex.com/bhi-scores-q3-2026.csv, and the seed for both the draw and the bootstrap is its first eight hexadecimal digits — a value fixed and published before any of this was contemplated, so that no number in the study was chosen by its author. What remains is not a document. It is evaluators, and that is a decision with a cost attached. Twenty-one more corporations mapped, and what the checker refused (2026-09-06 16:08 UTC) The ownership map covered sixty-two of the one hundred and eighty-four scored systems. It now covers eighty-three. The twenty-one added are the ones whose corporate perimeter can be read from a primary filing rather than reconstructed: eight financial institutions and market-data firms, five security platforms, and eight business-software platforms, each taken from the subsidiary exhibit its parent filed with the United States Securities and Exchange Commission for its most recent financial year. Every row cites that exhibit, and the checker read each exhibit's bytes. Two claims of the author's were refused by the checker during this work, and are recorded here because they would otherwise have been published. The first said that four of one platform's subsidiaries had been acquired within the two years to a stated month; a subsidiary exhibit records what is owned and not when it was bought, so the sentence had no source and the date is gone. The second gave a subsidiary a role invented from the resemblance between its name and the parent's; it now carries nothing beyond its jurisdiction. One further sentence was not written at all, for the same reason. A well-known consumer brand belongs to one of these groups, and the group's own exhibit names a legal entity while its annual report names the brand — but neither document joins the two. The entity is listed under its legal name with a neutral description, and the note beside it says plainly that no brand is mapped onto an entity here. It would have been easy to write the sentence everyone knows to be true. It is in neither source, so it is not on the page. The checker itself was corrected once. Its test for whether an entity is named in its own source searched for the name with any parenthesis stripped out. That is right when the parenthesis is a gloss and wrong when it is part of the legal name: it asked for a company that exists nowhere and so rejected three rows whose full names stand verbatim in the exhibits. The name as written now counts as well. That this was the third false rejection of the same shape in a week is why it is recorded here rather than quietly fixed. How each published figure was accepted (generated, not described): data: lib/group-contours.ts sha256:a22e9a7a68de71a1 92 row-scale 89 prose-proximity+meaning 57 own-filing-waiver 33 table-row 6 skipped:letter-adjacent 3 skipped:document-label 3 skipped:not-extracted 271 TOTAL accepted One hundred and one systems remain unmapped. Several have no subsidiary exhibit to read at all — a private partnership, a cooperative under central-bank oversight, a member-owned clearing corporation — and those will need a different kind of source and longer. Past the halfway mark, and three more of the author's sentences refused (2026-09-06 17:45 UTC) The ownership map now covers one hundred and five of the one hundred and eighty-four scored systems, up from sixty-two this morning. Twenty-two more corporations were added in this pass: laboratory and clinical-research groups, the two card schemes, a retail broker, four marketplaces, a delivery platform, three social platforms, and a set of software companies. As before, every row comes from the subsidiary exhibit its parent filed with the United States Securities and Exchange Commission for its most recent financial year, and the checker read each exhibit's bytes. Three sentences the author wrote were refused during this pass and are recorded because they would otherwise have been published. One dated a merger; the exhibit that was cited records what a company owns and carries no dates at all, so the year came from the author's memory rather than from the source and has been removed. One named a company that does not exist: the spelling of a payment subsidiary was corrected to the one word the filing uses. One added a corporation that already had a map, because the author did not check the coverage list before writing it — the duplicate was removed before anything shipped. None of the three would have been visible to a reader. That is the whole argument for a checker that reads the source rather than the claim: a false date and a misspelled company look exactly like true ones on the page. Some of what the exhibits show is worth stating plainly, because it is the kind of thing this index exists to make visible. One social platform's entire disclosed corporate perimeter is a single Irish company. One software company's subsidiary exhibit contains no subsidiaries at all and says so. A food-delivery platform now holds inside itself two European delivery companies that were once its competitors and each other's. A game engine's perimeter is half advertising. Two payment schemes of enormous reach have perimeters of four and eleven entities, cut by region rather than by acquisition — nothing like the software platforms of comparable size, whose exhibits read as lists of companies bought and kept alive as legal shells. How each published figure was accepted (generated, not described): data: lib/group-contours.ts sha256:7fdd557f91a0cec9 96 prose-proximity+meaning 94 row-scale 57 own-filing-waiver 33 table-row 8 skipped:letter-adjacent 4 skipped:not-extracted 3 skipped:document-label 280 TOTAL accepted Seventy-nine systems remain unmapped. The ones left are mostly the ones without a subsidiary exhibit to read: crypto protocols that have no corporate form at all, private companies, a cooperative under central-bank oversight, a member-owned clearing corporation, and several foreign issuers whose filings take a different shape. Those need a different kind of source, and the count will move more slowly from here. Every scored system now has a map, and eighteen of them are empty (2026-09-06 19:58 UTC) All one hundred and eighty-four systems in the release now carry an ownership map. This morning sixty-two did. The count is complete, which makes it worth saying exactly what completeness means here, because the honest answer is not that every system turned out to have a corporate perimeter. One hundred and forty of the maps are drawn from a primary document: a subsidiary exhibit filed with the United States Securities and Exchange Commission, the equivalent exhibit in a foreign issuer's annual filing, a company's own terms of service where it is private, or a regulator's complaint where that is the only document naming the parties. Every entity on every one of those maps was read out of the cited source by the checker, and every figure beside it was matched against the source's own bytes. Eighteen maps are empty, and they are the most informative part of this exercise. Some are empty because there is nothing to draw. The largest crypto asset in the index has no owning organisation at all: its own reference site answers the question directly, saying that nobody owns the network much as no one owns the technology behind email. Two entries are theoretical archetypes that name no real system and never did. One software company's subsidiary exhibit contains no subsidiaries and says so. One entry is a programme of a government aviation agency, so it has no shareholders and no exit at all; another is a regional transmission organisation whose members are the utilities on its own network. For these, emptiness is the finding. The rest are empty because the evidence is out of reach. A private partnership, a member-owned clearing corporation, a cooperative under central-bank oversight, several private companies and several foreign issuers publish no subsidiary list this checker can fetch: their own pages decline automated requests outright. Nothing is drawn for them and nothing is cited, because a claim about a structure that cannot be verified does not belong on a page whose whole argument is that its figures were checked. Where that is the reason, the note on the map says so in those words. The distinction between those two kinds of emptiness is stated on each map rather than left for a reader to guess, and it is the reason the count of maps is not being reported as an achievement. A map that says "this could not be verified" is worth publishing. A map that quietly implies a simple structure because the research was hard would not be. Two further things were fixed in the checker during this pass. A row whose figure carried no unit at all was going unchecked: a filing that prints a plain hundred in a percentage column is making a claim the moment the number is copied here, and ten such figures were being published without verification. A bare figure standing at the start of a row is now extracted and matched like any other, which is where the accepted count below rose from. The residue counter was then double counting those same figures as unextracted, and no longer does. How each published figure was accepted (generated, not described): data: lib/group-contours.ts sha256:32bde48cad23000a 115 prose-proximity+meaning 100 row-scale 64 own-filing-waiver 33 table-row 9 skipped:letter-adjacent 3 skipped:document-label 3 skipped:not-extracted 312 TOTAL accepted Twelve figures in the published rows are still not treated as claims, and the checker lists each one so the choice can be disputed: three are the number of an exhibit in a filing, and the rest are parts of product names or the date on which this index recorded that a transaction had not yet closed. Correction to the digest in the entry above (2026-09-07 04:50 UTC) The account published in the entry above pins the data file it describes, and that pin is wrong. It was captured, and the entry written, before one last edit to the file: a value had to be added to the list of relation types the file declares, because the build refused to compile without it. The edit was made, the build then succeeded, and the site was deployed — but the entry had already been written against the earlier state, so the digest it carries names a file that no longer exists in that form. The audit caught this on its next run, which is what the pin is for. It is recorded here rather than corrected silently, because a record that quietly acquires the right number afterwards is worth nothing. No row changed. The edit added a category name to a type declaration, not an entity, a figure or a source, and the counts the checker prints are identical to the ones published above, line for line. The correct pin is: data: lib/group-contours.ts sha256:9e28255b89829aa6 The lesson is narrow and worth stating: the account must be captured after the last edit, not before the last build. Capturing it earlier is exactly how a published description drifts from the thing it describes, which is the failure this pin was added to catch and has now caught once. A cited source moved, and the row it carried was wrong anyway (2026-09-07 05:00 UTC) The audit run that followed the correction above failed on a different row, and this one is not a bookkeeping matter. A cloud unit inside one of the Chinese groups was published under an English translation of its name. The page cited for it is the unit's own site, and that page was rewritten at some point since the row was written: what little of the name had appeared in Latin letters is now almost entirely in Chinese, so the checker could no longer find the entity in its own source and refused the row. Refetching showed something worse than a moved page. The name this map used appears nowhere at all — not in the current page, not in the Chinese text, not on the company's English site. It was a translation someone had made, and this index had published it as though it were the name. The row now carries the name the company uses, one word, and cites the site's front page where that name occurs. The old translation survives only inside the row's own description, which says what it was. The published account is regenerated with it. Nothing else moved: the counts are unchanged line for line, and the new pin is data: lib/group-contours.ts sha256:aa714955a5257649 Two failures in two consecutive audit runs, both caught by machinery rather than by reading, and both worth having. The first was a digest captured a moment too early. The second was a company named in a language its owner does not use for it, sitting on this map since the row was written and invisible until the source underneath it changed. Sent to the European Commission (2026-09-07 11:02 UTC) An information provision was sent today to the European Commission in connection with market investigation DMA.100236, the Article 19 investigation opened on 18 November 2025 into whether the obligations of Regulation (EU) 2022/1925 can address practices in cloud computing services. It went to the address the Commission's DMA team named in reply to a procedural enquiry of 1 September 2026 asking which channel was appropriate for unsolicited third-party research. That address is the general public mailbox for the Regulation, not a case-team address, so the letter asks that the document be forwarded and states its non-confidentiality in its first line, before anything needs to be opened. The document reports a structural measurement of switching capability and dependence across four cloud infrastructure providers under one provider-neutral frame frozen on 22 August 2026, before any provider parameter was scored. It is published in full at blackholeindex.com/BHI-Information-Provision-EC-DMA-100236-Cloud.pdf and hashes to sha256:b7f04cd6d3a663a99cd6c6dcd1d5da0d6598141aac70515fc82d6560fe6cdd4d. Two independent hostile audits were run against the draft before it went, and between them they found four claims that the underlying record does not support. All four were in the draft; none is in the document that was sent. They are recorded here because a submission to a regulator that quietly drops its errors is worth less than one that names them. The first was a measurement that was never made. The draft stated that the financial exit cost had been measured on a dated pricing snapshot and reported separately. The record says the opposite in its own words: the quantity is defined but no exit-cost figure has been computed, and the computation is a queued step. The sentence sat in the answer to one of the four questions the investigation was opened to ask. The document now says plainly that on fees and contractual terms it offers no measurement at all. The second was a universal negative that this project had already withdrawn. The draft said no provider documents an agent that changes a customer's estate without approval. A research issue recorded twelve days earlier had found a first-party page, inside the evidence cutoff, documenting a customer-selectable mode in which the agent acts without waiting for approval, and had concluded that the invariant was a default rather than an invariant. The draft repeated the frozen wording as though that note did not exist. The third was the word "published", used five times for a corpus that is private. The public file this record is part of says on its fourth line that snapshot contents remain private until quality assurance completes. A reader following the document's own citation would have found the contradiction in one click. The fourth was the exit criterion, paraphrased in a way that loosened it eightfold: the criterion binds the customer's engineering staff and a quarter of that headcount, not a quarter of the whole workforce, and the paraphrase had dropped the substantive half of the definition as well. Beyond those, the document that went names each provider in each finding rather than describing them anonymously beside a table that names them anyway; gives the scale and the direction of every parameter, which the draft omitted; and carries eight limitations the draft had left out, among them that the frame's own stipulations make the portability readings an upper bound, that no provider was contacted or given an opportunity to comment, that no page captures were retained, and that no dependency index, this one included, should carry decisive weight in a policy decision on its own. One process failure of our own is recorded with it. The letter carries a link to the document, and the document was published about a minute after the letter was sent rather than before it. The send sequence written for this despatch put publication first and verification of the live address second; that sequence was not followed. The window was short and the address now resolves to the file that was attached, byte for byte, but it was luck rather than method, and the method is the point. The site now dates itself from this file, not from its own build clock (2026-09-07 11:50 UTC) A footer on this site now carries two dates: the label of the release the scores belong to, and the date of the most recent entry in this record, linked to it. Both are content facts. Neither moves when the site is redeployed, and the second can be checked against this file in one click. The obvious way to build such a footer is to stamp the build time into the page. That was considered and rejected. A build stamp says the site was deployed, not that anything changed, and this site is deployed several times in a day for reasons no reader cares about — a checker fix, a renamed variable. A freshness claim that nothing verifies is the same defect this project spends its time removing from its figures, and it does not belong in the furniture either. The same reasoning removed something that was already here and wrong. The sitemap stamped every one of its two hundred and twenty-six addresses with the build time, asserting on each deploy that every page had changed. The search engine that consumes it states that it uses that value only where it is consistently and verifiably accurate; ours was neither, so the assertion is gone rather than corrected. The two ranking fields beside it are gone as well: the same documentation says plainly that they are ignored, and they encoded judgements about the relative importance of pages that nothing supported. Two files for machines were added. One tells crawlers what they may do, and says that nothing is blocked, including the crawlers that gather training data — because everything here is published under a licence that already permits copying and adaptation, and blocking a crawler while publishing under that licence would be a gesture without content. The other describes the project for language models in the format proposed for that purpose. No model provider has undertaken to read such a file; it is written on the chance that one does. That second file is a self-description, which is the genre this project has got wrong more often than any other: four hand-written accounts of its own checker were published and withdrawn in the space of a week. So it is not written. It is generated from the scored universe and from this record — every count and every date in it comes from the data — and the audit fails if the published file differs from what the generator produces. This entry is itself the test: appending it moves the date and the entry count, and the file that describes the project must move with it. Two errors on our own pages, found by checking a borrowed recipe (2026-09-07 12:22 UTC) A specification written for a different product, about being found and described accurately by search engines and assistants, was read against this site. Most of it transferred. Two of its factual claims did not survive checking, and checking it against our own pages turned up two errors of ours, which are the part worth recording. The first was ours and it was on the page describing the research paper. That page told search engines and social previews that the paper covers one hundred and eighty-four platforms across eighteen sectors. It does not. The paper reports a hundred platforms across twelve sectors and says so in a note on its own first page; the larger figures belong to the live index, which has moved past it. The description now says what the paper says. The second was the date the current release was published. It was written out by hand in three separate components, which is how three copies of one fact begin to disagree. It is now declared once and the three places derive from it. Two claims in the borrowed specification were wrong and are recorded so that nobody repeats them from here. It listed the crawlers of the large assistant providers, and omitted the two that actually govern whether a site can be cited in an answer — each provider distinguishes its search crawler from its training crawler, and it is the search one that decides citation. It also listed one provider's control token as though it were a crawler; that token sends no request of its own and only governs what may be done with content already gathered. What the site gained: structured description of the scored data as a dataset, with its licence, its creator identifier, the eleven measured variables and both published files, on the page that presents it; the citation metadata an academic index actually reads, which is not the structured data everyone reaches for first but a set of ordinary meta tags; and clean text versions of the paper and the two research notes, generated from the same files the pages render, so the two cannot come apart. Pages whose source is not already text have no such version, deliberately: a hand-made copy of a page is a second thing to keep true. Last, the validation roadmap. It carried a list of promises and, beside it, a separately written note saying which had been kept. Two hand-kept lists of one set of facts drift, and these had — the note still said a pre-registered study was pending after the study's plan had been frozen and its digest published here. Both lists are now one register, with a status, the date it was last assessed and the single fact it rests on for every promise, and the page renders from it. Of five promises made, two were not executed, one is prepared and not begun, and two are partly done. That is what the register says, and it is now the only place it is said. Seven pages told search engines they were copies of the home page (2026-09-07 12:51 UTC) Yesterday this project published a set of changes about being found and described correctly by search engines and assistants. Checking that work today, from the same borrowed specification, turned up a defect underneath it that had been live far longer and that none of yesterday's checks would have caught. A canonical link tells a search engine which URL is the real address of a page. Written once in the root layout, it is handed by the framework to every page that does not declare its own, with no warning of any kind. Seven pages declared none. Every one of them therefore carried — the instruction to treat the page as a duplicate of the home page and to attribute its content there. The pages were /robustness, the research-note index, both research notes, both CMA pages and the scoring editor. The same line sent an hreflang pair and an og:url to the home page from those pages as well. The four indexable ones are precisely the URLs this project asks to be cited by: /robustness is pinned by digest in submissions, and both research notes were listed in llms.txt the day before as citable addresses. They were advertised for citation while telling every crawler that they were not the address to cite. The other three ask not to be indexed at all, and a canonical on a page carrying noindex is a pair of conflicting instructions rather than a harmless extra. A second error, from yesterday and mine. Clean text versions of both research notes were generated and served, and neither page declared that its text version existed. Only the paper did. The check that was run yesterday was run against the paper, and its result was written down as though it covered all three. It did not. What changed. A canonical is now declared on the page it describes or not at all; the root layout declares none, and says why. Pages that ask not to be indexed declare none. Both research notes declare their text versions. The scoring editor is now marked noindex — nothing was exposed by its being crawlable, since its data sits behind a key checked on the server and the page shows only a prompt without one, but a login form has no place in a search result. A check of 133 assertions now enforces all of it, and it was tested by breaking six things on a copy and confirming each one was caught. The publication audit reads 54. A wrong number in a document already sent to a regulator (2026-09-07 13:55 UTC) The two pages carrying this project's submissions to the Competition and Markets Authority were marked noindex, with no reason recorded anywhere for it. They were absent from the sitemap and from llms.txt. So the one piece of independent corroboration this work has — the CMA published the July 2026 response in its record of responses to the Apple steering consultation, under "Feedback received", on its own servers, listed as "Black Hole Index" — was on this site and unreachable to anyone looking for it. That is the wrong way round: the check should be easier to reach than the claim. Opening those pages meant reading them properly first, and reading them properly turned up worse than a visibility problem. The April 2026 submission gives Netflix a B-score of 0.43 and calls it "near-zero lock-in". The instrument returns 1.00 on those same eleven parameters. 1.00 is the value in the paper published the month before the submission, in the released data file, and in the quarterly history for Q1-2026. The figure 0.43 appears nowhere else in this project and was never produced by any release. It was wrong when the document was written, and the document was sent to the CMA with it. At 1.00 Netflix sits on the Transition boundary, not near zero, so the argument built on it in section 8.3 does not hold as written. The document is left exactly as it was submitted. Editing it would make this site's own claim of faithful reproduction false, which is a worse fault than the one being corrected. The correction is published beside it, dated, and it is recorded here. Four further defects were found in the same pass, three of them ours. The response sent to the CMA in July tells the Authority it can verify a quarterly figure at a named address on this site. That address matched platforms by substring, so a request for "apple" returned Apple, Apple App Store and Apple Intelligence merged into one list — three different values for the same quarter with no field saying which belonged to which. The verification route offered to a regulator answered ambiguously. It now resolves exact identifiers only and returns the platform it resolved. The April page offered a "Save as PDF" button and hid both of its own warning notes from the printed copy. One click produced a clean, forwardable document carrying superseded figures, the wrong one among them, and no statement that the CMA never published it. The notes now print. The changelog said that submission had been "accepted" by the CMA. It was acknowledged: the Digital Markets Unit confirmed receipt on 13 April 2026 and nothing further was ever claimed privately. The word in public was stronger than the fact. Section 7.3 of the April document quotes a research group directly and gives no citation for it. The quoted phrases do not appear in that group's later published response, which this project holds. The quotation is reproduced as submitted, is flagged on the page as not re-verified, and should not be quoted onward from here. One correction to something published on this site earlier today. An entry made this morning said the CMA does not state the date 14 August 2026. It does: the update log on that consultation page records "14 August 2026 — Responses published". The claim was written after reading a summary of the page rather than the page itself, which is exactly the substitution this project spends its time removing from its own figures. The site now quotes the CMA's own line. What has not changed, and is now written next to every link rather than left implied: the CMA published all 37 non-confidential responses to that consultation, Apple's and Epic's among them. Publication means a response was received and published. It is not endorsement. No regulator has adopted, evaluated or expressed any view on this instrument, and no sentence on this site says otherwise. The copy the CMA published is not byte-identical to the copy served here: the Authority removed the author's contact e-mail from the letterhead and the signature block, and the text is otherwise identical word for word. Published copy sha256 926adba8f6a293d0518de5232ca70d2a607aa90ce7f3f72fbeb2b8f90a2caf49, 654,095 bytes, retrieved today. The publication audit now refetches that page on every run and fails if the listing, the quoted update-log line, or those bytes change. Six figures on one page had quietly stopped being true (2026-09-07 14:28 UTC) A decision was taken today that changes where the burden of correctness sits. Documents already submitted to an authority are not re-issued and are not followed by a correction letter. They stay reproduced exactly as they were sent, and when an error is found in one afterwards, the correction is published beside it and recorded here. That is the whole remedy. The consequence is that this site is now the only channel through which a correction reaches anyone, which makes staleness a defect rather than untidiness. So instead of reading the pages once and declaring them clean, the rule was written down as a check: every figure the site states in its own voice is either the value the data holds now, or it carries one line saying which release or cohort it belongs to. A figure whose justification nobody can write is a figure nobody should publish. Running it for the first time found six wrong statements, all on the executive brief, all in the present tense, none of them noticed by anyone who had read the page. Copilot was given as 4.73 where it is 5.23; ChatGPT as 1.96 where it is 2.25; Claude as 1.49 where it is 1.66; the ratio between the first two as 2.4 times where it is 2.3. The page said Claude had the lowest lock-in of the assistants, which stopped being true when platforms scoring below it were added. And it said only ten platforms sit below the point where capture equals escape, where twelve do. Every one of those was correct when written, against the Q1-2026 release. None of them was wrong in a way a reader would catch. They aged. They are not corrected by typing today's numbers in, which would only restart the same clock. The page now computes them from the released data when it is built, so the sentence and the figures cannot come apart again. Where a figure genuinely belongs to an earlier release — the paper's hundred-platform cross-section, the fourteen-platform reference cohort — it stays, with its reason recorded next to the rule rather than left to memory. The check does not cover the submission documents reproduced on this site. Those are historical artefacts, their figures are correct for the release they were written against, and where one is wrong the page carries a dated erratum instead. It also cannot tell whether a score belongs to the name printed beside it: it verifies that a figure exists in the release, not that it is the right one. That limit is stated in the check itself, and deriving figures from the data rather than typing them is the only real answer to it. The methodology page described a practice this project has not met (2026-09-07 16:17 UTC) Section 7 of the methodology page set out the scoring protocol in four points. Three of them were written in the present tense: that every parameter score is backed by a citation, that each score includes a prose justification explaining the evidence and its mapping to the rubric anchor, and that all scores, rationales and computed B values are published. Those three describe what the protocol asks for. They do not describe what has been published. The current release holds 184 platforms and 2,024 parameter scores. Exactly one of those platforms carries a stored rationale. None carries per-parameter citations, and the public interface returns neither. What is recorded is the eleven scores for each platform and, for every score changed through the quarterly process, the reason it changed. A reader can recompute B from the parameters and can see why a score moved. A reader cannot see the evidence behind a standing level, which is what those three sentences promised. This is a worse fault than any single wrong figure, and it is worth being precise about why. A wrong number is an error inside a method. A claim that the method is followed when it is not goes to whether anything published here can be checked at all — and it stood on the page that exists to tell a reader how to check. Where the protocol is met in full is in the determinations prepared for regulators. Each of those records, for every one of eleven parameters, the score, its evidence tier, the searches that were run, the sources retrieved, the primary source and the measurement provenance. The gap is not between the protocol and the project's capability; it is between the protocol and the dataset this site publishes as its main output. The page now separates the two: the four points state what the protocol requires, and a block beneath them states what is actually published, in the same words as this entry. The gap is also entered in the register of promises on the validation page, with its status and the date it was last assessed, so that it is counted rather than described. Nothing was withdrawn from the ranking and no score changed. The scores are what they always were: one assessor's structured judgement under a published rubric. What changed is that the site no longer claims more behind them than there is. This was found while assessing an unrelated proposal — to build a second universe from an externally defined list of the world's largest corporations. The assessment concluded that the proposal's value lies in its frame rather than in its five hundred scores, and that this gap should be closed first. That conclusion is a great deal less important than the gap.