Thesis 03
The Next Corpus Is Behind Closed Doors
Stack Overflow turned its first operating profit in the year to 31 March 2026 and its owner wrote the asset down to zero in the same accounts; two weeks ago a liquidated airline's internal record drew two qualified bids on a court docket. The corpus buyers now compete for was never published, and there are only three ways a closed door opens.
Two purchases, fourteen months apart, bought the same thing in two forms — and only one of them was ever for sale
In June 2025 Meta completed an investment in Scale AI, acquiring what its own Form 10-Q calls a "non-voting minority" of the equity, $13.79 billion of it booked as a non-marketable equity investment, and took Scale's founder in-house [1] — buying into the manufacturer of training data, outside this window and carried as an anchor. On 14 August 2026 Google was named successful bidder at $10,000,000 for the deidentified operational and employment record of a liquidated airline, with Mercor.io Corporation designated alternate at $7,500,000 — announced, contested and not approved, hearing 9 September 2026 [2]. That is buying the source: not what a company published, but the record it produced while working.
Settle one objection first. Is the airline's record not simply another archive, sold once and cheap? It is — and that proves the mechanism rather than breaking it. A living company's working record is not for sale at any price; it is the asset Thesis 02 called coordination capital, priced only as a flow while the company is alive. The only way to buy one whole is for its owner to die, and death re-prices it as an archive: $10m for 516 repositories, 372,585 commits, 43,170 pull requests with their review threads, 7,510,221,520 revenue transactions since May 2008 and — the part that matters most — the continuous-integration logs and test results that say which of those decisions worked [2]. Ten million dollars is not what that record was worth while the airline was flying. It is what it is worth now that nobody is adding to it. The price discovery in that auction, and the argument over who controls the de-identification, belong to Note 02.
Why buyers are at that door is visible on the other side of the ledger, and it is measurable — most sharply in the source every working programmer used.
The open corpus that taught a decade of programmers stopped being written: 207,189 questions a month at its peak, 1,346 in July 2026
Unit: questions posted per calendar month, UTC month boundaries · March 2014 – July 2026
Selected months, plotted on a linear scale with the baseline at zero: Mar 2014 207,189 · Mar 2018 172,891 · Mar 2020 155,906 · Oct 2022 105,892 · Mar 2023 87,176 · Mar 2024 44,727 · Mar 2025 15,086 · Jul 2026 1,346. Answers fell in step, from 317,154 to 3,137. Part of the decline is secular and began in 2014 with tighter duplicate-closing and moderation; what is not secular is the compounding acceleration — the March-on-March rate of decline went −29.3%, −48.7%, −66.3%, −78.6% across the four years to March 2026.
Executive synthesis
The price of a corpus follows its position on two axes — how much of it has been consumed, and whether anyone is still adding to it — multiplied by how much of it can be checked.
Raw public text became cheap in law and in economics at once. What buyers now pay for is narrower: a record whose outcomes were checked, a feed they must re-buy next year, and a corpus whose rows still know each other. Most of that supply was never published, because it was never meant to be read outside the company that produced it. Three numbers set the terms.
- $281m → $0 Carrying value of the goodwill on Stack Overflow in its owner's audited accounts, written off in full in the same year the business turned its first positive operating profit on revenue of $129m, up 12% Prosus FY2026 Annual Report, financial year ended 31 Mar 2026, published 27 Jun 2026 MEASURED
- $10m / $7.5m Two qualified bids for one company's deidentified operational and employment record, on the court record — the successful bid and the designated alternate ECF 1463, 14 Aug 2026 · status: announced, contested, sale not approved; hearing 9 Sep 2026 MEASURED
- 1.5 million Reinforcement-learning episodes across thousands of training environments reviewed in the final phase of one frontier model's training — the largest disclosed supply of this kind, and it was built, not bought Claude Opus 5 system card, 24 Jul 2026 COMPANY CLAIM
The honest one-line version, including the half that cuts against me. Supply is migrating from what was published to what was produced, and the third card is the best evidence that the migration has a ceiling: the largest disclosed corpus of checked outcomes in the window was generated inside a lab, not bought from anyone [4]. And the sell side of this market does not exist yet, which is why this is a channel and not a market.
Prior calls on this topic
This is the first thesis on data supply, so there is no scorecard to update — but Note 02 staked three dated tests on exactly this topic on 19 August 2026, and the series grades what it stakes. Status only, before any new argument.
| Test staked in Note 02 | Resolution date | Status, 24 Aug 2026 |
|---|---|---|
| No other bankruptcy estate markets a comparable corporate-data lot attracting two or more qualified bidders at $1m or more | 30 June 2027 | Open. A five-week review of claims-agent dockets and restructuring press from 15 Jul to 24 Aug 2026 found none MEASURED search-based negative, not a PACER census |
| Turing's Project Lazarus page still publishes a payment figure | 31 March 2027 | Open, not re-checked this cycle. Recorded here so the omission is visible rather than quiet |
| No buyer accepts a linkage-severing de-identification protocol without a price reduction | 31 December 2026 | Open. Both bidders in the 14 Aug auction signed the same referential-integrity clause, which is consistent with the test but does not resolve it MEASURED |
Verdict
"We've achieved peak data and there'll be no more." — and, in the same talk, "Pre-training as we know it will unquestionably end" and "We have to deal with the data that we have. There's only one internet."
Ilya Sutskever, OpenAI cofounder and former chief scientist — talk at the Conference on Neural Information Processing Systems, Vancouver, 13 December 2024, as reported by Kylie Robison for The Verge, 14 December 2024 [5] PRESS-REPORTED a reporter's transcription of spoken remarks; no published transcript exists
Refine — the stock of text did not run out; it stopped being what buyers pay for, because an archive amortizes into the weights while a checked outcome and a corpus nobody published do not.
The claim predates the evidence window, and one clause settles why it is adjudicated anyway: a verdict block adjudicates a standing public claim rather than an event, and this one was restated by its owner inside the window — "we're moving from the age of scaling to the age of research", 25 November 2025 [5].
An exhaustion frame predicts scarcity pricing across text supply. The filings show the opposite shape. Getty's "Other" revenue, which carries its data-access and AI-licensing agreements, fell 66.7% year over year in the quarter to 30 June 2026 [6]; Shutterstock's data offering fell 26% in the half beside a $173.7m impairment [7]; Taylor & Francis revenue fell 6.1% because a prior-year data contract did not repeat [8]; Reddit's long-duration licensing backlog fell 36% in six months [9]. Meanwhile a seller of the present added term: the New York Times' contracted licensing revenue for calendar 2027 grew 48% in the same six months [10]. That is amortization, not exhaustion. Epoch AI, whose 2024 estimates sit under the peak-data discourse, published both an exhaustion window of 2026–2032 and a 2030 supply envelope that does not bind [11]. The claim points at something real — the recipe stopped paying. Its framing, a stock that ran out, is wrong.
Inside the window the archive sellers shrank, one flow seller grew, and nobody published a realised per-use payment
Four issuers reported the same quarter in four ways, and the direction is not ambiguous. Getty Images' "Other" line — the bucket holding data access and AI licensing — came in at $5.238m against $15.716m a year earlier, the half at $13.9m against $25.0m, explained by the issuer as "lower volume and recognition timing of data access and/or licensing agreements" [6]. Shutterstock's Data, Distribution and Services segment fell 16% in the quarter and 28% in the half to $56.1m, the data offering alone down 26%, beside a $173.738m goodwill impairment [7]. Taylor & Francis revenue fell 6.1%, its parent restating growth "excluding non-recurring data contracts" and disclosing that the comparable half of 2025 carried "a significant non-recurring data contract" [8].
The counter-case is load-bearing and belongs here rather than in a footnote. The New York Times' contracted licensing revenue for calendar 2027 rose from $27m to $40m across two quarters while the company kept expensing "Generative AI Litigation Costs" as a separate line — $4.628m in the quarter, $8.840m in the half [10]. Litigating and licensing are complements for a party with pricing power. Wiley's AI revenue grew on the same pattern, bought by vertical applications rather than frontier labs [12].
Nobody stopped buying. Nobody is committing for long: the backlog shortens while the near tail holds
Unit: USD millions · 31 Dec 2025, 31 Mar 2026 and 30 Jun 2026, from six filings
Accent segments are the portion of each balance the issuer expects to recognise in calendar 2027; the full bar is the total obligation on contracts originally longer than one year. Reddit's total fell 36% in six months while its 2027 tail grew $5.3m across two quarters. The New York Times Company's totals are read from its disclosure tables and rounded to the nearest million.
Beneath the licensing tier sits the metered tier, and it is where the disclosure stops.
Four ways to price the same text, and only two have ever shown a realised payment
Unit: none — presence or absence of a published realised payment · read 24 August 2026
| Structure | Named instances | Has anyone published a realised payment? |
|---|---|---|
| Lump-sum archive | Getty, Shutterstock, Taylor & Francis, the 2024 Reddit agreements | Yes — in audited filings, and that is how we can see them shrinking [6][7] |
| Recurring feed or grounding licence | The New York Times, Wiley | Yes — as revenue and as remaining performance obligations [10][12] |
| Per-fetch toll | TollBit; Cloudflare's pay-per-crawl, still closed beta in its own documentation as of 28 Jul 2026; Stack Overflow adopted it on 19 Feb 2026 with no price disclosed | Only in the press — "hundreds to tens of thousands a month" for roughly 20% of about 7,000 participating sites, press-reported and unconfirmed by any party [13] |
| Per-use or revenue share | Microsoft's Publisher Content Marketplace (pilot, 3 Feb 2026, no dollar figure); ProRata's Gist 50/50 split; Perplexity's Comet Plus $42.5m pool at 80/20; Cloudflare's pay-per-use partners Ceramic.ai and You.com | Never, by anyone. A rate is not a payment, a pool is not a payout, and a split is not a sum [13] |
The mechanism: five kinds of source, two multipliers, and a price that follows position rather than volume
Treating supply as a token count is what produced four years of wrong forecasts. Sort every source by how much of it training has already consumed and by who is still adding to it, and a fifth kind appears with no natural origin at all.
Five kinds of source, sorted by what is left and by who is still adding to it
Unit: none — a classification, not a measurement · as of 24 August 2026
| Category | Consumed × still producing | Anchor evidence | What a buyer is actually paying for |
|---|---|---|---|
| 1 · Dead archives | Fully consumed, no longer produced | Getty "Other" −66.7% year over year; Taylor & Francis' non-recurring contract that did not repeat [6][8] | A one-time transfer into the weights. Once inside, it does not need buying again |
| 2 · Production-collapsed | Fully consumed, production near zero | Stack Overflow: 1,346 questions in Jul 2026 against 105,892 in Oct 2022; goodwill written to zero in a profitable year [3][14] | A shrinking option on a stock that stopped growing |
| 3 · Still producing | Heavily consumed, still generating | The New York Times' 2027 tail +48% in six months; Reddit's 26 billion posts and comments and its measured scarcity of machine text [10][15] | Duration and freshness — and, for one seller, the only large licensable place humans still go |
| 4 · Behind closed doors | Never consumed, produced continuously | 516 repositories, 372,585 commits, CI logs and 7.5bn transactions in one estate's asset schedule [2] | Process with recorded outcomes — the input reinforcement learning pays for |
| 5 · Manufactured for AI | Produced on demand, to specification | Expert labour at $85+/hr platform average; environment contracts at six to seven figures a quarter; 1.5m in-house RL episodes [4][16] | Difficulty that can be checked — and the vendor's margin, until the buyer builds it in-house |
Rows 1 and 2: what a consumed source is worth to its own owner
The cleanest evidence here is an owner's arithmetic about its own asset. In the year Stack Overflow first turned an operating profit, its owner wrote the remaining $280m of goodwill to zero and $54m of brand and customer intangibles off in full, because "no future cash inflows are associated with them", attributing it to "the disruption caused by generative AI", and cut the forecast average revenue growth behind that test from 16.7% to 6.6% [14]. Current cash improved and the future was priced away. No second bulk licence has been announced since early 2024 — an absence rather than a refusal — and the company has moved from bulk sale to metered access to an attempt to restart production itself [3]. The seller whose flow survived is meanwhile reported to be seeking a five- to eightfold increase on its roughly $60m-a-year agreement expiring in the first half of 2027 PRESS-REPORTED [9]. Pricing power tracks the liveness of the flow, not the size of the archive.
Row 3, and the one claim about Reddit I will make
Reddit's position is narrower than the version in circulation. It is not the last place humans talk — group chats and private servers carry more, none of it licensable. It is the only large place humans go that is. Machine-classified articles rose from 2.2% of newly published open-web articles in January 2020 to 51.7% by May 2025, while a deliberately conservative estimate across 51 subreddits over 2022–2024 finds machine text "marginally present", peaking "of up to 9%" [15] — an order-of-magnitude gap across two methods, and the Reddit measurement stops before agents became common. Against it: unauthorised researchers ran 34 synthetic personas through one subreddit for four months undetected PRESS-REPORTED [15], and logged-in daily unique growth in the United States has run at 1% for two quarters while the company stops reporting that split from the third quarter of 2026 [9]. A refuge that grows as a reading surface and stays flat as a writing surface is a weaker asset than its licence prices imply.
Row 4: the three doors
A closed door opens three ways, all three documented inside this window. The dead: an estate sells under court supervision, which produced the $10m and $7.5m bids [2]. The dying: a voluntary liquidation sells with no court order at all, at $10,000–$100,000 a lot, through a channel whose boilerplate covers "from codebases to databases to workspace data" [17]. The living: a contract — a buyer's page inviting operating companies to sell operational data at published six- and seven-figure tiers [18]; a lab's development-partner programme taking a bounded slice, "only your Claude Code input and output tokens from the first-party Claude API are provided to us"; and beneath both the default-on wave in the consumer and coding-assistant terms dated in section 07 [19]. Court, liquidator, contract: the mechanism changes and the direction does not.
What makes an estate corpus legible as a purchase is its vendor list — SAP, UKG, Navitaire, Coupa, ServiceNow, system by system — and its delivery clauses: "bare Git repositories or Git bundle files", "JSONL for metadata (commits, PRs, Issues, comments)", and "standard, easily restorable, and machine-readable formats" for the business systems [2]. A contract shows what a buyer asked to receive, not what it intends to do; the buyer's own words on the court record are "which can be helpful in improving our products and AI models" — improving, not training [20]. Two details outweigh the price. The alternate bidder's contract reproduces the schedule including the column header "Google's Data Purchase Request". And its Exhibit C, "Data Extraction", present in no other version, obliges the seller to transmit exports "to Buyer through Buyer's secure upload facility at data.mercor.com" [2]. A human-data manufacturer put its intake address in a bankruptcy exhibit.
Multiplier one: reinforcement learning pays for outcomes that can be checked
Four steps, one primary each. Reinforcement learning needs outcomes that can be checked, in the millions of episodes [4]. What is worth taking from a deployed model is behaviour under tools rather than prose: the lab that traced 16.55 million exchanges through fraudulent accounts named the targets itself — "agentic reasoning, tool use, and coding" [21]. A lab writes a bespoke contract only for an input it cannot generate, which is how narrow that development-partner scope is [19]. And a working company's record already carries outcome-labelled process: pull requests attached to commits, attached to test results, attached to a deployment that either held or did not [2]. The labels arrive free with the data, because one business produced both.
Multiplier two: synthetic generation multiplies volume; it does not replace supply
Synthetic data is usually presented as the alternative to buying data. It is closer to the opposite: a template is worth more when you can stamp it a thousand times. Phi-4 ran roughly 40% synthetic plus 15% web rewrites of about 10 trillion tokens in December 2024 — out of window, and still the only published mixture table at that detail; one public post-training corpus is 100% synthetically generated responses; and the largest study of the question, over 1,000 models and 100,000 GPU hours, finds roughly 30% rephrased synthetic optimal with a five- to tenfold speed-up [22]. What generation cannot do is invent the template. The market agrees: no current frontier system card discloses buying synthetic data from a third party, and the category's best-funded European independent wound down inside the window [23].
The scarcities, ranked, and the one that is false
First, verification — difficulty that a tool can check. This binds hardest through 2027, and it is my call. Its cost line is a credential-bound wage that rises with task difficulty: $15–20 an hour for entry annotation, $85 and up as a platform average, $100–200 and up at the frontier PRESS-REPORTED wide dispersion; no vendor publishes a rate card [24]. Its price ladder is the only one in public: six to seven figures a quarter per contract, $200–$2,000 a task, about $20,000 for a website replica and up to $300,000 for a complex product replica, exclusivity at four to five times non-exclusive [16]. The correction I owe my own draft: the same source records "substantially more in-housing" to avoid vendor margins. The cost is permanent; the vendor is not.
Second, freshness — the obligation to buy again. The only scarcity with disclosed, recurring, auditable prices; the balance-sheet split in section 03 is this scarcity in numbers.
Third, linkage — rows that still know each other. Real and priced: both bidders signed the identical clause requiring de-identification "while preserving referential integrity across the data set" [2]. It ranks third because it commoditizes — mature vendors already sell referential-integrity-preserving masking as a standard enterprise product [25], and two independent buyers accepting identical wording is what a standard term looks like. What is fresh is the mirror rather than the clause: two buyers' revealed preference, not one seller's marketing.
And the false scarcity: raw token volume. Epoch AI published both halves of the contradiction in 2024 — roughly 300 trillion effective tokens of human public text with an 80% confidence interval for exhaustion between 2026 and 2032, and roughly 500 trillion indexed plus 400 trillion multimodal supporting 6e28 to 2e32 FLOP through 2030 FORECAST [11]. Labs behaved as though a wall existed years before that arithmetic said one should, and the collapse literature does not rescue the volume story either: its three escape conditions — accumulate rather than replace, verify, rephrase — are the three every lab already runs [22]. Volume was never the binding term. Position was.
The legal lane, and two side passages
The $1.5bn Bartz settlement received final approval and judgment on 20 July 2026 [26]. The bifurcation underneath it — training on lawfully obtained books is transformative, downloading and retaining pirated copies is separately actionable — is a district-court holding, not circuit law, and the appeal that would have tested it fell with the settlement. The sum is adjudicated and collected in instalments to September 2027; it has not been paid. The doctrine is unsettled the other way too: Thomson Reuters v. Ross Intelligence was argued in the Third Circuit on 11 June 2026 and rests on harm to a market for licensing training data [26]. Liability attaches to the route rather than the content, so the cheapest compliant supply is a corpus acquired through a transfer approved in advance — and the voluntary rail delivers that with no court order [17].
Two smaller axes run alongside and point the same way. Openness commoditizes whatever it reaches: open weights disclose almost nothing about their inputs, and open data comes from outside the frontier — 10 trillion published pretraining tokens from a chip vendor, 9.3 trillion from an institute, 15 trillion of filtered web text from a platform — while the industry's measured transparency average fell from 58 to 40 out of 100 between 2024 and 2025, training data its most opaque dimension [27]. That vendor's own vice president of generative AI software stops one word short of the economics: "The model is the byproduct. It is not core to our business, which allows us to just open up the data, open up the recipes, open up everything" [27]. Calling that deliberate price suppression is my inference, not their claim INFERENCE. Free corpora ate the undifferentiated middle: only about 40% of 94 tracked licensing agreements now include training rights. The distillation axis runs the argument backwards: the deployed model became a raw material, and all three major United States providers now fence chain-of-thought behind encrypted blocks, a fence independent researchers priced on 10 August 2026 at roughly $720 per 10,000 traces [28].
Value map: the priced layers are running off, and the layer that would make a market has no sell side
Five layers sit between a company's record and a model's weights, and one of them is barely occupied.
-
Rights holders and archivesCONTESTED
Stock libraries, publishers, back catalogues — the sellers of what was already published
priced, disclosed and running off, as section 03 measures. In music, the counterparties took product control and equity instead of a per-unit price: no per-track or per-model rate has been disclosed in any AI music licence MEASURED absence
-
Flow and grounding sellersCLAIMED
News wires, newspapers, scholarly publishers — sellers of what happened today
the only sellers adding contracted term inside the window, and the ones with litigation budgets to match [10]
-
Human-data manufacturingCONTESTED
Expert labour, annotation, reinforcement-learning environments
labs are insourcing the control plane and outsourcing the payroll — and the insourcing is now stated by the vendors' own analyst [16][24]
-
Data-asset intermediationOPEN
Whoever stands between a company that holds a record and a lab that wants one
productized on the buy side only; nothing on the sell side MEASURED search-based negative, 24 Aug 2026
-
Environments and verificationOPEN
Sandboxes, graders, task sets — the machinery that turns a record into a reward
concentrated and unpriced in public; the only lab-published count is internal [4]
The fourth layer deserves the space, and the observation is not an indictment. An intermediary layer is forming, and every piece of it is on the buy side: one manufacturer's contract names its own upload facility in a bankruptcy exhibit [2], and two publish intake tiers of $100,000 to $1m-plus with referral fees up to $100,000 per company introduced [18]. A finder's fee is what a buyer pays when supply does not meet demand.
On the sell side there is nothing: no restructuring adviser, investment bank or claims agent markets corporate data corpora as a distinct offering anywhere in public as of 24 August 2026 MEASURED search-based negative; a private mandate would not appear. The market is buyer-pushed rather than seller-pulled — a channel, not yet a market. What a company here must own to survive annexation is neither the catalogue nor the rights but the request path or the labour graph: the flagship independent AI-data marketplace was absorbed by a network vendor that already owned the request path within about twenty months [29]. Whether data-asset intermediation becomes a durable category is an open question.
Where the economics work: five positions with real prices, and the risk inside each
These are not sectors but positions in the taxonomy above where money has actually moved. Companies named are publicly documented examples, not endorsements; I hold no position in, and no client relationship with, any of them.
Vertical corpora with a captive downstream application. Wiley reported $49m of AI revenue in the year to 30 April 2026, up 23%, lifetime past $110m, anchored by IQVIA and OpenEvidence [12]. What makes the position work is that the buyer is an application rather than a lab: it needs the corpus to answer questions today, so the licence has to renew. Risk: the buyer is one integration away from becoming a reseller, and the seller cannot see it coming.
Flow sellers with retrieval value. The New York Times Company added 48% to its 2027 contracted licensing in six months [10]. People Inc reported "Licensing and other" revenue of $47.0m in the second quarter of 2026, up 23.3%, a bucket that also holds brand licensing, so the AI share is unverified. News Corp's arrangement with Meta was reported at up to $50m a year for at least three years PRESS-REPORTED first reported by a title the seller owns [30]. Risk: every one is a renewal risk, annually.
The voluntary-liquidation channel. At the marketplace launched on 16 April 2026, a named seller — a thirteen-year transcription company — sold its internal messaging, email and issue history for "hundreds of thousands of dollars" [17]. Risk: two to three orders of magnitude below the court rail, no court order attaches — and the AI share of those closures is falling across three periods (17.7%, 15.9%, 14.4%), so this is not a rising tide of AI-company corpora.
Human-data firms that own a labour graph rather than a labelling tool. Both in-layer acquisitions in the window bought a faster way to acquire humans, not better tooling: Handshake–Cleanlab on 28 January 2026, whose cofounder said competitors "frequently use Handshake's platform to source human experts", and Labelbox–Upcraft on 10 February 2026, terms undisclosed in both. Surge reportedly runs above $1bn of revenue with roughly 110 employees PRESS-REPORTED [24]. Risk: software multiples on a staffing profit-and-loss, and a customer base that has started building the capability internally.
Transfer governance as a product. Three vendors now sell what the two bidders contracted for: Integral markets "entity-preserving" de-identification, Perforce Delphix claims deterministic masking that "ensures referential integrity across enterprise data estates", and Tonic.ai claims products that preserve "the relationships between tables" [25]. The honest caveat: the latter two are positioned mainly as test-data-management vendors, which strengthens the third-place ranking in section 04. Risk: a line item, not a business.
Market structure: twelve months in which the buy side productized and the sell side did not appear
Status matters more than announcement volume here: an auction result is not a sale and a pilot is not a price, so every row states what had actually happened.
-
28 Aug 2025
Anthropic updated its consumer terms so that Free, Pro and Max sessions became a training input by default, in its own words "including when you use Claude Code from these accounts" [19]
Shipped — policy in forceThe living door opens by contractCOMPANY CLAIM
-
15 Jan 2026
Cloudflare announced the acquisition of Human Native, the flagship independent AI-data marketplace, about twenty months after its launch [29]
Completed — price not disclosedThe request path buys the marketplaceCOMPANY CLAIM
-
28 Jan 2026
Handshake acquired Cleanlab, whose cofounder stated that competing vendors source human experts through Handshake's own platform [24]
Completed — terms not disclosedConsolidation at the manufacturerCOMPANY CLAIM
-
3 Feb 2026
Microsoft announced the Publisher Content Marketplace, a per-use arrangement between publishers and model providers, with no dollar figure attached then or since [13]
PilotR03.3 cohort member, still at zeroCOMPANY CLAIM
- 19 Feb 2026
-
23 Feb 2026
Anthropic disclosed roughly 16.55 million exchanges through about 24,000 fraudulent accounts attributed to three named actors, targeting "agentic reasoning, tool use, and coding" [21]
Published disclosureThe model becomes a source worth fencingCOMPANY CLAIM
-
4 Mar 2026
MOSTLY AI, the best-funded independent synthetic-data vendor in Europe, ceased operations after almost nine years; its founder attributed the failure to weakening privacy regulation, not to competition. Liquidation was dated 14 March 2026 [23]
Wound downThe compliance-synthetic market was never pointed at labsPRESS-REPORTED
-
16 Apr 2026
SimpleClosure launched Asset Hub, a marketplace for the assets of voluntarily dissolving companies, with workspace data in beta and buyers described as "AI labs, and reinforcement learning researchers" [17]
Launched — ~100 transactions reported at $10k–$100kThe dying door, with no court orderCOMPANY CLAIM
-
24 Apr 2026
GitHub's updated Copilot policy took effect, enumerating individual-tier interaction data as a training input by default [19]
Effective — announced 25 Mar 2026Behaviour under tools, collected at sourceCOMPANY CLAIM
-
27 Jun 2026
Prosus published its 2026 annual report: Stack Overflow's goodwill written down from $281m to nil, $54m of brand and customer intangibles written off in full, and the forward revenue-growth assumption cut from 16.7% to 6.6% — in the year the business first turned profitable [14]
Reported — audited financial statementsAn owner prices its own consumed corpusMEASURED
- 7 Jul 2026
-
20 Jul 2026
Final approval and judgment entered in the $1.5bn Bartz settlement; no class member or objector appealed within the thirty-day window, which closed 19 August 2026 [26]
Entered — adjudicated, collected in instalments to Sep 2027, no effective date announced as of 23 Aug 2026A price on the acquisition stepMEASURED
-
24 Jul 2026
Anthropic published the Claude Opus 5 system card: roughly 1.5 million reinforcement-learning episodes across thousands of environments in the final training phase, and four named data categories, none of them first-party agent usage [4]
PublishedThe largest disclosed supply is self-generatedCOMPANY CLAIM
-
14 Aug 2026
Auction results filed in the Spirit estate: Google successful at $10,000,000, Mercor.io Corporation designated alternate at $7,500,000, both agreements carrying the identical referential-integrity clause [2]
Announced, contested — sale not approved; hearing 9 Sep 2026The dead door, priced twiceMEASURED
-
19–20 Aug 2026
NVIDIA was reported to be in talks to invest in Mercor at a valuation of roughly $20bn — the candidate second instance of the Meta–Scale form [24]
Talks — preliminary, not agreed; no party has confirmed termsA dated falsifier, not an eventPRESS-REPORTED
-
21 Aug 2026
Springshot filed a limited objection in the Spirit case, contesting the ownership chain of part of the data and pointing at the resale right in the sale agreement [20]
Filed — unresolved at publicationWho owns a jointly produced recordMEASURED
Four paths, and the one I hold
Probabilities are my judgement, stated so they can be graded later; each path names one leading indicator mapped to a signal in section 10.
Two rails, both thin (~40%). The court rail stays a single-transaction curiosity and the voluntary rail keeps compounding at $10,000–$100,000 a lot, while the real supply arrives through contracts nobody files. I hold this path: it is the only one consistent with a productized buy side, an absent sell side, a five-week docket review that found no second estate lot, and a frontier lab whose largest disclosed corpus of checked outcomes is its own. Indicator: R03.1.
The channel institutionalises (~25%). A sell side appears — a restructuring adviser or claims agent markets corporate data as an asset class, comparable prices become citable, and diligence on corpus provenance becomes a standard workstream. This is the fastest way this thesis becomes obsolete, which is why CALL B is staked against it. Indicator: R03.6.
The labour gate closes it (~20%). Employee-data conditions attach to these sales — a carve-out, a review protocol, a use restriction — and shave the part of the corpus that made it worth buying. The union objection in the Spirit case is the first instance; a second, in a different estate, would make it a rule. Indicator: R03.1, on its conditions arm.
Null path (~15%). None of these layers becomes independently investable: labs keep building environments in-house for the margin reason their own analyst records, intermediaries get absorbed by whoever owns the request path, and the priced market stays a salvage market. In this path the mechanism is right and the investment implication is empty, which is a distinction worth stating plainly rather than blurring. Indicator: R03.5.
What others say, and where I differ
Pablo Villalobos and colleagues at Epoch AI — the stock-of-text estimates (6 June 2024 and 20 August 2024, both outside this window). Converges: the right question is a supply question. Departs: Epoch's own two reports give incompatible answers, and the organization's 2026 output has moved to compute and inference [11]. Side by side, the volume frame stops being the interesting one.
Shumailov and colleagues, "AI models collapse when trained on recursively generated data" (Nature, 2024, author correction 2025) — the model-collapse canon. Converges: recursive generation without correction degrades. Departs: Gerstgrasser and colleagues show a finite error bound when data accumulate rather than replace, Schaeffer and colleagues count eight conflicting definitions of collapse, and the largest study to date finds no degradation for rephrased synthetic at foreseeable scales [22]. Collapse is not why supply is tight.
Cloudflare's content-independence framing (1 July 2026). Converges: crawling is a poor proxy for value, and per-fetch metering cannot separate value from waste. Departs: the vendor sells the toll it measures, and its two same-day publications disagree on the denominator of their own re-fetch figure — the blog says "good bots", the release says "AI crawlers" COMPANY CLAIM no method, window or sample published [13].
Sacra's read of the human-data layer — gross margins of 30–40% and roughly 30 times annualised net revenue. Converges: the take rate is the tell, and the gross-to-net gap is where the analysis starts [31]. Departs: these are third-party estimates rather than disclosures, and the layer's attempted escape — buying corpora outright rather than reselling hours — changes the margin question rather than the multiple.
The intangible-asset accounting canon. IAS 38 states that "internally generated brands, mastheads, publishing titles, customer lists and similar items are not recognised as intangible assets", and a professional treatment extends the point to "digital assets including software, databases and domain names" [32]. Converges: nobody knows what a company's record is worth because the balance sheet is not allowed to carry it. Departs: a court-supervised sale manufactures exactly the observation accounting refuses to make, which is why bankruptcy became the accidental price-discovery venue for this asset — and the standard-setter has had the gap under review since 22 July 2026.
Signal scorecard
Six new signals join the ledger with this thesis, in a new cluster, scored against each signal's confirming condition on a −2 to +2 scale. Read the sign carefully on R03.5 and R03.6: a negative score on each supports this reading, and the rows say so. Two existing signals are referenced without re-scoring — the crawl concept keeps its own house, and R02.5's two-sided test has no branch for a lab that buys canonical state rather than building it, the structural gap this thesis exposes. The full ledger is at radar.ersiner.ai/signals.
| Ledger id | Score | Signal | Baseline (value, date) | Source & cadence | Confirms if | Weakens if |
|---|---|---|---|---|---|---|
| R03.1 | −1 | Salvage lot count — US Chapter 11 estates other than Spirit publicly marketing internal corporate data as a distinct asset lot | 0 additional estates, 24 Aug 2026, from a review of 15 Jul to 24 Aug 2026. Recorded beside it but not scored: roughly 100 voluntary transactions at $10k–$100k, and zero state or federal legislative response to the Spirit sale against 42 attorneys general in the 23andMe case | Free claims-agent dockets and restructuring trade press — monthly, and at every Scorecard [2][17] | 3 or more further estates market such a lot by 31 Mar 2027 | 0 by that date |
| R03.2 | +1 | Archive amortization — Reddit's disclosed long-duration remaining performance obligations, total and the portion contracted for calendar 2027 | Total $92.1m, 2027 portion $30.1m at 30 Jun 2026; prior readings $143.7m/$24.8m at 31 Dec 2025 and $120.6m/$29.0m at 31 Mar 2026 | Reddit, Inc. Forms 10-Q and 10-K on SEC EDGAR — quarterly [9] | the 2027 portion stays under $45m and the total keeps falling in the Q3 2026 10-Q | the 2027 portion exceeds $45m, or the total rises quarter over quarter |
| R03.3 | +1 | Realised-payment disclosure gap — how many of six named per-use or revenue-share programmes have published an aggregate payment actually made | 0 of 6, 24 Aug 2026. Cohort fixed at first reading: Microsoft's Publisher Content Marketplace, Cloudflare's pay-per-use partners Ceramic.ai and You.com, ProRata Gist, TollBit, Perplexity Comet Plus. A pool, a split, a rate and a participant count are none of them a payout | The six programmes' own announcements and product pages — quarterly [13] | still 0 of 6 at 30 Jun 2027 | any one publishes a realised aggregate payout |
| R03.4 | +1 | Channel crossing — a human-data manufacturer acquiring, or recorded bidding $5m or more for, a non-labour data asset | One court-recorded instance and two productized intake channels, 24 Aug 2026: Mercor.io Corporation designated alternate at $7.5m with an Exhibit C naming its own upload facility; micro1 publishing $100k/$500k/$1m-plus intake tiers with a $5m referral programme. Cohort fixed: Mercor, Surge, Handshake, micro1, Labelbox, Turing, Invisible | Court dockets, company announcements and primary reporting naming both sides — event-driven [2][18] | a second such acquisition or $5m-plus qualified bid by 30 Jun 2027 | none by that date |
| R03.5 | −1 | Agent-exhaust disclosure — how many of five frontier labs name first-party agent-product usage data as a training-data category in a published card a low count supports this thesis: usage is the specification, not the corpus | 0 of 5, 24 Aug 2026, across published cards for Anthropic, OpenAI, Google DeepMind, Meta and xAI. The Claude Opus 5 card names four categories — public internet, public datasets, private datasets, synthetic data generated by other models — and agent or user data is not among them, although the consumer terms have permitted it since 28 Aug 2025 | Published model and system cards — event-driven [4][19] | 1 or more labs name the category by 30 Jun 2027 | 0 of 5 at that date |
| R03.6 | −1 | Sell-side productization — any US restructuring adviser, investment bank or claims agent offering estate sales of corporate data corpora as a distinct public offering a minus says the asset class does not exist yet, which is the honest reading | 0, 24 Aug 2026. The strongest adjacent evidence is a law firm's own claim of "a genuine buyer market that did not exist five years ago" — demand-side attention, not a sell-side offering | Adviser and bank websites, restructuring-association conference programmes, named press comment — quarterly | 1 or more such public offerings by 30 Jun 2027 | none by that date |
Net across the six is zero, the correct reading of a market this early: the pricing and disclosure signals point one way, the institutionalisation signals the other, and nothing here is yet a trend.
CALL A — Reddit's contracted licensing revenue for calendar 2027 will still be under $45m in the third-quarter 2026 Form 10-Q. Confidence 80%. Resolution date 15 November 2026. Ledger signal R03.2. Grading source: Reddit's Form 10-Q on SEC EDGAR, expected around 30 October 2026 — after the October Scorecard's cutoff, so this call is graded on its own date with a log entry and a ledger reading. Being wrong would mean the archive is being re-signed at durable term.
CALL B — through 30 June 2027, no US restructuring adviser, investment bank or claims agent will publish a distinct sell-side offering for estate sales of corporate data corpora. Confidence 70%. Ledger signal R03.6. Grading: firm websites, restructuring-association conference programmes, named press comment. Being wrong would mean the asset class institutionalised a year earlier than this reading expects, and it is the single fastest way this thesis becomes obsolete.
CALL C — through 30 June 2027, no frontier lab will name first-party agent-product usage data as a training-data category in a published model or system card. Confidence 70%. Ledger signal R03.5. Grading: cards published in the window. This one runs against the popular version of my own argument: if a lab discloses it, "usage is the specification, not the corpus" needs rewriting.
CALL D — the Spirit deidentified-data sale will be approved by 30 November 2026 with at least one employee-data condition attached — a carve-out, a review protocol, or a use restriction — that is absent from the sale agreement filed on 14 August 2026. Confidence 60%. Ledger signal R03.1. Grading: the signed sale order on the case's public docket. This bets on the terms, not the outcome; if the sale is approved as filed, or not approved by that date, it is a miss and will be graded as one.
Regional translation: a corpus bought from an estate has to be summarizable to be usable in Europe
The provenance regime is where geography prices this market differently, and it is fully measurable. The European Union's general-purpose AI obligations, in force since 2 August 2025, bind every provider to a documented copyright policy and a public "sufficiently detailed summary" of training data; enforcement powers began on 2 August 2026, with penalties up to €15m or 3% of global turnover, and the Digital Omnibus in force since 27 July 2026 deferred the high-risk half of the framework while leaving the data half untouched [33]. The consequence is direct: a provenance cost the American rails do not price, falling hardest on exactly the supply this thesis says is scarce, because a bankrupt company cannot answer questions about what was in its own systems.
One pending case will decide whether the text-and-data-mining exception reaches generative training at all. C-250/25 Like Company, referred by the Budapest Környéki Törvényszék, was lodged on 3 April 2025, heard by the Court of Justice on 10 March 2026 and remains pending as of 24 August 2026, with no Advocate General's Opinion on the Court's register [33]. Anyone quoting an opinion date for it is quoting a practitioner alert rather than the Court.
The labour arithmetic is where this layer meets geography, and the constraint is not capital. The wage ladder in section 04 is credential-bound [24]: what limits supply is the rate at which universities and licensing bodies produce credentialed people. And the buy side publishes its own geographic filter: one intake page requires 30 or more employees, documentation in English, United States companies first, then the United Kingdom and Canada [18]. That is a measured exclusion, printed by a buyer — the third supply wave is being sourced in one language first.
Risk: one vendor's published criteria are not the market's, and the wage bands come from journalism with wide dispersion and no published sampling method.
Stakeholder lenses
Investors — mark the gross-to-net gap before anything else.
The headline here is gross flow through a labour marketplace; the business is the take rate underneath. Estimates put one manufacturer at roughly $2.0bn annualised gross against $180–250m net in the first half of 2026 and another at $1.1bn against roughly $450m, an implied take of about 41%, on multiples near 30 times annualised net revenue INFERENCE Sacra estimates; neither company discloses these splits [31] — a software multiple on a staffing profit-and-loss. Consolidation is running at the manufacturer layer, and the reported NVIDIA–Mercor talks are its candidate next instance PRESS-REPORTED preliminary, not agreed [24]. The second question: labour graph, or labelling tool.
Founders — your working record has a price, and a use restriction moves it.
The auction established what no accounting standard will: what a working record fetches depends on delivery format and on who controls the de-identification as much as on what it contains [2]. That price is now quoted outside bankruptcy: published intake tiers at $100k, $500k and $1m-plus COMPANY CLAIM [18], and an upload address written into a court exhibit [2]. Build-versus-sell is therefore a deliberate posture: make the record exportable while you are alive, then decide whether it is inventory or the company itself.
Enterprises — the enterprise carve-out is a purchase condition that is working.
Every agent vendor tracked in this series excludes enterprise tiers from training by default, so the default-on wave never reaches regulated, proprietary, large-scale engineering [19]. That is why the closed-door channel exists, and why an enterprise's own record is a decision rather than an asset at rest: sell it through the new channels [18], or hold it as the coordination capital Thesis 02 described. The terms are negotiable and the Spirit record set the precedent: de-identification "while preserving referential integrity", and the use restrictions argued over it [2]. The acquirer's version is the same question reversed: buy a data-rich target, or bid at an estate. And the schedule is precise: 97,500,000 customer profiles fenced, 7,510,221,520 revenue transactions sold MEASURED [2].
Careers — the opportunity and the exposure sit in the same evidence.
Both sides are here for anyone advancing or changing careers. Opportunities: expert hours at $85 and up as a platform average and $100–200 and up at the frontier PRESS-REPORTED wide dispersion; no vendor publishes a rate card [24]; the labs insourcing the control plane, where durable roles form [16]; data governance and de-identification, priced as a skill by the Spirit clause [25]; and verification work — difficulty a tool can check. Threats, same sources: that hourly headline is gross, with a take rate between it and the worker [31]; insourcing cuts the other way for vendor-side roles: the cost is permanent, the vendor is not [16]; and careers built on public contribution have lost the commons: 1,346 new Stack Overflow questions in July 2026 against 105,892 in October 2022 MEASURED [3].
Method, sources and disclosure
Evidence window: 28 August 2025 to 24 August 2026. The start is the date of the earliest load-bearing evidence here, the consumer-terms update that made assistant sessions a default training input. Every figure was re-checked against the linked primary source on 24 August 2026.
Measurement periods, kept separate from publication dates. Several load-bearing measurements precede the window and reach it through an in-window filing: Reddit's obligations at 31 December 2025 arrive in a 10-K filed 6 February 2026, Prosus reports a year ended 31 March 2026 in an annual report published 27 June 2026, and Getty's and Shutterstock's June quarters were filed in August. Items dated before the window — the Phi-4 report (12 December 2024), Epoch AI's two 2024 reports, the talk adjudicated above (13 December 2024), the Meta–Scale transaction (June 2025) and the 23andMe sale (2025) — are carried as historical anchors, never as evidence about the window.
Incentives, named. Cloudflare sells the toll it measures; Sacra sells research on the companies it estimates; Mercor, micro1, SimpleClosure, Turing and Integral sell into the market they describe, as does the analyst whose environment price ladder is the only one in public; NVIDIA publishes free corpora and sells the hardware their use requires; the estate's investment banker is a paid professional whose declaration supports approval of the sale he ran; practitioner alerts on European case law are law-firm marketing; and every capability statement about a model comes from the company selling it.
Evidence classes, one per stat card, chart and timeline entry: MEASURED an observed count, a filed financial statement or a document on a public record; COMPANY CLAIM a figure a company publishes about itself; PRESS-REPORTED a figure known only from journalism, which never travels without its outlet, its sourcing and the transaction's status; FORECAST a named forecaster's projection; INFERENCE my own arithmetic or judgement, inputs shown.
Vocabulary, attributed. "Referential integrity", "deidentification agent" and "remaining performance obligations" are quoted from filings. "Verified difficulty" and "working record" are mine, as is the five-category sort in section 04; every anchor under it is primary.
What could not be closed. The aggregate value of Reddit's 2024 licensing agreements is not restated in any in-window filing. An opinion date circulating for the pending European case is not supported by the Court's own register. And a widely repeated figure for one lab's annual spending on training environments reaches open sources only at second hand from a paywalled outlet. None of the three is printed above.
How this was made. This thesis began as my own draft position and every claim was treated as something to test. Two lost. The organizing tension I started with — that archives sell once and working records sell every year — could not carry its own evidence, since the measured repeat-sale series is one company deep, and the claim that estate transactions are increasing is not supported inside the window. Both are argued with in the text rather than deleted.
Disclosure. Written in a personal capacity, from public sources only. The author works within the Türkiye venture ecosystem; to avoid conflicts of interest, no Türkiye-based fund or startup is named or evaluated here. The author holds no positions in, and has no client relationship with, any company named.
This is Thesis 03
Third thesis in a recurring series. Theses are numbered across topics and published when ready, never on a calendar; a bi-monthly Scorecard re-scores every live signal no publication has touched and grades the dated calls in public — including the misses, which are the only ones that teach anything. The six signals scored here join the ledger at radar.ersiner.ai/signals, which this thesis also opens the data-deals registry on. First Scorecard: October 2026.
The Agents & Margins briefing is launching — until then, follow on LinkedIn
Sources
- 1.Meta Platforms, Inc. — Form 10-Q for the quarter ended 30 June 2025, source of "In June 2025, we completed an investment in Scale AI by acquiring a non-voting minority of its outstanding equity", of the $13.79 billion allocated to non-marketable equity investments, and of "as we do not have significant influence over Scale AI's operations"; repeated in the following quarter. With Scale AI's own announcement, Scale AI Announces Next Phase of Company's Evolution, 12 June 2025, which discloses a valuation above $29bn and "a minority" of equity and no percentage. The widely circulated "49%" appears in no primary; a press-reported total consideration of $14.3bn circulates alongside the filed figure.
- 2.ECF 1463, Notice of Auction Results for the Deidentified Data, filed 14 August 2026 in In re Spirit Aviation Holdings, Inc., et al., No. 25-11897 (SHL), United States Bankruptcy Court for the Southern District of New York — the results table ($10,000,000 successful, $7,500,000 alternate), the Google Sale Agreement at Exhibit A, the mirror Alternate Sale Agreement at Exhibit B carrying the identical §3(c) "while preserving referential integrity across the data set", the Mercor Exhibit C "Data Extraction" naming data.mercor.com, the delivery-format clauses, and the "SPIRIT DATA CATEGORIES" asset schedule reproduced identically at pages 18 and 33 including the column header "Google's Data Purchase Request". Counts quoted here are as disclosed in that schedule; the two largest are round where their siblings are exact and read as seller estimates. Filings are free at dm.epiq11.com/case/spirit. Sale not approved; hearing 9 September 2026.
- 3.Stack Overflow's production series and product record — monthly question and answer counts queried by the author against the Stack Exchange API 2.3 on 24 August 2026, UTC month boundaries, cross-checked against an independent extract of the same database to within 1.4% for 2022–2025 (2026 monthly values differ by up to 7% across sources because of deletion and snapshot timing). With the company's own posts — A new era of Stack Overflow, 30 December 2025; the pay-per-crawl adoption, 19 February 2026; Announcing Stack Overflow for Agents, 10 June 2026 — and the data-licensing page read 24 August 2026, which names only two customers and still advertises a posting rate roughly ninety times the measured one.
- 4.Anthropic — Claude Opus 5 System Card, 24 July 2026: roughly 1.5 million reinforcement-learning episodes across thousands of training environments in the final phase of training with about 400 full transcripts read, and the four named training-data categories (public internet, public datasets, private datasets, synthetic data generated by other models). Company claim.
- 5.Kylie Robison — OpenAI cofounder Ilya Sutskever says the way AI is built is about to change, The Verge, 14 December 2024, 00:34 UTC, reporting a talk given in Vancouver on 13 December 2024 at the Conference on Neural Information Processing Systems. The three quotations adjudicated above are the reporter's transcription; theverge.com blocks automated fetching, so the text was verified against the Wayback capture of 14 December 2024 and corroborated against a same-day news index. The in-window restatement — "we're moving from the age of scaling to the age of research" — is from the Dwarkesh podcast, 25 November 2025. Two sentences widely attributed to this talk are not verbatim in the report and are not printed here.
- 6.Getty Images Holdings, Inc. — Form 10-Q for the quarter ended 30 June 2026, filed 10 August 2026: "Other" revenue $5.238m against $15.716m (−66.7% year over year), half-year $13.9m against $25.0m, the issuer's own explanation quoted in section 03, and the termination of the Shutterstock merger.
- 7.Shutterstock, Inc. — Form 10-Q for the quarter ended 30 June 2026, filed 7 August 2026: Data, Distribution and Services $56.1m (−16% in the quarter, −28% in the half), the data offering −26% in the half, a $173.738m goodwill impairment and a $155.9m quarterly net loss.
- 8.Informa plc — 2026 Half-Year Results, 30 July 2026: Taylor & Francis revenue −6.1%, growth restated "excluding non-recurring data contracts", and the disclosure that the comparable 2025 half carried "a significant non-recurring data contract".
- 9.Reddit, Inc. — Form 10-Q for the quarter ended 30 June 2026, filed 31 July 2026 (remaining performance obligations $92.1m with $30.1m expected in calendar 2027; "Other" revenue $43.280m; the note that the balance "consists primarily of" long-term content licensing; the change in reporting of logged-in daily uniques from the third quarter of 2026), and the Form 10-K for FY2025, filed 6 February 2026 (balances at 31 December 2025, and the risk language quoted in section 03 and section 04). The reported renegotiation of a roughly $60m-a-year agreement expiring in the first half of 2027, and the five- to eightfold increase discussed by analysts, reach this piece only through business press of 22 July 2026 citing unnamed sources; neither party has published terms.
- 10.The New York Times Company — Form 10-Q for the quarter ended 30 June 2026, filed 5 August 2026: remaining performance obligations on contracts longer than one year, with $40m expected in calendar 2027 against $27m six months earlier, and "Generative AI Litigation Costs" of $4.628m in the quarter and $8.840m in the half. The bucket is described as digital archive and other licensing and certain advertising contracts, so it is not purely AI licensing.
- 11.Epoch AI — Pablo Villalobos and colleagues, Will we run out of data? Limits of LLM scaling based on human-generated data, 6 June 2024 (roughly 300 trillion effective tokens; 80% confidence interval for exhaustion 2026–2032), and Can AI scaling continue through 2030?, 20 August 2024 (roughly 500 trillion indexed plus 400 trillion multimodal tokens supporting 6e28 to 2e32 FLOP). Both outside the evidence window and cited as the origin of the peak-data arithmetic; the observation that the organization's 2026 output has moved to compute and inference is drawn from its own topic listing and is my inference.
- 12.Wiley — fourth-quarter and full-year fiscal 2026 results, 16 June 2026: $49m of AI revenue (+23%), lifetime above $110m, with IQVIA and OpenEvidence named. Company claim; the company does not disaggregate this line in its financial statements.
- 13.The metered and per-use tier, read 24 August 2026 — Cloudflare, Making AI search smarter and the same-day press release, 1 July 2026, source of the pay-per-use direction statement and of the re-fetch claim whose two official variants disagree on the denominator; the pay-per-crawl documentation ("closed beta", last updated 28 July 2026); Microsoft's Publisher Content Marketplace announcement of 3 February 2026; ProRata's published Gist 50/50 split; Perplexity's Comet Plus $42.5m publisher pool at an 80/20 split; TollBit's realised figures, which reach the public only as "hundreds to tens of thousands a month" for roughly 20% of about 7,000 participating sites, press-reported.
- 14.Prosus N.V. — FY2026 Annual Report, financial year ended 31 March 2026, published 27 June 2026: Stack Overflow revenue $129m (+12%), first positive adjusted operating profit, goodwill carrying value $281m → nil with a $280m impairment, $54m of brand and customer intangibles written off in full ("no future cash inflows are associated with them"), the stated cause "the disruption caused by generative AI", and the value-in-use assumptions (average forecast revenue growth 16.7% → 6.6%; pre-tax discount rate 18.2% → 19.1%). Comparatives from the FY2025 report. Neither report mentions the data-licensing product line by name in FY2026.
- 15.Machine-text prevalence, two measurements that are not like for like — Graphite (Jose Luis Paredes, Bevin Benson, Ethan Smith, Gregory Druck), More Articles Are Now Created by AI Than Humans, October 2025: 43,000 Common Crawl URLs, January 2020 to May 2025, AI-classified articles rising from 2.2% to 51.7% and plateauing from May 2024; detector false-positive rate 4.2%. And Lucio La Cava, Luca Maria Aiello and Andrea Tagarelli, Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit, arXiv:2510.07226, 8 October 2025: 51 subreddits, 2022–2024, machine text "marginally present" with peaks "of up to 9%" in some communities in some months, and the finding that such text draws comparable or higher engagement. Different methods, different periods, different units. The unauthorised persona experiment of November 2024 to March 2025 is press-reported, from the Washington Post, 30 April 2025.
- 16.Epoch AI (JS Denain and Chris Barber) — An FAQ on Reinforcement Learning Environments, 12 January 2026: contract sizes of six to seven figures per quarter, $200–$2,000 per task, roughly $20,000 for a website replica and up to about $300,000 for a complex product replica, exclusive deals at four to five times non-exclusive, and the observation of "substantially more in-housing" driven by avoiding vendor margins, maintaining confidentiality and using in-house domain expertise. The publisher sells research into the market it describes; it does not quantify the in-house share.
- 17.SimpleClosure — Asset Hub launch, 16 April 2026, source of the roughly 100 transactions, the $10,000–$100,000 band and the buyer description "AI labs, and reinforcement learning researchers", with Anna Tong's same-day report in Forbes naming the transcription company and the "hundreds of thousands of dollars" price; and the H1 2026 shutdown report, 20 August 2026, source of the 27.3% software share, the $11,900 median cash at closure, the highest-volume claim, and the falling AI share of closures (17.7% in 2024, 15.9% in 2025, 14.4% in H1 2026). The August release carries no Asset Hub transaction figures; those remain sourced to April. Company claims throughout.
- 18.The buy-side intake pages, read 22–24 August 2026 — micro1's data-partnerships page (tiers at $100k, $500k and $1m-plus; a $5m referral programme paying up to $25,000 per company introduced; eligibility criteria of 30 or more employees, documentation in English, United States first, then the United Kingdom and Canada) and Mercor's data page (referral payments up to $100,000). Company claims; both companies buy the supply they describe.
- 19.The contract door, three primaries — Anthropic, About the Development Partner Program, 22 May 2026 ("Only your Claude Code input and output tokens from the first-party Claude API are provided to us", two-year retention, organization-admin opt-in, no compensation stated); Anthropic, updates to consumer terms, 28 August 2025, with the accompanying documentation covering assistant sessions "including when you use Claude Code from these accounts"; and GitHub, updates to the Copilot interaction-data usage policy, announced 25 March 2026 and effective 24 April 2026. All company claims.
- 20.The rest of the Spirit record — ECF 1508, Limited Objection of Springshot, Inc., filed 21 August 2026, which quotes the assets definition verbatim, raises the §1(c) resale right, and puts the buyer's own words on the record at page 7 note 4 ("which can be helpful in improving our products and AI models"); ECF 1470, the investment banker's declaration, recording that a first bid sought the customer list and that the most competitive first-round bids moved to a schema excluding personal information; and ECF 1446, the consumer privacy ombudsman's report of 12 August 2026, a scanned filing read visually. The separate marketing of the customer list to "third-parties in hospitality or travel industries" is press-reported, 19 August 2026; no docket entry for that lot existed as of 24 August 2026.
- 21.Anthropic — Detecting and preventing distillation attacks, 23 February 2026: roughly 16.55 million exchanges across about 24,000 fraudulent accounts attributed to three named actors, with the extracted capabilities named as "agentic reasoning, tool use, and coding". Company claim, and the company sells access to the models it is protecting.
- 22.Synthetic generation at frontier scale — Microsoft, Phi-4 Technical Report, arXiv:2412.08905, 12 December 2024 (40% synthetic plus 15% web rewrites of about 10 trillion tokens; outside the window, carried as an anchor); NVIDIA's published post-training corpus, whose responses are 100% synthetically generated; and Kang and colleagues (Meta FAIR and Virginia Tech), Demystifying Synthetic Data in LLM Pre-training, arXiv:2510.01631 — more than 1,000 models and over 100,000 GPU hours, finding roughly 30% rephrased synthetic optimal with a five- to tenfold speed-up. The collapse canon it answers: Shumailov and colleagues in Nature (2024, author correction 2025), with Gerstgrasser and colleagues (arXiv:2404.01413) on accumulation and Schaeffer and colleagues (arXiv:2503.03150) counting eight conflicting definitions of collapse.
- 23.The synthetic-data vendor category — MOSTLY AI's closure after almost nine years, reported 4 March 2026 by an Austrian technology outlet with the founder on record attributing it to weakening privacy regulation, liquidation dated 14 March 2026 in the commercial register; Syntho's announcement of 9 June 2026 that it had acquired the brand; NVIDIA's 2025 acquisition of Gretel, whose price was never confirmed and whose widely repeated "$320m" is a prior valuation rather than a price, with the brand, pricing page and open-source kit gone by 2026; and Prime Intellect's Series A, 8 July 2026, the largest verified revenue in the reinforcement-learning environment business, selling to enterprises training their own agents rather than to frontier labs.
- 24.The human-data manufacturing layer, company announcements and reporting — Handshake's acquisition of Cleanlab, 28 January 2026, terms not disclosed, with the cofounder statement that competitors "frequently use Handshake's platform to source human experts" (single-sourced); Labelbox's acquisition of Upcraft, 10 February 2026, terms not disclosed; Surge's reported revenue above $1bn on roughly 110 employees with no outside capital through 2024; micro1's reported $500m gross run rate (TechCrunch, 20 August 2026); the reported NVIDIA–Mercor talks at a roughly $20bn valuation, 19–20 August 2026, preliminary and not agreed; and the wage bands for annotation and expert work, which appear across trade coverage with wide dispersion and no published sampling method. Press-reported unless the company published it.
- 25.Transfer-governance and masking vendors, read 24 August 2026 — Integral's "entity-preserving" de-identification (announced 4 May 2026); Perforce Delphix, claiming deterministic masking such that "the same input always produces the same masked output across every system, environment, and data pipeline" and that the product "ensures referential integrity across enterprise data estates"; and Tonic.ai (Tonic Structural and Tonic Textual), claiming products "built to de-identify sensitive data and generate realistic synthetic data while preserving the relationships between tables". All company claims; the latter two are positioned mainly as test-data-management vendors.
- 26.The legal lane — Bartz v. Anthropic PBC, No. 4:24-cv-05417 (N.D. Cal.), docket entry 680, order granting final approval and entering judgment on the $1.5bn settlement, 20 July 2026 (order PDF mirror); the docket count showing no class member or objector appeal within the window that closed 19 August 2026, the two notices concerning only the $101,561,111 fee pool, and the absence of any announced effective date or payment as of 23 August 2026. The bifurcation it rests on is a district-court holding; the interlocutory appeal that would have tested it was withdrawn with the settlement. Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc. was argued in the United States Court of Appeals for the Third Circuit on 11 June 2026 and remains undecided.
- 27.The open-data axis — NVIDIA's Nemotron-CC-v2 (about 6.6 trillion tokens, 18 August 2025) and Nemotron-CC-v2.1 (2.5 trillion tokens, 15 December 2025), both under NVIDIA's own data agreement for model training; Ai2's Dolma 3 (about 9.3 trillion tokens, ODC-BY) published with OLMo 3; Hugging Face's FineWeb family (15 trillion tokens); Kari Briski, vice president of generative AI software at NVIDIA, interviewed by The Deep View, 6 April 2026, source of "The model is the byproduct. It is not core to our business, which allows us to just open up the data, open up the recipes, open up everything"; the Foundation Model Transparency Index (December 2025), whose industry average fell from 58 to 40 out of 100 with training data the most opaque dimension; and the count of 94 tracked public licensing agreements of which about 40% include training rights. The complementary-goods reading of NVIDIA's behaviour is mine and is not the company's stated rationale.
- 28.The distillation axis — Alexander Panfilov and colleagues, Stealing Reasoning Traces from Proprietary LLM APIs, arXiv:2608.09867, 10 August 2026: all three major United States providers return reasoning as encrypted blocks held by the client, the blocks are cross-compatible within a provider's ecosystem, and decoding 10,000 traces at 12,000-token windows costs roughly $720. With the sworn testimony of 30 April 2026 in which Elon Musk conceded xAI had "partly" used OpenAI technology and offered the defence "Generally A.I. companies distill other A.I. companies"; Google's Gemini API terms effective 23 March 2026 forbidding attempts to "extract or replicate any component of the Services"; and the White House memorandum of 23 April 2026 on adversarial distillation, which is an image-only scan and is therefore press-reported here rather than quoted.
- 29.Cloudflare — the acquisition of Human Native, announced 15 January 2026, price not disclosed. Company claim; the acquirer already owned the request path the marketplace depended on.
- 30.Other flow sellers — People Inc's second-quarter 2026 results, "Licensing and other" revenue $47.0m (+23.3%), a bucket that also holds brand licensing and a news-aggregation partnership, so the AI share is a subset and is unverified; and the reported News Corp arrangement with Meta at up to $50m a year for at least three years, first reported by a title the seller owns, with no published terms from either party.
- 31.Sacra — research on the human-data layer, read 24 August 2026: gross-versus-net estimates for the largest manufacturers (roughly $2.0bn annualised gross against $180–250m net in the first half of 2026 at one; $1.1bn gross against roughly $450m net at another, an implied take of about 41%), gross margins of 30–40% and multiples of roughly 30 times annualised net revenue. These are the estimator's own figures, not company disclosures, and the estimator sells research on the companies it estimates.
- 32.The intangible-asset canon — IFRS Foundation, IAS 38 Intangible Assets ("Internally generated brands, mastheads, publishing titles, customer lists and similar items are not recognised as intangible assets", with the stated reason that the cost of generating one internally "is often difficult to distinguish from the cost of maintaining or enhancing the entity's operations or goodwill"); Gary Anders, Accounting for intangibles: IAS 38 review explained, INTHEBLACK (CPA Australia), 1 November 2025, extending the point to "digital assets including software, databases and domain names" and quoting Michael Masterson of Intrinsika; and the IASB's own Intangible Assets work-plan project, at the "Decide Project Direction" stage after the meeting of 22 July 2026.
- 33.The European provenance regime — the general-purpose AI obligations in force since 2 August 2025, the copyright chapter's requirement of a documented policy and a public "sufficiently detailed summary" of training data, enforcement powers from 2 August 2026 with penalties up to €15m or 3% of global turnover, the 2 August 2027 date for models placed before the regime began, and the Digital Omnibus in force from 27 July 2026 deferring the high-risk provisions to 2 December 2027 while leaving the data provisions untouched. With the Court of Justice's own register for Case C-250/25 Like Company, read 24 August 2026: referred by the Budapest Környéki Törvényszék, lodged 3 April 2025, heard 10 March 2026, pending, with no Advocate General's Opinion listed and none scheduled in the Court's five-week judicial calendar.