The Private Data Gap: What Labs Cannot Get from the Open Web

The open web has been read. What frontier models need next was never published, and it will not be scraped. It will be licensed.

September 2026

For a decade the working assumption of the AI industry was that the web would keep providing. More pages, more posts, more repositories: the corpus was treated as a renewable resource. That assumption has quietly expired. The models grew to the size of their food supply, the high-quality portion of the open web has largely been read, and the frontier's appetite now exceeds what publication produces. What follows is not a slowdown. It is a change in where the next data comes from.

The web was finite after all

The exhaustion is structural rather than numerical. Repetition returns diminishing gains: training again on what a model has already absorbed sharpens very little. Synthetic data helps at the edges and inbreeds at the center. And the open web's growth is increasingly the output of models themselves, which makes fresh scraping partly an exercise in models reading their own exhaust. The practical consequence inside the labs is that gathering has given way to sourcing. The question is no longer what is out there. It is who holds what the model needs, and on what terms it can be licensed.

The symptoms are visible from outside the labs. Crawlers meet more refusals every quarter as publishers and platforms fence off what remains. Litigation over scraped corpora has turned unlicensed material into a balance-sheet risk rather than a bargain. And the labs' own sourcing teams have gone from an afterthought to a discipline, with budgets that did not exist a few years ago. None of that happens in a world where the open web still feeds the frontier.

What the open web never contained

The deeper problem is not that the web ran out. It is what the web never had. The web is the published record: conclusions, announcements, final drafts, accepted answers. Work itself happens elsewhere, in the private record. The issue threads behind the code. The calls behind the contracts. The drafts behind the reports. The tickets behind the resolved complaint. The operational tables behind the quarterly summary. Models trained only on the published record learn what finished work looks like without learning how work gets finished.

The capability the frontier wants next, systems that can carry multi-step work to completion, is taught by exactly the material that was never published. That material exists at extraordinary scale. It sits inside operating companies that never thought of it as an asset, in formats no scraper will ever reach, behind rights no scraper could ever clear.

The licensing turn

The buy side's response is already visible in the deals that reached the public record. Major publisher archive deals put recurring terms on decades of edited text. Forum-corpus licensing by large platforms priced threaded human judgment on an annual basis. Per-title book licensing by a big-five publisher put a unit price on the professional long form. Even a bankrupt airline's operational archive found a buyer at auction, which is the purest demonstration on record that ordinary operational material clears at a real price. Around the modalities, standing markets have formed: licensed video trading per minute, speech datasets sold in tiered bundles by the hour.

The pattern across all of it fits in one sentence: labs are paying for what scraping cannot legally or physically reach.

Discretion is the unlock

If the demand is real and the supply is vast, why has the market not flooded? Because the holders are operating companies, and operating companies have excellent reasons not to license in public. A public listing tells competitors what you hold. It tells clients that their vendor licenses data. It invites every question a general counsel dreads, before any money is on the table. The private data gap persists not because holders lack data but because licensing has looked like exposure.

This is why the mechanics matter as much as the market. A workable private-data deal is discreet end to end: the asset is described rather than displayed, matched buyers review under NDA, the seller is named only with consent and usually only at contract time, and what changes hands is a license rather than the asset itself, so the business keeps what it built. Discretion is not a nicety of the process. It is the technology that converts holders into sellers.

Discretion also returns control to the seller in the places that matter commercially. You decide the scope of the corpus, and exclusions are normal. You decide between exclusive terms at a higher price and non-exclusive terms that can be licensed again. You decide when, and whether, your name enters the room. A market that asks holders to give up that control stays empty. One that preserves it fills.

The gap is the market

Between what the frontier needs and what the open web can still give sits the private archive of essentially every operating company: the code, the documents, the images, the recordings, the footage, the workflows, the tables. That gap is not going to close on its own, because publication habits will not change and the appetite will not shrink. It closes one private transaction at a time, on the seller's terms or not at all.

The companies that treat their archive as an asset, and license it the way assets are licensed, with rights in order and discretion intact, are the supply side of the next phase of AI. The only open question for any particular company is whether the archive keeps sitting in storage, or starts earning.

You hold what the web does not.

Free instant valuation. Private licensing. Paid when it closes.

Get your free valuation

All insights