The internet that built GPT-4 is closing shop.
Between mid-2025 and early 2026, Cloudflare reported blocking approximately 416 billion AI bot requests—a figure that, if accurate, suggests the era of freely scraping the web is effectively over. At the same time, projections from research outfit Epoch AI indicate that the world's supply of high-quality, non-repetitive human text might run dry somewhere between 2026 and 2032, depending on how fast the labs keep scaling their models. The math is unforgiving: AI systems are growing more voracious just as the buffet is being cleared away.
What's emerging to fill the gap isn't simply a legal maneuver. It's an entirely new infrastructure for training data—one built not on algorithmic scraping, but on consent, provenance, and the deliberate capture of real-world human activity. If the last generation of models was trained by crawling Wikipedia and hoovering up Reddit threads, this one is being shaped by distributed contributor networks, licensing agreements worth hundreds of millions of dollars, and startups positioning themselves as something like "the API for real-world data." The result is a market that mixes genuine technical innovation with a heavy dose of compliance theater and old-fashioned data ops. It's messy. It's expensive. And it's moving faster than most people outside the labs realize.
A Market in Three Tiers
The training data economy has split into distinct layers, each with its own economics and players.
At the top sit the licensing megadeals—agreements that amount to Silicon Valley's version of paying rent after years of squatting. Reddit, in its February 2024 S-1 filing, disclosed aggregate contract value of $203 million across AI partnerships. News Corp reportedly inked a five-year deal with OpenAI somewhere in the neighborhood of $250 million. Stack Overflow and Axel Springer have signed similar pacts. These are the marquee names, the structured, high-value corpora that labs are now willing to pay eight or nine figures to access without legal exposure.
Below that, the annotation and labeling layer hums along with less fanfare but no less importance. Scale AI—perhaps the most visible player here—raised a $1 billion Series F in May 2024 at a $13.8 billion valuation, with backing from Amazon and Meta. That round underscored continued appetite for managed data operations: the RLHF feedback loops, evaluation datasets, and quality assurance pipelines that turn raw data into something usable for model training. Competitors like Labelbox, LXT (which folded Clickworker into its operations in 2025), and Defined.ai (which has reported strong revenue growth) are all vying for pieces of the same pie. According to Grand View Research, the data annotation tools market stood at around $1.03 billion in 2023 and is projected to hit $5.33 billion by 2030—a compound annual growth rate in the mid-twenties, if projections hold.
The third tier is newer, less consolidated, and considerably more experimental: commissioned multimodal capture. Academic projects like Ego4D (more than 3,600 hours of egocentric video) and Ego-Exo4D V2 (1,286 hours with multimodal sensor fusion) offer a glimpse of the modalities frontier labs are chasing—first-person video, inertial measurement unit data, depth sensing, gaze tracking. But these public datasets are orders of magnitude smaller than what's needed for a frontier training run. That gap has created an opening for startups that promise distributed, consented collection at commercial scale. Whether they can deliver is another question.
What's Driving This
The technical pressure is straightforward, if not exactly simple. Multiple studies have documented the risks of "model collapse"—the degradation that occurs when models are trained on synthetic or recursively generated data. A 2024 explainer in Nature and several peer-reviewed papers confirm that while carefully curated mixtures of real and synthetic data can mitigate some of these risks, there's still no substitute for a foundation of authentic, human-generated material. Synthetic data has a ceiling. It can supplement, but it can't fully replace the messy, diverse inputs that come from real human activity.
Meanwhile, the legal and regulatory ground is shifting beneath the entire enterprise. The EU AI Act's provisions for general-purpose AI took effect in August 2025, requiring transparency reports, copyright policy disclosures, and training-data summaries in a format the European Commission approves. Pre-existing models have until August 2027 to comply, but enforcement powers reportedly went live in mid-2026. ISO/IEC 42001, an AI management systems standard published in 2023, is gradually becoming a checkbox item in enterprise procurement. The NIST Generative AI Profile, released in July 2024, emphasizes data provenance and documentation—the kind of paperwork that many early AI companies never bothered with. Buyers are now asking for per-row audit trails and evidence of contributor consent, questions that would have seemed absurd just a few years ago.
In the United States, the regulatory picture is messier. Executive Order 14110 was revoked in January 2025, but federal agencies still lean on OMB guidance and the NIST AI Risk Management Framework in their compliance planning. California's CPRA enforcement and the Data Broker Delete Act are building out registry and audit requirements that extend through 2028. GDPR's Article 9 restrictions on biometric data processing continue to complicate operations in Europe. And litigation—most prominently, The New York Times v. OpenAI and Microsoft, which survived a motion to dismiss in April 2025—has made it clear that the copyright risks of web scraping are both real and potentially ruinous.
Then there's the infrastructure crunch. Memory shortages, datacenter power constraints, and bottlenecks in HBM and DRAM supply chains are pushing labs toward data efficiency over brute-force scaling. When compute becomes the limiting factor, data quality matters more than it used to. And when the open web is increasingly locked behind Cloudflare's default-block settings and robots.txt policies enforced at CDN scale, scraping becomes cost-prohibitive even when it's technically legal. The old model—crawl everything, sort it out later—just doesn't work anymore.
Case Studies from the Field

Niantic Spatial offers a useful precedent, though not one most people think of as a training data company. The firm reportedly uses over 30 billion geolocated images—captured via Pokémon GO and similar apps—to power visual positioning systems and robotics localization. That scale wasn't achieved through scraping. It came from incentive-driven crowdsourcing: users contributed data in exchange for gameplay, and Niantic aggregated it into a commercial product. It's a model that trades incremental user value for massive, diverse, real-world datasets. Whether it can work outside gaming is an open question.
Hub.xyz, a Palo Alto-based startup from the Spring 2026 Y Combinator batch, is attempting something similar for training data more broadly. The company bills itself as "the API for real-world training data," operating what it describes as a distributed contributor network across 150 countries. Its pitch centers on provenance: every dataset supposedly comes with per-contributor audit trails, consent documentation, and GDPR and CCPA compliance built in from the start. The platform offers off-the-shelf multimodal datasets—figures cited include 54,000 hours of egocentric video, 2.4 million image frames, 70,000 video clips, 80,000 hours of audio across 47 languages—as well as bespoke capture through what the company calls contributor and SMB networks.
Hub.xyz raised $1.7 million in a pre-seed round led by SwissBorg in September 2025, opened a U.S. headquarters in Palo Alto the following month, and hired a head of sales from stints at Google and Bright Data. The company's LinkedIn presence through early 2026 has repeatedly articulated the thesis: the open web is "dying" for AI ingestion, synthetic data has inherent limits, and enterprise buyers now demand provenance and consent as table stakes. The team size is listed as 10 people on the YC directory. Third-party API gateway provider Lava has integrated a Hub.xyz dataset endpoint, which suggests at least some level of technical validation. But the company has not disclosed named customers, independent audits of dataset quality, or updated funding figures beyond that initial announcement. In other words, it's still early—very early.
The competitive landscape is crowded but segmented in interesting ways. Scale AI, Labelbox, LXT, and Surge AI dominate the managed labeling and RLHF operations side. Defined.ai emphasizes custom speech and NLP collection. Appen, a legacy vendor that once dominated this space, has faced revenue declines and client concentration issues over the past year or so. Mozilla's Common Voice project continues expanding long-tail language coverage, tracking over 250 languages as of late 2025. What sets newer entrants like Hub.xyz apart isn't necessarily superior annotation tooling—competitors have that well in hand—but rather the focus on fresh capture, contributor networks, and compliance-first provenance rather than post-hoc labeling of existing web data. Whether that's a durable advantage or just good positioning remains to be seen.
What Comes Next
The convergence of data scarcity, regulatory pressure, and infrastructure constraints suggests this market isn't going away anytime soon. Epoch AI's projections remain the most frequently cited quantitative benchmark: depending on quality filters and scaling assumptions, labs could exhaust the easily accessible stock of human-generated public text somewhere between 2026 and 2032. (Those are wide error bars, which tells you something about the uncertainty involved.) Multimodal training—especially for embodied AI and robotics—will require even more specialized capture: first-person video, sensor fusion, environment scans. The academic datasets that exist today are proof-of-concept, not commercial scale.
The licensing economy for existing high-value corpora—news archives, forum discussions, code repositories—is now established. The open question is whether distributed, consented capture can scale to meet frontier training demands at a price labs and enterprises are willing to pay. If contributor networks can deliver quality, diversity, and audit trails at something resembling API-level convenience, they'll capture real value. If the overhead of provenance and compliance makes costs prohibitive, labs will lean harder on synthetic augmentation and smaller, more curated datasets. That's the bet everyone's making, one way or another.
For founders building AI products, the implications are practical and immediate. Procurement teams are asking new questions: Where did this data come from? Who consented to its use? Can you provide an audit trail? The EU's transparency requirements and ISO 42001 adoption are turning these questions from nice-to-haves into hard requirements, particularly for enterprise sales. Startups that can answer them convincingly—whether through licensing deals, commissioned capture, or contributor networks—will have an edge in sales cycles that are already long and complex.
For investors, the data infrastructure layer is splitting in two. Scale AI's billion-dollar round signals continued demand for managed operations at the high end, where margins and integration complexity justify premium pricing. But there's room for differentiation below that: companies that can unlock hard-to-access modalities (egocentric video, long-tail languages, sensor-rich environments) or solve compliance and provenance at the dataset level rather than the annotation level. The market is early enough that positioning matters as much as execution—maybe more.
The Reckoning

The web that trained the last generation of models isn't coming back. Cloudflare's default-block policies, publisher paywalls, and the looming exhaustion of public text have created a new kind of scarcity—one that's both technical and legal. What replaces it will be more expensive, more regulated, and more deliberately structured. Whether that's an improvement or just a more bureaucratic version of the same scaling logic depends largely on who's building it and what incentives they're responding to.
But one thing seems certain: the rush is on. And this time, the data comes with receipts.
