When Runway released Gen-4.5 this past January, the company made the expected promises: longer clips, crisper resolution, native audio. Google's Veo kept shipping updates through last year. Luma unveiled Ray3 in December, integrating real camera hardware. Each launch delivered better outputs. Each also exposed a more fundamental constraint.
The question everyone building generative video started asking wasn't about compute anymore—though compute still matters, of course. It was simpler to state, harder to solve: Where do you actually get enough video?
Text models had Common Crawl, that vast repository of scraped web pages. Image models hoovered up the open internet until the lawsuits started arriving, then scrambled for licensing deals with Shutterstock and Getty. Video, though, presents a different animal entirely. It's heavier by orders of magnitude, harder to label meaningfully, locked behind platform terms of service that explicitly prohibit bulk scraping, and increasingly tangled in regulatory tripwires. As the race to ship video foundation models intensifies, a parallel scramble has begun—quieter, less celebrated, but possibly just as consequential—to build the data infrastructure that actually feeds them.
The Problem With Pixels at Scale
The numbers sketch the outlines. WebVid-10M, released back in 2021 and widely used in early video model research, offered roughly 52,000 hours of captioned clips scraped from across the web. By last year, academic papers were openly criticizing its noise and quality problems. Panda-70M arrived with 70 million clips curated from HD-VILA-100M, a notable step up. OpenVid-1M appeared aiming to improve on WebVid at research scale. ViMix-14M appeared late last year, promising "crawl-free access" and longer-form captions.
Valuable as these datasets are for academic work, they remain snapshots. They age poorly. They lack the granular labels that model builders increasingly demand: hand-object interactions, scene continuity, engagement patterns, creator intent. And perhaps more critically, they don't solve the licensing problem.
YouTube's CEO warned in April 2024 that training on platform videos without permission violates terms of service. (Google itself acknowledged in June last year that it uses YouTube content to train Gemini and Veo—a privilege available to the platform owner but not to outside labs.) When Cloudflare began blocking AI crawlers by default last July, the message became unmistakable. The era of freely scraping the web for training data, already fraying, is closing.
A Category That Didn't Exist Two Years Ago
Enter what some founders are calling "video data infrastructure"—a category that barely existed 18 months back. Shofo, a startup in Y Combinator's Winter 2026 batch, bills itself as "Common Crawl for Video." The company maintains what it describes as "the world's largest index of short-form videos," running human-in-the-loop pipelines to clean, segment, and label content from social platforms. Their public dataset on Hugging Face, updated late last month, reveals the schema: transcripts, hashtags, engagement metrics, fps, resolution, duration, flags for AI-generated content and ads, language tags. The fields tell you what video model builders need beyond raw pixels.
Shofo isn't working in isolation. VIDEO for AI, established last year, operates as a marketplace facilitating licensing deals between catalog owners and AI companies. Luel, also from YC's winter batch, takes a contributor model—providing rights-cleared video and audio datasets, with particular focus on robotics applications. LXT offers video data collection across more than 150 countries for enterprise compliance use cases. Acuity AI positions itself around real-world robotics video.
The approaches vary—indexing and curation, licensing marketplaces, contributor networks, compliant collection—but they share a recognition. Assembling large, labeled, legally defensible video corpora requires actual infrastructure, not just clever scraping scripts and storage buckets.
The Legal Thicket Keeps Growing

The backdrop here is a compliance environment that keeps thickening. The EU AI Act's general-purpose AI obligations took effect last August for new models. Providers must now submit training-data summaries and demonstrate copyright compliance. Models released before that date have until August 2027 to comply, but the trajectory is set. If you're training a foundation model for distribution in Europe, provenance isn't optional.
In the US, the legal picture remains messier but no less fraught. Getty Images largely lost its UK copyright claims against Stability AI last November—a ruling that narrowed liability theories for training on images—but refiled in California back in September. The litigation grinds forward. Federal legislation like the TAKE IT DOWN Act, signed last May, criminalizes non-consensual deepfake intimate imagery and mandates 48-hour takedowns. Multiple states have enacted or expanded deepfake laws in the past year. The regulatory pressure isn't fading.
Meanwhile, reported incidents keep climbing. Gartner research cited last September found that 62 percent of organizations had experienced deepfake incidents. A study released by Runway in January found that humans correctly classify real versus AI-generated videos only 57.1 percent of the time—barely better than coin flips. The trust and safety implications ripple backward to training data. If you can't verify the provenance of your corpus, how do you prevent poisoned inputs or CSAM contamination? How do you demonstrate compliance to enterprise customers, or to regulators who are increasingly paying attention?
These aren't hypothetical concerns for AI labs. They're practical blockers. Shutterstock and OpenAI signed a six-year licensing deal back in July 2023 covering images, video, and music libraries—a recognition that compliant access beats legal exposure. Runway struck a deal with Lionsgate in September 2024 to train custom models on the studio's catalog. Adobe integrated Runway's Gen-4.5 into Firefly this past December, showcasing tight model-toolchain partnerships that sidestep messy licensing questions.
What Winning Looks Like
So what does it actually take to win in this emerging category? The technical requirements alone are steep. You need pipelines that can ingest, deduplicate, and segment video at scale. You need embeddings and semantic search so customers can query for "50,000 cooking videos with hand-object interactions and no scene changes," as Shofo's marketing suggests. You need continuous updates—models trained on 2021 internet video struggle with 2026 visual culture and aesthetic norms. You need labeling that goes beyond captions: activities, objects, interactions, scene coherence, audio-visual sync.
But the harder moat, perhaps, is legal and operational. Platform terms of service prohibit bulk scraping of YouTube, TikTok, Instagram. Compliant collection requires either negotiated API access, creator opt-ins, or licensing deals with rights holders. That's relationship-heavy work, not just engineering.
Twelve Labs, which brought its video-understanding foundation models to AWS Bedrock last April after signing a three-year strategic collaboration agreement, illustrates the enterprise pathway: deep integration with cloud platforms, focus on searchability and retrieval-augmented generation, emphasis on provenance tracking. It's a playbook that prioritizes trust over growth hacking.
For robotics and embodied AI applications, the data needs diverge further. Ego4D, a large egocentric video benchmark, reached roughly 3,670 hours in its second version, with challenges running into last year. Follow-on datasets like EgoExo4D and Apple's EgoDex—829 hours captured on Vision Pro, released last April—expand high-fidelity interaction data. Research groups are developing methods to convert human-hand videos into robot-centric views. Tesla's petabyte-scale video corpora for autonomous driving, processed on custom Dojo hardware, hints at the compute and data scale required for real-world perception tasks.
The Market Signals Are Accumulating

Menlo Ventures estimated $18 billion in generative AI infrastructure spend for 2025, double the 2024 figure. Vendor estimates for the broader video processing platform market range from something like $8.76 billion this year to a projected $22.6 billion by 2034, though methodologies vary widely and should be taken with appropriate skepticism. The training data platform market itself is pegged somewhere around $2.35 billion in 2025, according to one market research estimate.
Scale AI, the data labeling incumbent, saw reported valuations swing wildly across 2024 and into last year—some outlets cited figures from $14 billion to $29 billion depending on the round and reporting date. The company secured a strategic investment from Meta last June, underscoring the value large labs place on data infrastructure. Surge AI serves major frontier customers including OpenAI, Google, Microsoft, Meta, and Anthropic. Appen, the legacy player, posted revenue declines through fiscal 2025—$233.4 million in revenue with a 9 percent higher net loss year-over-year—suggesting the market is shifting toward newer entrants and more specialized providers.
What Comes Next
What happens next feels predictable in outline, if not in specifics. Training-data procurement will keep shifting from opportunistic crawling toward a mixed model: licensed studio and catalog deals, creator opt-ins and marketplace platforms, targeted contributor collection for specialized domains like robotics or industrial applications. Platforms with proprietary video at scale—YouTube, TikTok, Instagram—become gatekeepers rather than open resources. Closed APIs and enforced terms of service mean negotiated access, which elevates vendors who can deliver rights-cleared, continuously updated video corpora with robust provenance and labeling.
The safety and provenance stack becomes table stakes. Watermarking, C2PA-style provenance logging, consent and indemnity tracking, automated filtering for disallowed categories—all of this infrastructure matters under EU obligations and enterprise security policies. Model builders will pay for it because the alternative is regulatory exposure or reputational disaster. Or both.
Shofo, for its part, remains early. The company was founded last year, went through Y Combinator this past winter, and has not disclosed customers, revenue, or funding beyond YC backing. Its four founders—Bryan Hong, Andre Braga, Braiden Dishman, and Alexzendor Misra—previously built Correkt, a multimodal search tool that reportedly had 43,000 users before they pivoted. Their public dataset on Hugging Face shows roughly 58,000 rows of TikTok metadata spanning videos from March 2017 through January, with fields for transcripts, engagement, AI-generated flags, and licensing marked as "other."
Whether Shofo specifically becomes the Common Crawl of video or gets outcompeted by better-funded players remains an open question. What feels less debatable is that video foundation models need better data infrastructure than what exists today—fresher, cleaner, more compliant, more richly labeled. The companies that solve that problem won't just support the model builders. They'll shape what kinds of models get built, and by whom.
That's the less visible race happening right now, while everyone watches the model releases.
