The demos look effortless. Type a few words, wait thirty seconds, and watch as an AI conjures a photorealistic video from nothing—a sunset over Tokyo, a cat riding a skateboard, anything you can imagine. But behind the seamless interfaces of Google's Veo and Runway's latest models lies a problem the industry doesn't advertise: they're running out of video to feed the machines.
It's a strange kind of scarcity. The internet is drowning in video content—hours upon hours uploaded every minute. Yet the specific, high-quality, legally defensible footage needed to train state-of-the-art AI models? That's become harder to find than venture capital at a down-market pitch session.
Enter Shofo, a four-person startup from Y Combinator's Winter 2026 batch that wants to become what founder Bryan Hong calls "Common Crawl for Videos." Launched publicly in late January, the San Francisco company promises something AI labs are desperate for: custom video datasets, millions of hours deep, cleaned and tagged and delivered in days instead of months. It's an audacious bet on infrastructure at a moment when everyone else is chasing the next flashy model.
Whether Shofo can pull it off is an open question. But the fact that they're trying—and that they have plenty of company—says something about where the AI industry finds itself in early 2026.
The Acceleration Nobody Expected
The pace has been dizzying, even by Silicon Valley standards. Google unveiled Veo 3 in May 2025 with native audio generation, then pushed out Veo 3.1 this January with vertical video support and transparency tooling. Runway iterated from Gen-4 in March 2025 to Gen-4.5 by December, alongside what it calls a "world model" for generating longer, more coherent sequences. OpenAI struck a billion-dollar deal with Disney in December, licensing more than 200 characters from Marvel, Pixar, and Star Wars—the first major IP partnership for a frontier video model.
Perhaps more telling was YouTube's move last September to integrate Veo into Shorts, allowing ordinary users to generate eight-second clips with text prompts. That kind of platformization means distribution at scale, but it also means data requirements at scale. And moderation headaches, too, though YouTube hasn't said much about that part.
The technical benchmarks keep rising. Academic research like Helios 14B, published in early March, pushes toward real-time, long-form generation with temporal coherence—a fancy way of saying the model doesn't forget what happened three seconds ago. New evaluation frameworks like LongVideoBench stress the need to process more frames, more tokens, longer contexts. Better models demand better data, which enables better models, which demand even richer data. It's a loop, and it's tightening.
Why Video Data Is Different
Text was easy. Scrape the web, filter out the junk, tokenize billions of words, and you've got training data. Images? A bit harder, but manageable. Video is another beast entirely.
State-of-the-art text-to-video models need millions of clips spanning an absurdly long tail of scenarios. Open research datasets exist—WebVid-10M, InternVid's 760,000 hours released at ICLR 2024—but they come with obvious gaps. Long-form supervision. Safety edge cases. Industrial footage. Dense annotations. And critically, licensing certainty.
Papers published throughout 2024 and 2025 consistently flag collection friction. LongVideoBench, released in July 2024, highlighted the scarcity of hours-scale, long-context datasets. SurveillanceVQA-589K and iSafetyBench, both published in 2025, focused on anomaly detection and safety scenarios where generic corpora fall short. Even VideoUFO, a million-clip dataset from March 2025 emphasizing Creative Commons YouTube content, doesn't solve what one AI researcher called "the bespoke data problem."
Here's what that means in practice: an AI lab working on cooking videos might need 100,000 hours showing hand-object interactions, no scene changes, good lighting, multiple camera angles. That level of specificity doesn't exist in off-the-shelf datasets. You have to build it, which takes time and money—lots of both.
Epoch AI's analyses, updated through 2025, project exhaustion of high-quality human text data "before around 2026" in pessimistic scenarios. Vision data has longer runways—2030 to 2060 depending on modality and whether you're willing to reuse synthetic content. But video data, being richer and bulkier and harder to clean, may be the first casualty. The industry is staring at a finite resource, and the clock is ticking faster than most companies admit publicly.
The Legal Minefield

Then there's the compliance problem, which has gone from theoretical to very real in the past eighteen months.
YouTube's CEO made clear back in April 2024 that using YouTube videos to train third-party models would violate the platform's terms of service. The company introduced creator opt-in settings in December 2024, but unauthorized scraping remains explicitly prohibited. That's not just a policy preference—it's enforceable, and YouTube has lawyers who like to remind people of that fact.
The European Union's AI Act brought transparency obligations for general-purpose AI models into effect in August 2025. Dataset providers serving European clients now need stronger provenance documentation, training-data summaries, and verifiable opt-out compliance. The European Commission's consultation on machine-readable text and data mining protocols, running from late 2025 into early 2026, is expected to yield standardized formats for rightsholders to signal reservations. The 2019 Digital Single Market Directive already allows rightsholders to opt out via machine-readable means under Article 4(3). A German court ruling in November 2024—Kneschke v. LAION—emphasized enforceability of those opt-outs, signaling that judicial patience for "we didn't see the opt-out flag" defenses is wearing thin.
In the U.S., SAG-AFTRA's 2025 contracts for interactive media and commercials include explicit consent and disclosure requirements for digital replicas. The union negotiated controls ensuring performers' data can't be used to train AI without consent, with formalized digital replica compensation. That matters for any dataset containing identifiable performers, which is most video datasets outside heavily anonymized or synthetic corpora.
A February 2026 audit of license integrity in the AI supply chain found that 96.5 percent of datasets and 95.8 percent of models lack required license text. That's not an academic footnote. It's a due-diligence nightmare for any lab buying or building training data, especially one eyeing an eventual IPO or acquisition.
The Infrastructure Play

This is the opportunity Shofo and a cohort of competitors are chasing. The founders—Bryan Hong (CEO, Berkeley dropout), Andre Braga (Head of AI, formerly MIT), Braiden Dishman (COO, previously at AWS), and Alexzendor Misra (CTO, previously CEO of Correkt, which reached 43,000 users)—started building in San Francisco in 2025. Their pitch centers on aggregating public and private video sources into what they describe as a continuously updating index of "billions of videos," then delivering custom, cleaned, structured datasets in days rather than the months it typically takes labs to do it themselves.
Early evidence of execution includes a verified Hugging Face organization with a dataset of roughly 58,000 TikTok videos, updated in mid-February 2026. The dataset includes transcripts, hashtags, language metadata, resolution, frame rate, and an "is_ai_generated" flag—signs of a multi-modal pipeline at work. The license is marked "other," which likely means downstream buyers will expect explicit rights provenance and compliance documentation before signing checks.
Shofo isn't operating in a vacuum. Troveo claims to deliver "training-ready video data" to top labs, with over 500,000 ten-second clips delivered and content licensing deals with creators. Sieve positions itself as a "video data research lab" with "hundreds of petabytes of curated video" and Fortune 100 clients, though the specifics remain vague. Versos offers "rights-cleared archives" with provenance tracking and chain-of-custody, and published a blog post on video data licensing in mid-February. Storyblocks, Defined.ai, and Protege each tout ethically sourced or rights-managed video for AI training. Veritone reported processing five trillion tokens from premium video and audio by the second quarter of 2025.
The established players are moving, too. Shutterstock expanded its multi-year licensing partnership with OpenAI in May 2023 to include video and music, not just images. Getty Images is actively commercializing AI data licensing and provides free sample datasets, though it's simultaneously embroiled in IP litigation against Stability AI in the UK—a delicious contradiction that underscores the industry's legal messiness. Synthesia licensed Shutterstock's video library in April 2025 to improve avatar realism. And GoPro launched an opt-in "AI Training Licensing" program in August 2025, collecting over 125,000 hours from subscribers in roughly two weeks with a 50 percent revenue share to contributors.
The infrastructure layer is fragmenting, but a common thread runs through it all: verifiable provenance, explicit licensing, compliance tooling. The era of "just scrape it and figure it out later" is ending, at least for companies that plan to operate at scale in regulated markets.
The Economics Are Tricky

Storage and egress costs for petabyte-scale video corpora aren't trivial. AWS S3 Standard runs roughly $0.023 per gigabyte per month for the first 50 terabytes, with egress commonly around $0.09 per gigabyte for the first 10 terabytes per month. Cloudflare R2 positions itself as egress-free, which could matter for high-access-rate use cases like continuous labeling and curation. Metadata-heavy catalogs can leverage serverless querying tools like Amazon Athena, priced at $5 per terabyte scanned. The unit economics demand efficiency at scale, which means Shofo and its competitors need volume to make the math work.
The bigger question is whether this infrastructure layer becomes defensible. A widely cited 2023 analysis from Andreessen Horowitz argued that generative AI platforms have few structural moats, and that control of distribution and sustained iteration might matter more than static model performance. The same logic could apply to data infrastructure. If Shofo or its competitors can build network effects—more data sources opting in, better curation pipelines, faster turnaround, higher trust—they might entrench. But if data curation becomes commoditized, or if labs decide to vertically integrate by licensing directly from platforms and creators, the window could close quickly.
The market opportunity looks substantial, at least on paper. Grand View Research estimated the AI training dataset market at $3.195 billion in 2025, with forecasts reaching $16.32 billion by 2033—a 22.6 percent compound annual growth rate. The World Trade Report 2025 and World Bank Digital Progress and Trends 2025 cite similar trajectories, with training datasets expected to hit $17.04 billion by 2033. Video is the fastest-growing segment within that market, driven by model improvements, platform integrations, and enterprise adoption.
But growth will be uneven. Companies that can navigate what might be called the "compliance premium"—verifiable chain of custody, machine-readable rights, opt-in pipelines—will likely capture a disproportionate share of demand from risk-averse labs and enterprises. Academic and open datasets will continue to grow, but they're unlikely to cover long-video, high-resolution, interaction-specific, or tightly controlled IP domains. Those will require paid, licensed sources, which is where infrastructure startups hope to win.
What Happens Next
For now, the race is on, and nobody's sure who's winning. Video AI is moving too fast for labs to wait months assembling custom datasets. The compliance burden is too high to ignore provenance and hope regulators look the other way. And the technical demands—long-form coherence, character consistency, domain specificity—are only getting steeper.
Whether Shofo becomes the Common Crawl of video or just one more vendor in a crowded field will depend on execution, partnerships, and whether it can deliver datasets that labs trust both legally and technically. The team is small, the competition is fierce, and the incumbents have deep pockets and existing relationships. But the underlying trend seems clear enough: infrastructure for AI video models is no longer a nice-to-have. It's the next bottleneck, and someone's going to solve it.
The question is whether solving it builds a durable business, or just opens the door for the next wave of commoditization. In an industry that moves as fast as this one, four people with a good idea and Y Combinator backing might be enough. Or it might not. We'll know soon enough—probably sooner than any of us expect.
