Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 20, 2026

Mining Parasite Biology for Breakthrough Autoimmune Therapies

Mining Parasite Biology for Breakthrough Autoimmune Therapies
YcDrug Discovery+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 20, 2026

The Race to Build Digital Humans: Inside the $7B In-Silico Trial Boom

The Race to Build Digital Humans: Inside the $7B In-Silico Trial Boom
Drug DiscoveryClinical Trials+3
SaaS iconSaaS
February 20, 2026
YcArtificial IntelligenceData InfrastructureComputer VisionTraining Data

Shofo Builds 'Common Crawl for Video' to Feed AI Model Hunger

YC-backed startup indexes 100M+ short-form videos as AI labs scramble for training data amid platform crackdowns and regulatory pressure on web scraping.

Shofo Builds 'Common Crawl for Video' to Feed AI Model Hunger

There's a particular irony to training artificial intelligence on videos of people doing everyday things. The AI labs building tomorrow's video generation models desperately need footage of humans cooking, dancing, arguing, and teaching their dogs to skateboard. But the platforms where billions of those moments live—TikTok, YouTube, Instagram—are increasingly off-limits, legally speaking.

Which is precisely the problem Shofo thinks it can solve.

The four-person, Y Combinator-backed startup bills itself as something like "Common Crawl for video," though that comparison carries more weight than the founders might intend. It's offering AI companies what they can no longer easily grab themselves: a curated, indexed repository of more than 100 million TikTok videos, delivered through SQL-ready APIs and priced by the record. Pre-cleaned. Pre-segmented. Ready to feed into training pipelines.

Whether that business model survives contact with platform terms of service, emerging regulations, and the legal minefields of scraping social media remains an open question. But Shofo's pitch reflects a broader truth reshaping the AI industry: the training data feeding frenzy that powered text and image models is colliding with reality in the video space.

When the Firehose Gets Shut Off

The math is straightforward, if staggering. YouTube sees more than 20 million video uploads daily. Shorts alone generate over 200 billion views every 24 hours. That's an almost incomprehensible volume of human activity, most of it exactly the kind of diverse, real-world footage that video AI models crave.

But volume doesn't equal access.

YouTube CEO Neal Mohan put it plainly in 2024: using the platform's content for AI training without permission "would be a problem" and a clear violation of terms of service. TikTok's privacy center outlines active anti-scraping enforcement. Both platforms offer research APIs, technically, but with restrictions that make them impractical for the scale most AI labs require. It's a little like offering someone dying of thirst a thimble of water.

The crackdown is measurable. Cloudflare's 2025 Radar report found that AI bots now represent roughly 4.2 percent of HTML requests, with rapid growth from identifiable crawlers like GPTBot and ClaudeBot. Between July and December 2025 alone, Cloudflare blocked 416 billion bot requests. The gap between what AI companies want to crawl and what platforms will allow is widening, fast.

Regulatory pressure is mounting, too. The EU AI Act's obligations for general-purpose AI models took effect August 2, 2025, requiring detailed training data summaries and mechanisms to respect copyright opt-outs. The Digital Services Act's Article 40 researcher data access provisions are now live. Even preliminary breach findings against Meta and TikTok signal European regulators aren't bluffing about transparency requirements.

The old playbook—"just scrape it"—is breaking down. Perhaps faster than the industry expected.

The Hunger Pains

Video models consume data and compute at rates that would have seemed absurd a few years ago. Research published at CVPR 2025 confirmed what insiders suspected: scaling laws for video diffusion transformers follow predictable patterns. More data, cleaner data, diverse data—all translate directly into better output quality. The Open-Sora 2.0 project argued that a commercial-grade video model could be built for around $200,000 if you optimize the data pipeline properly.

That kind of efficiency drives intense demand.

Nearly 40 percent of video advertisements are expected to use generative AI by 2026, up from 22 percent in 2024, according to IAB and Insider Intelligence. Google's Veo has gone through three major iterations in rapid succession, adding vertical 9:16 support and slashing API prices to enable scale. Runway is on Gen-4.5 and building world models. Kuaishou's Kling, Pika, and Luma's Dream Machine are all racing to claim market share.

The industry is eating itself trying to keep up. Andreessen Horowitz's 2026 State of Generative Media report found that enterprises use a median of 14 different generative media models—far more fragmented than the concentrated large language model market. That fragmentation creates demand for specialized datasets. A model optimized for cooking tutorials needs different training material than one focused on fashion, or sports highlights, or how to fix a leaky faucet.

But where does that material come from when the platforms have locked the gates?

Enter the Intermediaries

Digital illustration for article section "Enter the Intermediaries" in "Shofo Builds 'Common Crawl for Video' to Feed AI Model Hunger" - A surreal blue dreamscape visualizing the concept of data intermediaries and massive digital dataset...

Shofo's public sample on Hugging Face offers a glimpse of the value proposition: 50,000 TikTok videos under an MIT license, complete with metadata, transcripts, comments, and engagement metrics. The company describes this as a subset of an index containing "hundreds of millions" of videos. Pricing is structured around APIs rather than bulk downloads, with endpoints covering feed data, profiles, hashtags, and comments. Custom datasets can be ordered by topic—"cooking videos with hand-object interactions from 2025" is one example the founders cite.

The founders frame the problem in practical terms on their Y Combinator profile: too many teams are "spending months turning unstructured social content into model-ready data." Shofo wants to shortcut that process. It's a model that mirrors the trajectory of text data providers, adapted for a medium that's considerably harder to wrangle. Video files are larger, metadata is richer, and the filtering required—NSFW content, illegal material, privacy risks involving faces and voices—is vastly more complex.

The focus on short-form social content is strategic. These platforms generate massive volumes of diverse footage across nearly every conceivable human activity. They're also the platforms most aggressively blocking scrapers, which creates the market gap Shofo is trying to fill.

But here's the tension: Shofo is scraping platforms that explicitly prohibit automated access. Whether the company has licensing agreements or is relying on aggressive interpretations of fair use and research exceptions remains unclear. The public materials emphasize accessing "hard to get social media posts" without detailing the legal framework.

The Fragmented Landscape

Shofo enters a market already populated with alternatives, though most pursue different strategies.

Academic datasets have been foundational. HowTo100M offers 136 million clip-caption pairs from YouTube instructional videos, representing roughly 15 years of video time. WebVid-10M provides 10.7 million video-text pairs across 52,000 hours and was used to train Meta's Make-A-Video. Ego4D offers 3,670 hours of first-person footage under a consortium with strict privacy controls. ACAV100M delivers 100 million 10-second clips focused on audio-visual correspondence.

These datasets suffer from predictable problems, though. Link rot is endemic—the Kinetics-700 dataset from DeepMind acknowledges persistent issues with missing videos as YouTube links break over time. Academic datasets also lag platform trends. Short-form vertical video from TikTok, Instagram Reels, and YouTube Shorts represents a format shift that older datasets don't capture well. The world moved on; the data didn't.

Licensed sources offer legal certainty at a price. Shutterstock signed a six-year deal with OpenAI covering images, video, music, and metadata. Adobe's Firefly Video Model markets itself as "commercially safe," trained exclusively on licensed Adobe Stock footage and public domain content. Adobe and Runway announced a partnership in late 2025, integrating Runway's Gen-4.5 into Firefly workflows.

These deals provide peace of mind but at significant cost, and with content that skews heavily toward professional stock footage rather than organic social media. There's a world of difference between a polished Adobe Stock clip of someone cooking and a genuine TikTok video of someone burning dinner while their cat knocks over a plant in the background. AI models trained on the latter will understand the world better than models trained on the former.

Open-source tooling has matured, at least. LAION's video2dataset demonstrated the ability to assemble 590 million video-text pairs using 16-core EC2 instances, processing WebVid-10M in roughly 12 hours. That pipeline includes deduplication, optical flow calculation, and scene detection. But even sophisticated tooling doesn't solve the core problem: accessing the source material without violating terms of service or running afoul of regulators.

The Legal Minefield

Digital illustration for article section "The Legal Minefield" in "Shofo Builds 'Common Crawl for Video' to Feed AI Model Hunger" - A conceptual and surreal illustration depicting the complex legal landscape of web scraping, featuri...

The hiQ v. LinkedIn case offered some clarity in 2022, with the Ninth Circuit ruling that scraping public web pages likely doesn't violate the Computer Fraud and Abuse Act. But that decision doesn't shield against contract claims based on terms of service, and later consent judgments restricted hiQ's activities anyway. The legal landscape is muddier than most AI companies would prefer.

YouTube has introduced an opt-in program allowing creators to license their content for third-party AI training, but it's off by default. How many creators will actively enable that setting? Early indications suggest not many.

TikTok faces geopolitical risk in the United States following the Protecting Americans from Foreign Adversary Controlled Applications Act, which creates uncertainty about relying on TikTok data pipelines long-term. Building a business on scraping a platform that might be banned—or dramatically restructured—is risky.

Then there's the Common Crawl controversy from 2025, which serves as a cautionary tale. Investigations alleged that the nonprofit's archives included paywalled news content, highlighting the risks of "open crawl" pipelines. For video, those risks amplify. Faces, voices, minors, personally identifiable information—all raise stakes that go beyond copyright infringement into privacy law territory.

What Happens Next

Digital illustration for article section "What Happens Next" in "Shofo Builds 'Common Crawl for Video' to Feed AI Model Hunger" - A conceptual illustration depicting the fragmentation of the AI industry as it seeks video training ...

The question facing AI labs isn't whether they need more video training data. They clearly do. The question is how to get it legally, ethically, and at scale—preferably all three simultaneously.

The industry is fragmenting along predictable lines. Some companies are pursuing aggressive scraping strategies and hoping the legal system moves slowly. Others are negotiating licensing deals with platforms and stock providers. A third group is betting on synthetic data or carefully curated public domain sources. And then there are the intermediaries like Shofo, positioning themselves as the solution to a problem they didn't create but are happy to profit from.

Shofo's success will depend on threading a very narrow needle: delivering data that's genuinely useful while navigating contractual risks, regulatory scrutiny, and the possibility that platforms crack down even harder. The company is operating in a legal gray zone—not quite criminal under CFAA, perhaps, but contractually questionable at best.

What seems certain is that the training data landscape for video AI has shifted permanently. The open frontier that powered early text and image model development is closing, if it isn't already shut. Licensed corpora, controlled APIs, creator opt-ins, and compliance-focused intermediaries are becoming the new normal, whether AI companies like it or not.

Video model developers face a choice, really. Build relationships with rights holders. Pay for curated datasets. Or risk legal exposure by scraping anyway and hoping the platforms don't notice, the regulators don't care, and the lawsuits don't land.

The data gold rush for video isn't over. But the rules of the game are changing faster than most participants expected, and the winners will be whoever adapts quickest—or lawyers up best.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • Mining Parasite Biology for Breakthrough Autoimmune Therapies
  • The Race to Build Digital Humans: Inside the $7B In-Silico Trial Boom
  • Contra Launches AI Agent Commerce Platform, Eyes Creator Economy
  • Google Launches Gemini 3.1 Pro with Doubled Reasoning Performance
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.