Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSMarch 25, 2026

The Race to Archive Human Skill for Robot Training

The Race to Archive Human Skill for Robot Training
RoboticsTraining Data+3
Healthtech & Biotech iconHealthtech & BiotechMarch 25, 2026

AI Models Predict Missing Biology Data to Speed Drug Discovery

AI Models Predict Missing Biology Data to Speed Drug Discovery
Drug DiscoveryAi+3

Founders Mentioned

Bryan Hong

Shofo

saas icon
SaaS

Braiden Dishman

Shofo

saas icon
SaaS

Alexzendor Misra

Shofo

saas icon
SaaS

Bryan Hong

Shofo

saas icon
SaaS

Braiden Dishman

Shofo

saas icon
SaaS

Alexzendor Misra

Shofo

saas icon
SaaS
SaaS iconSaaS
March 25, 2026
YcTraining DataAi InfrastructureComputer VisionB2b Saas

Shofo Launches 'Common Crawl for Videos' for AI Training Data

YC W26 startup debuts platform indexing billions of social videos to deliver custom training datasets for AI labs, addressing surging demand for multimodal model data.

Shofo Launches 'Common Crawl for Videos' for AI Training Data

Training AI on video is expensive, time-consuming, and—increasingly—a data scavenger hunt. Not just any footage will suffice for the teams racing to build the next generation of video-capable models. They need specific clips: particular activities captured from particular angles, stripped of watermarks, tagged with metadata, ready to feed into a training pipeline.

Enter Shofo, a four-person outfit fresh out of Y Combinator that's positioning itself as something between a video librarian and a data broker. The San Francisco startup, which emerged publicly in late January, has built what it describes as "Common Crawl for Videos"—a continuously updated index of billions of short-form social clips that can be sliced, filtered, and packaged into custom training datasets on demand.

The pitch is appealingly simple: Tell us what you need, and we'll pull it from our archive, clean it up, and ship it back model-ready within days. Want 100,000 hours of cooking videos featuring hand-object interactions, no scene changes, all from this year? Shofo says it can do that.

Whether the company can actually deliver at that scale, and at what cost, remains to be seen. But the ambition reflects a broader shift in how AI labs are thinking about training data—and how much they're willing to pay to skip the grinding work of assembling it themselves.

Social Video as Infrastructure

Shofo started by indexing TikTok and is now expanding across major short-form platforms: Instagram Reels, YouTube Shorts, and the like. The system ingests massive volumes of video, then runs it through what the company calls an enrichment pipeline. That means object detection, activity recognition, semantic search, creator demographic filtering, automatic transcription, engagement metrics—the works.

The result is a queryable library that lets teams carve out datasets by content type, language, visual features, even the number of likes a video received. Behind the scenes, the company uses techniques like hashtag-based discovery, platform API indexing, speech recognition for transcripts, and Redis-based deduplication to keep the archive from collapsing under its own weight.

In mid-February, Shofo released a sample dataset on Hugging Face to show off the machinery. The "Shofo TikTok General Small" collection contains 50,000 TikTok videos—roughly 500 gigabytes—alongside metadata that includes transcripts, comments, engagement counts, hashtags, language tags, and flags for AI-generated content or advertisements. The schema is thorough: video paths, creator info, posting dates, resolution, frame rate, duration, even regional language data for comments and overlay text.

According to the Hugging Face repository, the dataset has been downloaded more than 36,000 times since its last update. Shofo describes it as a "curated subset" of a much larger TikTok index containing what the company claims are "hundreds of millions" of videos. Production datasets sold to customers, the team notes, come with heavier labeling and curation than this public research sample.

A Crowded but Fragmented Market

Digital illustration for article section "A Crowded but Fragmented Market" in "Shofo Launches 'Common Crawl for Videos' for AI Training Data" - A conceptual representation of a crowded and fragmented video data market, featuring a central, intr...

Video training data has become a surprisingly hot space. Oxylabs advertises access to 4 million original videos from a million unique channels. NetNut pitches YouTube datasets tailored for enterprise AI work. Waffle Video offers on-demand footage with built-in quality assurance. Most of these focus on YouTube or rely on creator partnerships.

Open research datasets have long been the fallback option. Google's YouTube-8M, released back in 2016, offered 8 million videos and 500,000 hours of content; it's still cited as a benchmark. WebVid-10M, from 2021, became a go-to resource for video-text pretraining. More recent efforts like VideoUFO and HD-VILA-100M have pushed sample counts into the millions, though the latter has drawn scrutiny for allegedly including publisher content without clear authorization.

Shofo's angle is different: short-form social video, specifically. Jared Friedman, a Y Combinator partner, said in a LinkedIn post this spring that he believes Shofo has assembled "the largest [short-form] dataset commercially available." The company itself uses language like "world's largest library of videos" and "billions of videos indexed," though these are self-reported claims without independent verification.

Still, the focus on social platforms makes strategic sense. Short-form video dominates engagement across demographics, and the content is visually dense, heavily tagged, and rich with metadata that traditional datasets lack. For labs training models on human activity, object interactions, or culturally relevant behavior, TikTok and its peers offer diversity that a curated YouTube collection simply doesn't.

The Team Behind the Index

The Shofo founders—Bryan Hong (CEO), Andre Braga (Head of AI), Braiden Dishman (COO), and Alexzendor Misra (CTO)—all have ties to UC Santa Barbara or Berkeley. Braga brings a background in statistics and data science, with prior work at MIT. Dishman came from AWS. Misra was previously CEO of Correkt, a multimodal AI search engine that reportedly reached 40,000 users before the team pivoted.

That pivot is central to understanding what Shofo actually is. Correkt was about indexing and searching multimodal content for end users; Shofo is essentially the same infrastructure, repurposed for AI training data instead of consumer search. The team already knew how to scrape, index, and process huge volumes of social media content at scale. Shifting the customer from everyday users to AI labs appears to have been the move that caught Y Combinator's eye.

Public business records show the company incorporated in California in May of last year. Since then, it's been heads-down building, emerging publicly only with the Y Combinator launch.

The Copyright Question No One Wants to Answer

Digital illustration for article section "The Copyright Question No One Wants to Answer" in "Shofo Launches 'Common Crawl for Videos' for AI Training Data" - A conceptual and minimal 3D rendered image representing the complex issue of video copyright, featur...

The unavoidable issue for any company indexing and reselling social video is copyright—and platform terms of service. Shofo's Hugging Face dataset includes a carefully worded disclaimer: the company "does not claim ownership of the underlying video content," and users "must ensure compliance with applicable copyright laws and platform terms."

It's a common stance among dataset vendors, but it doesn't exactly eliminate risk. A report last fall by iMEdD Lab found that major news publishers, including The New York Times, had their YouTube videos scraped and included in large training datasets without authorization. The Times explicitly stated it did not consent to third-party use of its content for AI training purposes.

As regulatory scrutiny around training data intensifies, dataset vendors may face mounting legal or reputational pressure over sourcing practices. Shofo hasn't publicly detailed how it handles licensing or permissions for the social content it indexes. The company's materials emphasize that datasets are intended for research and that users bear responsibility for compliance—a strategy that may work for now, but could become less tenable if governments or platforms decide to crack down.

What You Can Actually Buy

Digital illustration for article section "What You Can Actually Buy" in "Shofo Launches 'Common Crawl for Videos' for AI Training Data" - A single, beautifully crafted retro-futuristic mechanical box sitting on a clean, uncluttered surfac...

Shofo is taking requests at [email protected]. The company's homepage and launch materials emphasize fast turnaround on custom datasets, though pricing and minimum order details aren't publicly available. Earlier LinkedIn posts from the founders hinted at pricing around $0.0001 per post for raw social data, but that may have been tied to the earlier Correkt product or pilot programs rather than current offerings.

For teams building multimodal models, the value proposition is access to a type of data that's genuinely hard to assemble in-house: diverse, engagement-filtered, pre-tagged short-form video at scale. Whether Shofo's index is truly the largest commercially available is difficult to verify. But for labs that need tens of thousands of hours of specific footage delivered quickly, it's one of the few vendors claiming it can pull that off.

The company's Hugging Face presence suggests a commitment to technical credibility. The sample dataset is well-documented, with clear schema definitions and transparent collection methodology. That level of detail matters to data scientists evaluating vendors, and it's a baseline Shofo seems intent on maintaining.

The bet here is straightforward: As video-capable AI models proliferate, demand for custom training corpora will surge. The infrastructure is built. The index is growing. Now it's just a question of how many labs are willing to pay for it—and whether the legal ground beneath the entire industry stays solid long enough for Shofo to find out.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • The Race to Archive Human Skill for Robot Training
  • AI Models Predict Missing Biology Data to Speed Drug Discovery
  • The Quick Commerce Shakeout: Why India Scaled While Europe Retreated
  • Open-Source Visual Testing Tools Emerge for AI Coding Agents
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.