Training AI on video is expensive, time-consuming, and—increasingly—a data scavenger hunt. Not just any footage will suffice for the teams racing to build the next generation of video-capable models. They need specific clips: particular activities captured from particular angles, stripped of watermarks, tagged with metadata, ready to feed into a training pipeline.
Enter Shofo, a four-person outfit fresh out of Y Combinator that's positioning itself as something between a video librarian and a data broker. The San Francisco startup, which emerged publicly in late January, has built what it describes as "Common Crawl for Videos"—a continuously updated index of billions of short-form social clips that can be sliced, filtered, and packaged into custom training datasets on demand.
The pitch is appealingly simple: Tell us what you need, and we'll pull it from our archive, clean it up, and ship it back model-ready within days. Want 100,000 hours of cooking videos featuring hand-object interactions, no scene changes, all from this year? Shofo says it can do that.
Whether the company can actually deliver at that scale, and at what cost, remains to be seen. But the ambition reflects a broader shift in how AI labs are thinking about training data—and how much they're willing to pay to skip the grinding work of assembling it themselves.
Social Video as Infrastructure
Shofo started by indexing TikTok and is now expanding across major short-form platforms: Instagram Reels, YouTube Shorts, and the like. The system ingests massive volumes of video, then runs it through what the company calls an enrichment pipeline. That means object detection, activity recognition, semantic search, creator demographic filtering, automatic transcription, engagement metrics—the works.
The result is a queryable library that lets teams carve out datasets by content type, language, visual features, even the number of likes a video received. Behind the scenes, the company uses techniques like hashtag-based discovery, platform API indexing, speech recognition for transcripts, and Redis-based deduplication to keep the archive from collapsing under its own weight.
In mid-February, Shofo released a sample dataset on Hugging Face to show off the machinery. The "Shofo TikTok General Small" collection contains 50,000 TikTok videos—roughly 500 gigabytes—alongside metadata that includes transcripts, comments, engagement counts, hashtags, language tags, and flags for AI-generated content or advertisements. The schema is thorough: video paths, creator info, posting dates, resolution, frame rate, duration, even regional language data for comments and overlay text.
According to the Hugging Face repository, the dataset has been downloaded more than 36,000 times since its last update. Shofo describes it as a "curated subset" of a much larger TikTok index containing what the company claims are "hundreds of millions" of videos. Production datasets sold to customers, the team notes, come with heavier labeling and curation than this public research sample.
A Crowded but Fragmented Market

Video training data has become a surprisingly hot space. Oxylabs advertises access to 4 million original videos from a million unique channels. NetNut pitches YouTube datasets tailored for enterprise AI work. Waffle Video offers on-demand footage with built-in quality assurance. Most of these focus on YouTube or rely on creator partnerships.
Open research datasets have long been the fallback option. Google's YouTube-8M, released back in 2016, offered 8 million videos and 500,000 hours of content; it's still cited as a benchmark. WebVid-10M, from 2021, became a go-to resource for video-text pretraining. More recent efforts like VideoUFO and HD-VILA-100M have pushed sample counts into the millions, though the latter has drawn scrutiny for allegedly including publisher content without clear authorization.
Shofo's angle is different: short-form social video, specifically. Jared Friedman, a Y Combinator partner, said in a LinkedIn post this spring that he believes Shofo has assembled "the largest [short-form] dataset commercially available." The company itself uses language like "world's largest library of videos" and "billions of videos indexed," though these are self-reported claims without independent verification.
Still, the focus on social platforms makes strategic sense. Short-form video dominates engagement across demographics, and the content is visually dense, heavily tagged, and rich with metadata that traditional datasets lack. For labs training models on human activity, object interactions, or culturally relevant behavior, TikTok and its peers offer diversity that a curated YouTube collection simply doesn't.
The Team Behind the Index
The Shofo founders—Bryan Hong (CEO), Andre Braga (Head of AI), Braiden Dishman (COO), and Alexzendor Misra (CTO)—all have ties to UC Santa Barbara or Berkeley. Braga brings a background in statistics and data science, with prior work at MIT. Dishman came from AWS. Misra was previously CEO of Correkt, a multimodal AI search engine that reportedly reached 40,000 users before the team pivoted.
That pivot is central to understanding what Shofo actually is. Correkt was about indexing and searching multimodal content for end users; Shofo is essentially the same infrastructure, repurposed for AI training data instead of consumer search. The team already knew how to scrape, index, and process huge volumes of social media content at scale. Shifting the customer from everyday users to AI labs appears to have been the move that caught Y Combinator's eye.
Public business records show the company incorporated in California in May of last year. Since then, it's been heads-down building, emerging publicly only with the Y Combinator launch.
The Copyright Question No One Wants to Answer

The unavoidable issue for any company indexing and reselling social video is copyright—and platform terms of service. Shofo's Hugging Face dataset includes a carefully worded disclaimer: the company "does not claim ownership of the underlying video content," and users "must ensure compliance with applicable copyright laws and platform terms."
It's a common stance among dataset vendors, but it doesn't exactly eliminate risk. A report last fall by iMEdD Lab found that major news publishers, including The New York Times, had their YouTube videos scraped and included in large training datasets without authorization. The Times explicitly stated it did not consent to third-party use of its content for AI training purposes.
As regulatory scrutiny around training data intensifies, dataset vendors may face mounting legal or reputational pressure over sourcing practices. Shofo hasn't publicly detailed how it handles licensing or permissions for the social content it indexes. The company's materials emphasize that datasets are intended for research and that users bear responsibility for compliance—a strategy that may work for now, but could become less tenable if governments or platforms decide to crack down.
What You Can Actually Buy

Shofo is taking requests at [email protected]. The company's homepage and launch materials emphasize fast turnaround on custom datasets, though pricing and minimum order details aren't publicly available. Earlier LinkedIn posts from the founders hinted at pricing around $0.0001 per post for raw social data, but that may have been tied to the earlier Correkt product or pilot programs rather than current offerings.
For teams building multimodal models, the value proposition is access to a type of data that's genuinely hard to assemble in-house: diverse, engagement-filtered, pre-tagged short-form video at scale. Whether Shofo's index is truly the largest commercially available is difficult to verify. But for labs that need tens of thousands of hours of specific footage delivered quickly, it's one of the few vendors claiming it can pull that off.
The company's Hugging Face presence suggests a commitment to technical credibility. The sample dataset is well-documented, with clear schema definitions and transparent collection methodology. That level of detail matters to data scientists evaluating vendors, and it's a baseline Shofo seems intent on maintaining.
The bet here is straightforward: As video-capable AI models proliferate, demand for custom training corpora will surge. The infrastructure is built. The index is growing. Now it's just a question of how many labs are willing to pay for it—and whether the legal ground beneath the entire industry stays solid long enough for Shofo to find out.
