There's a quiet desperation spreading through the world of artificial intelligence, and Hub.xyz thinks it has found a way to profit from it.
The startup—a recent graduate of Y Combinator's latest batch—is selling something increasingly precious in the age of large language models: fresh training data. Not scraped from the public internet, already exhausted by every major AI lab. Not synthetic, the controversial workaround that some researchers warn could poison future models. Real data. Video of people cooking breakfast. Audio recordings in Swahili and Tamil. Images captured at 1080p resolution, rights cleared, ready to download through a POST request.
The pitch is elegantly simple. Developers call an API. Hub.xyz returns structured datasets within minutes—or so the company claims. Egocentric video, multimodal audio, hundreds of thousands of image frames spanning dozens of environments. The kind of embodied, real-world content that text-focused foundation models can't scrape from Wikipedia or Reddit.
Whether anyone is actually buying remains an open question.
Names Without Proof
According to posts on the company's LinkedIn account earlier this year, Hub.xyz lists Google, Adobe, and DeepMind among its early customers, alongside smaller outfits like Rime and LiveKit. It's an impressive roster—if true.
No case studies have appeared. No press releases from the customers themselves. No independent confirmation of any kind.
This isn't necessarily damning. Startups often work with major enterprises under NDAs long before they can announce partnerships publicly. And it's entirely plausible that a scrappy 10-person team could land pilot projects with AI labs desperate for novel training data. But absent corroboration, the customer list exists in a familiar gray zone: too specific to ignore, too unverified to accept at face value.
What can be confirmed is more modest. Hub.xyz raised $1.7 million in a pre-seed round as of September 2025, with no additional public updates by May 2026. The company expanded to Palo Alto last fall, citing proximity to Fortune 500 partners. YC partner Harj Taggar backed the company. The team includes co-founders Armin Kiani and Tim Sprecher, who serves as CEO, along with Jeff Kanai—a former Google and Bright Data executive hired as Head of Sales during the U.S. push.
The funding is now nearly a year old. No follow-on round has been announced post-Demo Day, which could mean the company is raising quietly or hasn't needed to. Or it could mean something else entirely.
The Catalog

Hub.xyz's website offers a window into what it's actually selling, even if the specifics remain somewhat opaque. The company advertises 54,000 hours of egocentric footage across 84 environments. Another 2.4 million image frames. Seventy thousand video clips totaling 18,000 hours. Eighty thousand hours of audio in 47 languages.
Impressive numbers—though these are marketing figures without independent verification. The company treats them as promotional claims rather than audited metrics, which is standard practice for early-stage startups but leaves room for inflation or creative accounting.
The API itself is documented—barely—through third-party integration platforms like Lava Gateway, which lists a single example endpoint: POST https://api.hub.xyz/v1/datasets. Developers supply credentials; the gateway proxies requests. Beyond that? Opacity. No public OpenAPI specification. No sandbox keys available without contacting sales.
For a company positioning itself as API-first, the lack of a robust developer portal is conspicuous. Perhaps it's protecting trade secrets. Perhaps the product is still too nascent for full public documentation. Either way, it raises questions about how ready the platform really is for enterprise-scale adoption.
The Contributor Model

Here's where Hub.xyz diverges from competitors like Scale AI or Appen: a distributed network of individual contributors, supposedly spread across 150 countries, who submit images, audio clips, and video through an internal platform in exchange for USDC payments.
The company also claims partnerships with 73 small and medium-sized businesses across 10 industries, though again, no list of participants has been made public.
There's a wrinkle. Different versions of Hub.xyz's marketing copy have cited varying contributor numbers—ranging from "500K+ contributors across 100+ countries" to "verified active contributors across 150 countries"—without reconciling the discrepancy or providing auditable counts. Is the higher number aspirational? Outdated? Conflating waitlist signups with active participants? The inconsistency suggests marketing materials updated in pieces rather than a cohesive, fact-checked narrative.
The company opened a waitlist for contributors earlier this year and has signaled plans to expand access broadly "within weeks," though no specific timeline has materialized.
On compliance, Hub.xyz says the right things. Its Master Service Agreement, last updated in September, outlines roles under GDPR, UK GDPR, and CCPA. Site copy promises "rights-cleared datasets," full provenance tracking, and consensus-based quality assurance involving both AI and human annotators. These are contractual commitments, not third-party certifications—important, but not independently verified.
A Crowded Field

Hub.xyz is hardly the only company chasing this opportunity, and perhaps not even the best-positioned.
Scale AI launched its Physical AI Data Engine in March, announcing partnerships with robotics manufacturers like Universal Robots and touting a global collector network. Appen, a veteran in the data annotation space, continues to offer multimodal datasets with an emphasis on licensing and provenance—two things enterprises care deeply about. Human API, which debuted earlier this year, markets itself as the first platform where AI agents can hire humans directly for data tasks.
Then there are the scrappier plays. Reporting from earlier this year described startups like Micro1 recruiting thousands of contractors across dozens of countries to film themselves performing household chores—160,000 hours of egocentric video per month, according to coverage syndicated by CNN.
The underlying thesis uniting all of them is the same: text is running out. Epoch AI research estimates the stock of quality public text at roughly 300 trillion tokens and projects exhaustion sometime in the next several years under current scaling trends. Web scraping worked brilliantly for GPT-3 and its descendants. It won't work forever.
The next generation of models demands something richer. Video. Audio. Embodied data capturing how humans navigate physical spaces, manipulate objects, speak across dialects. The kind of information you can't pull from a Common Crawl dump.
Hub.xyz's bet is that developers would rather call an API than build collection infrastructure themselves. That's probably correct, assuming the API delivers what it promises.
What's Missing
A few things remain stubbornly unclear.
The customer list is self-reported and unverified. The contributor count fluctuates across marketing surfaces without explanation. The API documentation available publicly consists of a single endpoint example on a third-party gateway—not exactly confidence-inspiring for developers evaluating whether to integrate.
Pricing is undisclosed. Service-level agreements aren't published. Data quality benchmarks? Absent. The company hasn't shared how it validates contributor submissions, enforces rights clearances at scale, or handles edge cases like deepfakes or synthetic media creeping into the pipeline.
And then there's the fundraising silence. An eight-month-old pre-seed round suggests Hub.xyz is either operating on remarkably lean burn or preparing to raise soon. For a capital-intensive business model—paying thousands of contributors, maintaining API infrastructure, hiring sales teams—$1.7 million doesn't go far.
A Functional Bet
What's not in dispute is the problem. Data scarcity is real. The solutions being built—whether by Scale, Appen, or upstarts like Hub.xyz—represent infrastructure bets on how the next generation of AI models will be trained.
Hub.xyz's API-first, contributor-driven model appears elegant in theory. Execution is another matter. Can the company scale collection without sacrificing compliance? Maintain quality as contributor volume grows? Deliver on the promise of "two minutes to first data" under enterprise query loads?
For now, the platform is live—at least in some form. Developers willing to sign up can test whether the API holds up in practice, assuming they can navigate the sparse documentation and opaque pricing.
The company's real challenge isn't technical. It's trust. In a market where enterprises are already skittish about data provenance, liability, and regulatory exposure, unverified customer claims and inconsistent marketing copy aren't reassuring. Hub.xyz has the right narrative. Whether it has the substance to match remains to be seen.
