The lawsuits arrived in Illinois courtrooms this May—nine class-action complaints, all filed under the state's notoriously strict biometric privacy law, all alleging the same basic wrong. Companies from ElevenLabs to Amazon, the suits claim, "stole" voiceprints to train AI models without bothering to ask permission first. Meta, Google, Apple, Microsoft, NVIDIA, Samsung, Adobe—the defendant list reads like a who's-who of tech giants suddenly facing the awkward question of where, exactly, all that training data came from.
Into that legal thicket walks Datoric, a two-person startup fresh out of Y Combinator's Summer 2026 batch, with a pitch that sounds almost quaint: What if you just paid for clean data instead?
In late July, Datoric launched its first catalog of licensed training datasets—voice, video, robotics data, all of it collected with what the company insists is explicit contributor consent and documentation to prove it. The timing, to put it mildly, isn't accidental. Nor is the company's core promise: buy datasets where the provenance and licensing terms are bundled from the start, rather than bolted on later when the lawyers start circling.
Whether that proposition finds buyers is the open question. But the legal backdrop is certainly getting noisier.
The Catalog: Four Datasets, Each With a Story
Between July 19 and July 21, Datoric released four production datasets, each targeting a different corner of the AI buildout. The specs, verified on the respective datasheets, suggest the company is thinking beyond voice cloning tutorials.
The Computer-Use Agent Traces dataset is the most unusual. It contains 250,000 traces spanning roughly 15,000 hours of screen activity—around 10 million timestamped UI actions across more than 100 task categories. Think autonomous agents, GUI automation, anything that requires teaching a model how humans actually interact with software. It's niche, but it's also the kind of data you can't easily scrape from the open web without creating a surveillance nightmare.
The TTS Conversational Voice set is more familiar territory: 20,000 hours of 48 kHz WAV audio covering 30-plus languages, complete with transcripts, speaker labels, and technical metrics like signal-to-noise ratio. This is the stuff text-to-speech models eat for breakfast, and it's also the category where the Illinois lawsuits hit hardest.
For multimodal work, Datoric offers the Audio-Video Conversation dataset—4,000 hours of 1080p-plus video synced with high-fidelity audio in over 20 languages. It's labeled for facial visibility, gestures, emotion, laughter, overlapping speech. The kind of granular annotation that makes or breaks a model's ability to understand context beyond words.
Then there's the Egocentric Residential Video collection, which might be the most ambitious: 100,000 hours of first-person footage capturing household tasks across 80 categories that aren't cooking. (Why not cooking? The company doesn't say, but one imagines legal or practical constraints around food-handling data.) The dataset targets robotics and embodied AI, annotated for task phases, hand-object interactions, state changes—plus privacy redaction where needed, because filming inside people's homes for 100,000 hours raises obvious questions.
All four live on Hugging Face behind gated repositories. You request access; Datoric decides whether to grant it. Production delivery happens per engagement, not open download. It's a far cry from the free-for-all ethos of early machine learning.
Consent Before Recording, Not After
Datoric's pitch hinges less on scale—Scale AI isn't exactly losing sleep over a two-person competitor—and more on workflow. The company describes a six-step pipeline that's deliberately unsexy: define the data spec, source consented contributors, collect "at the origin" (no scraping), conduct review and QA, package a pilot, deliver with licensing rights attached.
That "at origin" collection is the legal linchpin. Rights documentation is generated alongside the data itself, creating what the company calls a chain of custody that holds up under audit. According to Datoric's data-rights page, every sample includes traceable origin records, clear licensing terms, and documented redaction steps where personally identifiable information appears.
It's a departure—perhaps more than the founders expected when they started—from the retrofitted licensing approach that's become industry standard. Typically, datasets get assembled first, often from scraped sources, and consent frameworks are negotiated later. If at all. Datoric's stated methodology flips that: lock down contributor agreements before the first frame is recorded, then attach those agreements to the delivered data packages like a legal receipt.
Whether that's paranoia or prudence depends on how the Illinois cases shake out.
Security as Sales Pitch

Datoric also bills itself as "security-first training data," which translates to architecture choices that sound almost paranoid. No public task listings. No contributor APIs exposed to the open internet. No shared annotation pools where your proprietary dataset might mingle with someone else's.
Instead, the company operates isolated, customer-specific workspaces for each project. Dataset access on Hugging Face requires approval requests; PII redaction protocols are published; a security contact sits at [email protected]. The privacy and terms-of-service pages—both last updated April 7, 2026—detail data handling, retention policies, and quality metrics under Delaware law.
That private-by-design posture serves dual purposes. It reduces the attack surface for data leakage, obviously. But it also creates a compliance-friendly audit trail for companies navigating FTC guidance on voice and video training—guidance that got considerably sharper after the agency's 2023 action against Amazon over Alexa data retention. Having documented consent and redaction processes baked into the supply chain may prove more defensible than post-hoc compliance bolted onto scraped datasets.
At least, that's the bet.
A Crowded Field, or a New Category?
Datoric isn't alone in chasing licensed training data, though the landscape is still taking shape. Scale AI and Universal Robots announced plans in March 2026 to release a large-scale industrial dataset collected on UR robots later this year. Defined.ai expanded its "physical AI" robotics offerings in January. Mozilla Common Voice released new speech datasets—Scripted v26.0 and Spontaneous v4.0—in June, governed by community licensing. Voices.com has positioned itself as a provider of "governed" AI voice datasets with consent and traceability, trying to reframe rights-cleared voice data as an industry standard rather than a premium offering.
What sets Datoric apart, in theory, is bundling. Most providers specialize—voice or robotics or video, rarely all three with agent traces thrown in. Datoric's bet is that teams building multimodal foundation models or embodied agents want fewer contracts to negotiate, not just cleaner data. One vendor, four modalities, single compliance framework.
The company also released research alongside its datasets, perhaps hedging against the "just a data vendor" label. VideoTruth-Bench, published June 24 with a July 21 report edition, evaluates six frontier vision-language models—Claude 4.5, Gemini 2.5 Flash and Pro, GPT-4o among them—on contradiction detection, temporal ordering, and hallucination resistance. The benchmark flags significant sycophancy gaps: 24 to 60 percentage points, suggesting current models struggle badly with adversarial or contradictory video inputs.
For Datoric, publishing evaluation tools positions the company as a partner in model safety, not just a supplier. Whether customers see it that way is another matter.
Two Founders, Zero Named Customers

Nikhil Reddy (CEO) and Jeffrey Lin (CTO) founded Datoric in 2025 before entering Y Combinator's Summer 2026 batch. Reddy studied math, economics, and computer science at the University of Chicago and lists prior quant and software engineering internships. His Y Combinator profile includes a colorful detail: he "bypassed Google OAuth 2.1 and QA testers to automate 15,000-plus hours on paid data annotation sites in early 2023." It's the kind of hustle that suggests firsthand familiarity with the labor economics—and the legal gray zones—of data collection.
Lin brings a background in math, computer science, and robotics from NYU, plus AI/ML and software engineering experience. The team size was listed as two on Y Combinator's directory as of Summer 2026, though that figure may well have changed since. Harshita Arora is listed as the primary YC partner.
No customers, partners, or external funding rounds have been disclosed beyond the Y Combinator investment. The company has published no formal press release or Launch HN post—its public presence consists of the Y Combinator listing, a polished website, and the July dataset releases.
One detail worth noting: the Hugging Face dataset contact lists an email domain that differs from Datoric's primary branding. It could be a placeholder, a personal alias, or just startup scrappiness. Hard to say.
Infrastructure or Premium Niche?
The broader question Datoric raises is whether licensed training data becomes infrastructure—something every serious AI lab budgets for—or remains a premium niche for the compliance-obsessed.
Academic research published in February 2026 found widespread non-compliance with attribution and license propagation in datasets and models. A U.S. Copyright Office report from 2025 noted the growth of negotiated licensing in AI training, though it stopped short of predicting how fast or how far that trend would go. If litigation costs and regulatory scrutiny push more companies toward provenance-documented data, Datoric's early positioning could pay off handsomely.
But if fair-use defenses hold—and if most teams continue scraping with legal risk budgeted in as a cost of doing business—the market for consent-first datasets may stay narrow. Perhaps very narrow.
For now, Datoric is live. The datasets are gated but accessible. The legal backdrop is getting louder, and the Illinois lawsuits are winding their way through discovery. Whether that translates to traction for a two-person startup betting on compliance over scale will depend on how many teams decide clean licenses are worth the premium.
And, perhaps more importantly, whether the courts agree that scraping voiceprints without permission is actually theft.
