The timing was either impeccable or reckless, depending on who you ask.
In mid-February, as the New York Times copyright lawsuit against OpenAI entered a more contentious phase and Anthropic agreed to settle a books-related case for $1.5 million, William Namgyal and Inigo Lenderking—both Berkeley dropouts—stepped out of Y Combinator with a pitch that felt almost too obvious: What if AI companies just bought data they were legally allowed to use?
The same month, an arXiv paper landed with a uncomfortable statistic for the industry: 96.5% of AI training datasets apparently lack proper license text or meaningful copyright compliance. For Namgyal and Lenderking, that wasn't just an academic finding. It was a market gap.
Their company, Luel, sells multimodal training data—speech, sensor readings, video—packaged with something rarely seen in the wild-west world of AI datasets: consent logs, chain-of-title documentation, and audit trails that, in theory, could withstand courtroom scrutiny. No scraped web content. No legal ambiguity. No midnight calls from corporate counsel three years down the line.
Whether it works at scale remains very much an open question.
The Mechanics of Consent
Luel operates as a two-sided marketplace, though calling it that undersells the compliance apparatus bolted onto each transaction. AI teams submit specifications—modality, scenario, device requirements, quality benchmarks. The company posts those requests to what it claims is a network of more than 3 million vetted contributors. Those contributors then upload content or record conversations, but only after obtaining explicit consent from everyone involved.
Multi-stage QA follows, using tools like Google Vertex AI for automated content analysis. What arrives on the other end: structured datasets with direct download links, consent releases, PII audits, manifests, transcripts, quality scores. The whole package.
Pricing models vary—flat fees, per-minute rates, revenue-sharing arrangements—though Luel has not disclosed specific numbers. Contributors get paid through Venmo, PayPal, Stripe, or Wise, typically within 2–7 business days. Minimum payout threshold: $5.
On paper, it's elegantly simple. In practice, the legal architecture reveals just how thorny data provenance has become.
The Fine Print

Contributors grant Luel a perpetual, worldwide, transferable license to sell their content. The kicker: it's exclusive and sublicensable, covering everything from biometric data to voice recordings. Buyers, meanwhile, receive rights to use the data for model training and derivatives, but they cannot redistribute the raw datasets. Sales are final. Irrevocable.
The privacy policy—last updated February 16—explicitly describes the handling of biometric signals, deduplication hashes, and data flows. SOC and AICPA compliance badges appear on the site, though no third-party attestation links were visible during a review of the pages.
To demonstrate what's possible, Luel published a sample dataset on Hugging Face called Ego-Realm: 1,082,530 frames across 108 clips, just over 10 hours of 1080p footage captured with a Pineye camera system. The company describes this as roughly 1% of a larger enterprise collection, built by more than 400 contributors.
The catalog now lists dozens of speech datasets and multiple sensor and video collections. Y Combinator highlighted examples like patient-doctor conversations recorded in South Asia and gemstone footage intended for robotics applications—use cases specific enough to suggest early customer conversations, if not signed contracts.
A Landscape in Flux

Luel enters a data infrastructure market that has been quietly reordering itself.
Scale AI, long the gravitational center of data labeling and collection, underwent significant restructuring after Meta acquired a 49% stake for roughly $14–15 billion last June. Founder Alexandr Wang moved to Meta. Scale laid off 14% of its workforce—around 200 employees—the following month.
Meanwhile, established licensing players have been building AI-specific pipelines. LiveRamp expanded its data marketplace into an AI training hub in January. The Copyright Clearance Center announced its AI Systems Training License last March and recently launched internal AI reuse rights for U.S. higher education. Shutterstock, an early mover, expanded its OpenAI partnership back in 2023 and introduced broader AI services this past October.
What's different about Luel, perhaps, is the bet that consent-first infrastructure can move fast—"within days" rather than the weeks or months traditional vendors require, according to the company. That's a bold claim for a two-person team that recently posted its first software engineer role on LinkedIn.
A third-party tracker lists Luel at $500,000 raised, consistent with Y Combinator's standard deal, though the company has not confirmed any funding figures beyond its accelerator participation. The YC launch post, published about a month ago, picked up 761 upvotes—a respectable showing, though hardly viral.
Namgyal, who studied in Berkeley's M.E.T. program before dropping out, and COO Lenderking position the service as a response to what they see as a fundamentally broken supply chain. There's an appealing clarity to the argument: if the foundation is compromised, everything built on top eventually cracks.
The Unanswered Questions

Can a contributor network—no matter how vetted—actually scale to meet the appetites of frontier models? GPT-4 reportedly trained on trillions of tokens. Luel's Ego-Realm sample, while meticulously documented, represents 10 hours of video. The gulf between demonstration and deployment is wide.
And there's the matter of price. Consent costs money. Compliance costs money. AI companies have grown accustomed to essentially free data, scraped from the open web with the legal questions deferred to some hypothetical future. Convincing them to pay a premium—even to avoid litigation—requires either regulatory pressure or a high-profile legal loss that makes the risk calculation obvious.
Both may be coming. But neither has fully arrived.
For now, Luel has identified a clear opening: AI builders who would rather pay upfront than defend a subpoena later. Whether that's a niche or a new normal will depend on how the next few copyright cases play out—and whether the industry's tolerance for legal ambiguity finally runs dry.
