The pitch is simple, almost bracingly so: what if you could train your AI model on data you actually owned?
BeatpulseLabs, a relative newcomer to the AI infrastructure space, raised $1.8 million in a pre-seed round on June 8, 2026. The round—co-led by UK-based Arāya Ventures and Czech firm Lighthouse Ventures, with backing from Alumni Ventures and Mexico's Avalancha Ventures—comes at a moment when enterprise AI builders are growing skittish about the legal exposure that comes with web-scraped training sets.
It's not hard to see why. The New York Times, Getty Images, and a parade of authors and artists have sued major AI labs over copyright infringement. Meanwhile, lawmakers in Brussels and Washington are circling the question of fair use in machine learning. In this climate, the promise of "rights-cleared" datasets carries weight.
Whether it carries $1.8 million worth of weight remains to be seen.
The Inventory
BeatpulseLabs sells training data across four verticals: speech, video, music, and radar. The company claims its offerings are not just licensed but entirely absent from the open web—a selling point aimed squarely at enterprises wary of regulatory blowback or shareholder scrutiny.
The speech catalog runs to over 1 million hours, spanning 75 languages and produced, the company says, by 250 in-house linguists and voice actors. Video holdings—more than 1 million hours of broadcast-grade footage—come from a network of 700 professionals and what BeatpulseLabs describes as exclusive partnerships with global broadcasters, though specific partners are not named.
On the music side: 850,000 owned assets, created by 3,000 classically trained contributors. The radar dataset, perhaps the most niche of the bunch, comprises 10 terabytes of live operations data annotated by what the company claims are 50 NATO-cleared specialists. That last offering targets defense and aerospace clients—a signal that BeatpulseLabs is thinking beyond consumer-facing AI applications.
The emphasis throughout is on provenance. In an industry where foundation models have been trained on everything from copyrighted novels to unlicensed YouTube videos, owning the underlying content outright is less a luxury than a risk management strategy.
Who's Behind It

Co-founders Jason Rieff and Nikolay Vitanov are steering the operation from a Delaware incorporation with a London mailing address. Rieff, a South African entrepreneur, lists himself on an undated Web Summit profile as a three-time founder with one exit. Vitanov, Bulgarian by origin, reportedly spent six years as a director at Citi before co-founding BeatpulseLabs in mid-2024, according to an unverified profile on The Org.
The team is hiring across music annotation, speech and linguistics, video processing, and radar intelligence, with roles listed on the company careers page. A recent LinkedIn post seeking an APAC Head of Sales hints at geographic ambitions beyond the UK and European base.
According to Tech.eu, the company posted 10x revenue growth in H1 2026, though absolute figures remain under wraps. Whether that growth stems from a handful of large contracts or a broadening customer base is unclear.
The Bigger Picture

BeatpulseLabs is positioning itself as a counterpoint to the "move fast and scrape everything" ethos that has defined much of the AI training data economy. But it's entering a market that's already getting crowded—and sophisticated.
Origin Lab pulled in $8 million earlier this year to broker video game data for world model builders. Suno, the AI music generator, raised over $400 million at a $5.4 billion valuation in early summer. The common thread: as models grow more specialized, the quality and legal standing of training data matter as much as—perhaps more than—sheer volume.
BeatpulseLabs says it plans to use the $1.8 million for platform expansion and customer acquisition. The company also stresses that it scales through "subject matter expert performance and dataset fidelity" rather than headcount—a pointed rejection of the call-center model that dominates much of the data labeling industry.
Whether that approach can generate the margins investors expect from a venture-backed startup is another question entirely. Curating high-quality, rights-cleared datasets is expensive. So is employing 250 linguists and 3,000 musicians. The bet here is that enterprises will pay a premium to sleep soundly at night, legally speaking.
If they do, BeatpulseLabs may have found a durable niche. If not, it will become a footnote in the ongoing scramble to feed the AI hunger for data—cleanly, legally, and at scale.
