Lionel Messi's venture capital firm has placed a bet on voice synthesis technology—not exactly the most obvious pivot from the pitch. But Play Time, the investment vehicle backed by the Argentine forward, participated in a seed round for Fish Audio dated May 1, 2026, a scrappy voice AI startup that's built a following around open-source speech models and aggressive pricing.
The deal closed earlier this year, according to filings tracked by CB Insights and PitchBook, though neither party has publicly disclosed the amount or issued the customary press release. What we do know: Fish Audio claims it hit $13 million in annual recurring revenue and 7 million users by spring, up from $10 million ARR and 5 million users just a quarter earlier—self-reported figures that haven't been independently verified. That's the kind of growth trajectory that gets attention in a market increasingly skeptical of AI hype but hungry for alternatives to the category's dominant player.
A Play for the Infrastructure Layer
Play Time invests where sports, media, and technology collide—think athlete-driven content platforms or fan engagement tools. Voice AI might seem tangential, but the firm appears to be betting on the underlying infrastructure that could power everything from personalized sports commentary to automated highlight reels. It's early-stage capital, the kind meant to fuel product iteration and market validation rather than fund splashy ad campaigns.
Fish Audio operates under the parent entity Hanabi AI Inc., with Shanghai Qita Dynamic Technology Co., Ltd. listed in recent terms of service updates as the operational backbone. CEO and co-founder Rissa Cao shared the revenue and user figures in a LinkedIn post several months back—a casual disclosure that's become standard practice among founders courting developer communities rather than traditional press cycles.
For context: ElevenLabs, the 800-pound gorilla in this space, raised $500 million in February 2026 at an $11 billion valuation. Fish Audio is playing a different game entirely, one that hinges on being faster and cheaper than incumbents while maintaining acceptable quality. Whether that's enough to carve out durable market share remains an open question.
Speed, Price, and the Developer Pitch
The company's technical thesis rests on its S2 model, detailed in an arXiv paper published in March. Co-authored by chief scientist Shijia Liao—formerly a researcher at NVIDIA—the paper claims a real-time factor of 0.195 and time-to-first-audio latency under 100 milliseconds. In plain English: the model processes and generates speech faster than you can blink, with streaming inference capabilities across more than 80 languages.
Third-party developer blogs and scattered social media threads have amplified unverified claims that Fish Audio runs "2x faster than Cartesia" and costs roughly a sixth of what ElevenLabs charges. The pricing structure is credit-based, which makes apples-to-apples comparisons tricky, but external trackers peg Fish Audio's API costs around $15 per million characters. ElevenLabs, depending on tier and model, ranges from $60 to $165 per million characters.
None of this has been independently verified through peer-reviewed benchmarks, and the startup world is littered with performance claims that don't hold up under scrutiny. Still, the developer chatter suggests Fish Audio has built something real enough to gain traction among engineers hunting for cost-effective alternatives.
The competitive landscape is crowded and well-funded. Cartesia, which raised a $27 million seed led by Index Ventures, is chasing similar real-time synthesis goals. Paris-based Gradium secured a $100 million seed from Nvidia in July 2026—a war chest that dwarfs Fish Audio's undisclosed round. These are not weekend hackathon projects; they're serious infrastructure plays backed by serious institutional money.
Open Source as Strategy—and Liability

Fish Audio's open-source roots distinguish it from most rivals. The company's fish-speech repository on GitHub has collected more than 31,000 stars, with model weights freely available on Hugging Face. A July integration with audio.cpp, a C++ inference library, brought native support for Fish Audio's S2 Pro model—the kind of low-level tooling that appeals to developers building voice features into larger applications.
But the open platform has become a double-edged sword. Voice artists and talent agents have taken to LinkedIn to complain about unauthorized voice cloning on Fish Audio's platform, which the company reports hosts upwards of 2 million user-generated voice models. It's a moderation nightmare that every generative AI platform eventually confronts: how do you police a user base that's creating content faster than you can review it?
Cao has stated in LinkedIn posts her commitment to takedowns and stricter enforcement for content that violates intellectual property rights. The company's June terms of service include the standard safety and liability clauses you'd expect from a generative AI service, though the sheer volume of user-uploaded models presents an ongoing headache. Perhaps more than the founders anticipated.
One veteran voice-over artist, speaking on condition of anonymity, put it bluntly: "These platforms act like they're just neutral infrastructure, but they're enabling IP theft at scale." Fish Audio isn't unique in facing this criticism—ElevenLabs, Respeecher, and others have navigated similar controversies—but for a smaller startup courting enterprise customers, perception matters.
Shipping Fast, Scaling Faster

Fish Audio has kept up a brisk product cadence, pushing weekly updates through its Reddit community. Recent releases include S2.1 Pro, a free API trial window in late June, and integrations with Telnyx and Story Studio. A mobile app appeared on Google Play a few weeks back. In a LinkedIn post from several months ago, the company announced emotion-aware speech-to-text, claiming support for more than 100 languages with paralanguage tagging—industry jargon for capturing sighs, laughter, and other non-verbal cues. These claims have not been independently verified.
The reported jump from $10 million to $13 million ARR in a single quarter, with roughly half the revenue mix coming from enterprise customers, suggests the company has found product-market fit with a specific cohort: developers and businesses willing to trade brand recognition for performance and cost savings. That's a real segment, though not an infinite one.
What Comes Next

Whether Fish Audio can sustain this momentum depends on variables that have little to do with model latency or pricing spreadsheets. IP enforcement will remain a persistent challenge as the platform scales. Competitive pressure from billion-dollar rivals with deeper pockets and more established sales organizations will intensify. And there's a broader, unresolved question hanging over the space: can open-source strategies actually win in real-time voice infrastructure, or will proprietary, tightly controlled systems ultimately dominate?
Messi's investment firm is betting on the former. So are the developers who've starred the GitHub repo and integrated the API into production systems. But in venture capital, as in soccer, early leads don't guarantee final outcomes. The game, as they say, is still very much in play.
