Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Fintech iconFintechJune 4, 2026

Drafted Raises $16M to Democratize Home Design with Generative AI

Drafted Raises $16M to Democratize Home Design with Generative AI
YcGenerative Ai+3
SaaS iconSaaSJune 4, 2026

French Quantum Startup Quobly Raises €115M Series A Led by Bpifrance

French Quantum Startup Quobly Raises €115M Series A Led by Bpifrance
Quantum ComputingDeeptech+3

Founders Mentioned

Cassidy Dalva

Miso Labs

saas icon
SaaS

Aoden Teo

Miso Labs

saas icon
SaaS

Cassidy Dalva

Miso Labs

saas icon
SaaS

Aoden Teo

Miso Labs

saas icon
SaaS
SaaS iconSaaS
June 5, 2026
YcVoice AiOpen SourceB2b SaasEnterprise Ai

Miso Labs Launches Open-Source Voice AI Faster Than Humans

YC S26 startup releases 8B-parameter MisoTTS with 110ms latency—faster than human reaction time. Open-source weights now available; commercial API waitlisted as it targets enterprise.

Miso Labs Launches Open-Source Voice AI Faster Than Humans

Cassidy Dalva and Aoden Teo want you to believe their voice AI model can respond faster than you can blink. Actually, faster than you can process what someone just said to you and formulate a reply—at least according to their internal tests.

The claim: 110 milliseconds. That's the latency Miso Labs—the two-person startup the pair founded out of Y Combinator's Spring 2026 cohort—says its newly released voice model achieves, though this figure comes from the company itself and hasn't been independently verified. For context, cognitive scientists peg the typical human reaction time to speech somewhere around 160 milliseconds. If Miso's figures hold up beyond controlled demos, that would put conversational AI into territory that feels genuinely instantaneous.

The model, called MisoTTS (or "Miso One," depending on which piece of marketing collateral you're reading), went live on Hugging Face earlier this month. All 8 billion parameters of it. The weights are open, the downloads are free, and developers willing to wrangle 33 gigabytes of model files can start experimenting immediately.

The commercial API? Still waitlisted. Miso Labs says it's "securing the compute to serve our API at scale." Translation: the infrastructure isn't ready yet, but the code is out there for anyone to test.

Whether 110 milliseconds is marketing spin or legitimate breakthrough remains an open question. The company hasn't released methodology details or submitted to independent benchmarking. What they have released is a model that's generating early interest—1,600 GitHub stars and 128 forks as of early June—and a bold on-premises pitch aimed squarely at enterprises wary of sending voice data to third-party clouds.

An 8-Billion-Parameter Bet on Speed

Under the hood, MisoTTS is a text-to-dialogue RVQ Transformer. Miso Labs describes the architecture as "Sesame-style CSM" built on a Llama 3.2-style backbone, with roughly 300 million additional parameters dedicated to the audio decoder. It uses the Mimi audio codec, operates with 32 codebooks, and handles sequences up to 2,048 tokens. English only, at least for now.

The full model file weighs in at 32.8 GB, shipped in bfloat16 precision by default. That's a hefty download, even by the inflated standards of modern AI. But for developers accustomed to pulling multi-billion-parameter models, it's workable.

One feature stands out: one-shot voice cloning from a 10-second audio sample. Feed the model a short clip, and it can generate speech in that voice. That capability puts MisoTTS in direct competition with ElevenLabs and similar commercial offerings, though with a crucial difference—you can run it locally, no API fees required.

There's a wrinkle, though. Miso Labs includes watermarking via Sony's SilentCipher as an available option. The documentation urges users to configure a private watermark key in production deployments, a detail that matters as regulators zero in on voice cloning risks. U.S. senators sent inquiries to voice-synthesis firms in April, pressing them on deepfake safeguards and scam prevention. Miso seems aware of the scrutiny.

Open Weights, With a Catch

The license is a modified MIT, which sounds permissive until you read the fine print. Products or services crossing 50 million monthly active users or $10 million in monthly revenue must display "Miso Labs" prominently in their user interface. The copyright belongs to Kamino Learning, Inc., the Delaware entity incorporated last August that operates as Miso Labs.

So it's open-weights, yes. Purist open-source? Not quite. Small-scale deployments get a free pass. But hit certain scale thresholds and brand attribution becomes mandatory. That licensing structure will matter for anyone evaluating MisoTTS for production use at companies already operating at scale.

The On-Prem Angle

Digital illustration for article section "The On-Prem Angle" in "Miso Labs Launches Open-Source Voice AI Faster Than Humans" - A sleek, minimalist glass and brushed steel vault resting on a polished pedestal, securely enclosing...

Miso Labs is pitching hard on local deployment. The company's website emphasizes self-hosted inference and offers enterprise support contracts for organizations that want voice processing kept in-house. That's a direct play for healthcare systems, financial institutions, and government agencies—sectors where data residency rules and privacy concerns outweigh the convenience of managed cloud APIs.

The comparison chart on Miso's site lists their model at 110 milliseconds, human reaction time at 160, Sesame at 300, ElevenLabs at 700. These are vendor-supplied figures that should be regarded as estimated values without independent validation. Real-world latency depends on hardware, pipeline design, network conditions—variables the chart doesn't address. It's marketing, not peer review.

Still, Dalva and Teo seem serious about building a team around the technology. Their careers page lists openings for speech synthesis researchers, low-latency GPU kernel engineers, and a founding go-to-market lead focused on developer partnerships. The hiring plan from a two-person team suggests they aim to expand rapidly across R&D and commercial functions, assuming the model gains traction.

The API Problem

Here's the gap: you can download the weights today, but you can't use the managed API unless Miso Labs approves your access request. The waitlist page frames it as a compute constraint—they're still lining up servers to handle API traffic at scale.

It's a familiar playbook for small teams shipping large models. Release open weights first to build developer mindshare and gather feedback. Launch the paid API later, once infrastructure is solid and you've worked out the kinks. In the meantime, anyone comfortable with Python and a GPU can clone the repo and start generating audio. You just can't rely on Miso to host it for you.

A Market Getting Crowded

Digital illustration for article section "A Market Getting Crowded" in "Miso Labs Launches Open-Source Voice AI Faster Than Humans" - A modern, conceptual representation of a crowded audio AI market featuring sleek, abstract 3D soundw...

Miso isn't entering empty space. Mistral released Voxtral TTS in March, positioning it as an open alternative to ElevenLabs on naturalness and expressiveness. Kyutai's Moshi—released back in September 2024—demonstrated real-time full-duplex dialogue with latencies hovering around 200 milliseconds, a benchmark that felt impressive at the time. ElevenLabs remains the commercial standard, though latency claims for its API vary wildly depending on who's measuring and under what conditions.

Miso's wager seems to be that combining 8 billion parameters with sub-200-millisecond latency and local deployment flexibility will carve out a defensible niche. The pitch centers on "the most emotive foundation models for voice," a phrase that appears in marketing materials but lacks empirical backing. Emotive quality—prosody, tone, naturalness—is notoriously subjective and difficult to benchmark.

No publicly named enterprise customers or partnerships have been announced as of early June. Miso Labs' funding details outside Y Combinator's involvement remain undisclosed. Miso Labs is early-stage, two people, and just put weights into the world. Whether that's enough to compete against well-funded incumbents or gain traction in a market already saturated with voice models is an open question.

Worth a Download?

Digital illustration for article section "Worth a Download?" in "Miso Labs Launches Open-Source Voice AI Faster Than Humans" - A sleek, minimalist three-dimensional acoustic waveform shaped into a perfectly smooth, polished sph...

For developers evaluating voice models, MisoTTS offers something useful: the ability to test latency and quality claims locally before committing to any API contract. No vendor lock-in, no usage fees at small scale, full control over the inference pipeline. That's appealing if you're building conversational AI and want to understand what's possible outside managed services.

For enterprises, the on-prem angle is intriguing, particularly in regulated industries. But you'll want to run your own benchmarks, stress-test the latency claims, and parse the licensing terms carefully before deploying at scale.

Either way, the weights are out there now. Thirty-three gigabytes and a 10-second voice sample away from finding out whether 110 milliseconds is engineering reality or optimistic marketing. Dalva and Teo have made their bet. The next few months will reveal whether the market agrees.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Drafted Raises $16M to Democratize Home Design with Generative AI
  • French Quantum Startup Quobly Raises €115M Series A Led by Bpifrance
  • YC-Backed Juno AI Reaches 90K Patients, $70K MRR Since Launch
  • Fuse Energy Raises $30M Extension at $5B Valuation to Scale UK Model
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.