Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Healthtech & Biotech iconHealthtech & BiotechMarch 26, 2026

Inside the Race to Build Digital Twins of Human Bodies

Inside the Race to Build Digital Twins of Human Bodies
Digital TwinsPrecision Medicine+3
Healthtech & Biotech iconHealthtech & BiotechMarch 26, 2026

Qualified Health Pursues $100M Raise After Major Health System Wins

Qualified Health Pursues $100M Raise After Major Health System Wins
HealthtechAi+3

Founders Mentioned

Amit Yadav

Ishiki Labs

saas icon
SaaS

Robert Xu

Ishiki Labs

saas icon
SaaS

Amit Yadav

Ishiki Labs

saas icon
SaaS

Robert Xu

Ishiki Labs

saas icon
SaaS
SaaS iconSaaS
March 26, 2026
YcVoice AiConversational AiSales IntelligenceB2b Saas

Teaching AI When to Shut Up: Ishiki Labs' Breakthrough in Voice Timing

While competitors race to make AI talk more, this YC-backed startup solved the opposite problem: knowing when to stay silent. Inside the multimodal breakthrough reshaping sales calls.

Teaching AI When to Shut Up: Ishiki Labs' Breakthrough in Voice Timing

Watch any voice AI demo and you'll see the same choreography unfold: the bot waits, perfectly patient, for you to finish speaking. Then it processes. Then it responds. Polite, certainly. Predictable, absolutely.

And utterly inhuman.

The problem isn't what these systems say—most can hold their own on facts and context by now. It's their complete inability to grasp something every toddler learns before kindergarten: sometimes the smartest move is to keep your mouth shut.

That timing gap—the awkward, half-second dance of who speaks when—has emerged as the defining constraint of conversational AI, even as the technology barrels toward ubiquity. Gartner projected worldwide AI spending would hit $2.52 trillion in 2026, a 44% jump from 2025, with billions flowing into voice and agentic systems. Yet for all that capital and compute, most voice assistants still feel robotic. They can't read the room. Or more precisely, they can't read the millisecond pauses, the overlapping speech, the social cues that govern how humans actually talk to each other.

Which brings us to Ishiki Labs.

A Winter 2026 Y Combinator company, Ishiki is taking what might seem like a counterintuitive approach: while competitors race to make AI faster and more capable, founders Amit Yadav and Robert Xu are building AI that knows when not to talk. Their product, Fern, is a real-time sales assistant designed around what they call "socially conscious AI"—systems that understand turn-taking, context, and the subtle art of strategic silence.

It sounds almost trivial until you consider how few voice systems actually pull it off.

Why Timing Is Everything (and Everything Is Hard)

The technical challenge here is deceptively nasty. Current voice AI operates in what researchers call "half-duplex" mode: the system waits for a pause, interprets that pause as permission, then generates a response. Fine for "What's the weather?" Catastrophic for complex, multi-party conversations where pauses carry meaning and interruptions signal engagement rather than rudeness.

Research published in March frames the problem starkly. One paper, "Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue," argues that voice agents must decide at every conversational beat whether to speak or hold back, using the full context—not just the last thing someone said. Another March study, RESPOND, critiques pause-and-respond systems as "stiff and robotic," calling instead for predictive orchestration that enables natural interjection.

In production environments—sales calls, customer service, high-stakes meetings—the stakes become visceral. Sales reps report that delays above 500 milliseconds feel noticeable. Anything over a second breaks the spell entirely.

Yet typical voice pipelines chain together speech-to-text, large language model inference, and text-to-speech, each adding latency. Worse, these systems struggle with barge-in scenarios (when the user interrupts mid-response) and often lack the echo cancellation needed for full-duplex audio. Practitioner communities on Reddit and developer forums consistently cite interruption handling as the single biggest predictor of perceived quality—sometimes mattering more than which model powers the thing.

Meta's Multimodal DNA

Yadav and Xu bring unusual pedigree to this particular problem. Yadav spent years at Meta on the LLaMA team and in Reality Labs, where he worked on multimodal and video assistants for smart glasses—technology designed for continuous, socially embedded interaction rather than request-response exchanges. He holds a PhD. He's published more than 20 papers.

Xu came from Reality Labs as well, focused on systems and infrastructure, with a prior stint at Citadel Securities.

Their academic work telegraphs the thesis. In their paper "Beyond Words: Multimodal LLM Knows When to Speak," they describe MM-When2Speak, a model that predicts when and how to respond using audio, vision, and text signals. The authors report up to 4× improvement in response-timing accuracy compared to baseline LLMs. Another paper, M-CALLM, achieved sub-35 millisecond latency for conversational predictions with 96% accuracy—though those numbers come from academic settings with limited sample sizes, the usual caveats about lab-to-production translation applying.

In a LinkedIn post two months ago, Yadav framed the mission bluntly: today's multimodal assistants aren't human-like because they're "optimized for 1-to-1 request-response, don't understand timing, don't know when to stay silent, don't act proactively." He shared demo videos showing "model knows when to stay silent" in action.

The Y Combinator profile makes the commercial strategy explicit: "We originally built socially conscious AI—an AI that knows when to talk and when to stay silent—then applied it to sales." Sales, they reasoned, is the most timing-critical conversational domain. Miss the moment to surface a key data point and the deal stalls. Interrupt the prospect mid-thought and trust evaporates.

It's a narrow wedge with broad implications.

Two Modes, One Philosophy

Digital illustration for article section "Two Modes, One Philosophy" in "Teaching AI When to Shut Up: Ishiki Labs' Breakthrough in Voice Timing" - A minimalist and conceptual 3D scene representing a silent copilot in "Shadow Mode," featuring a fac...

Fern operates in two distinct configurations.

Shadow Mode is invisible: it joins sales calls but doesn't speak, instead providing a private, real-time feed of live summaries, suggested questions, and context about who's on the line. Think of it as a silent copilot, streaming intelligence that only the rep can see. No one else on the call knows it's there.

Full Presence Mode lets the AI participate actively—speaking, answering questions, pulling company knowledge. Even here, the company's positioning emphasizes that the system is designed to interject appropriately rather than dominate. The product site promises features like role-playing tough calls and turning "every call into team intelligence" by integrating with company data.

Ishiki recently opened the technology to developers through an API published at api.ishikilabs.ai. The platform supports real-time WebRTC streaming with sub-second latency and configurable agent identity and capabilities. Pricing tiers range from Free to Max, with included minutes and overage rates—suggesting a dual strategy: selling Fern directly to commercial teams while also courting developers building their own real-time voice applications.

The Real-Time Land Grab

Ishiki isn't operating in a vacuum. The incumbent players in sales and meeting intelligence are all pivoting from post-call analytics toward real-time guidance, and fast.

Fireflies.ai, long known for meeting transcription and summaries, launched Live Assist in October 2025—a feature that surfaces suggestions and answers during meetings across Zoom, Teams, Webex, and Google Meet. Cresta, entrenched in contact-center agent assist, positions around real-time prompts and coaching during live conversations, emphasizing agentic workflows in production settings. Even Gong, historically focused on post-call conversation intelligence and pipeline forecasting, announced expanded real-time elements in February as part of its "Mission Andromeda" push into enablement.

The shift is enabled by infrastructure improvements from the hyperscalers. OpenAI's Realtime API, updated for production in late 2025, provides WebRTC and WebSocket support for voice, audio, and vision with significantly reduced latency. Azure OpenAI offers multiple realtime model SKUs. Anthropic began rolling out voice modes for Claude in early March, though practitioners often note OpenAI maintains an edge in latency and backchannel responsiveness.

Perhaps the broader trend is captured in Gartner's forecast from last August: by the end of this year, 40% of enterprise applications will feature task-specific AI agents, up from less than 5% in 2025. Voice is becoming one modality among many for these agents, but it remains the most timing-sensitive—and the hardest to get right.

Regulatory Headwinds

Digital illustration for article section "Regulatory Headwinds" in "Teaching AI When to Shut Up: Ishiki Labs' Breakthrough in Voice Timing" - A sleek, minimalist 3D pop art designer toy conceptualizing the tightening regulation of voice AI, f...

The regulatory environment is tightening around voice AI, particularly in ways that affect emotion detection and automated calling.

The EU AI Act's transparency rules take effect in August, with prohibitions on certain biometric categorization and emotion recognition already active since February 2025. Deployers must inform users when they're interacting with AI and disclose any use of emotion inference. That matters for products claiming to read conversational affect or stress—a capability that sounds useful until you confront the compliance overhead.

In the U.S., the FCC ruled in February 2024 that AI-generated voice calls fall under the Telephone Consumer Protection Act, making robocalls with cloned or synthetic voices illegal without explicit consent. Enforcement has been aggressive. High-profile cases involving fake political robocalls have made clear the agency isn't playing around. For companies building outbound voice agents, that regulatory line is sharp and consequential.

Inbound assist tools like Fern sidestep much of that risk, but meeting assistants still navigate all-party consent requirements—California's CIPA, for instance, requires consent from all participants to record calls, not just the person who invited the bot.

Despite those constraints, the momentum is unmistakable. SoundHound launched generative AI voice assistants in select Jeep vehicles in Europe last August, bringing natural language conversation to automotive. Hume AI, which raised $50 million in March 2024, is building empathic voice interfaces that adapt responses based on emotional cues, with production case studies in HR and recruiting. PolyAI raised $86 million in December 2025, according to Forbes, as enterprise voice assistant deployment accelerates.

The academic frontier is racing ahead too. Recent research on continuous interjection and full-duplex dialogue management points toward systems that predict not just what to say but when to jump in—moving beyond pause-detection heuristics to genuine conversational orchestration. Another paper describes real-time detection of customer questions during sales calls, with automatic retrieval and dashboard presentation of answers—essentially the use case Fern is commercializing.

The market opportunity is large enough to attract this level of attention. The call-center AI segment alone is projected to grow from around $3.25 billion in 2025 to over $4 billion this year, according to The Business Research Company. A 2022 Gartner forecast predicted conversational AI would reduce contact-center labor costs by $80 billion in 2026—a figure that remains directionally plausible even if the exact mechanics prove messier than anticipated.

The Restraint Bet

Digital illustration for article section "The Restraint Bet" in "Teaching AI When to Shut Up: Ishiki Labs' Breakthrough in Voice Timing" - A minimalist, conceptual representation of conversational restraint featuring a single, oversized, g...

What Ishiki Labs is wagering on is that the next wave isn't just about making AI more capable or more knowledgeable. It's about making AI more conversational—which paradoxically requires teaching it restraint.

A system that knows when to shut up, when to wait for the right moment, when to let a pause breathe? That's more valuable in high-stakes conversations than one that answers every silence with reflexive verbosity.

That insight, applied rigorously, could reshape how teams collaborate with AI in real time. Or—and this is the riskier possibility—it could turn out that most buyers still prefer the simpler, faster, louder option. The one that feels like a tool rather than a participant.

Either way, the company's explicit focus on social timing represents a bet that the awkwardness of current voice AI isn't a minor UX problem. It's the core constraint holding the technology back from truly ambient integration.

And if that's right, the race isn't to make AI talk more.

It's to make AI time better.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Inside the Race to Build Digital Twins of Human Bodies
  • Qualified Health Pursues $100M Raise After Major Health System Wins
  • Octen-8B Tops New RTEB Benchmark, Disrupting AI Search Standards
  • Synthetic Biology's Commercial Moment: Products Hit Shelves in 2026
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.