The Meta Ray-Bans perched on your nose can identify Taylor Swift from across a crowded room. They'll translate your Italian waiter's specials into English, mid-sentence. They even know that song drifting out of the corner coffee shop is "Cruel Summer."
What they can't reliably figure out? Whether you're actually speaking to them.
It sounds almost absurdly basic. And yet this singular problem—what researchers call "addressee detection"—has become the most expensive unsolved riddle in the multimodal AI gold rush. Voice agents that pipe up mid-argument with your spouse. Meeting bots that stare blankly through direct questions. Wearables that sleep through your third frantic "Hey Meta" or jolt awake because someone on TV said "a city." As these systems migrate from the relative safety of smartphones into our living rooms, conference rooms, and eventually, strapped to our faces, their inability to read social cues isn't just annoying. It's an existential threat to the category.
Consider the data point that should terrify every smart speaker manufacturer: a 2020 study from Northeastern University and Consumer Reports clocked roughly one misactivation per hour during ordinary TV watching. Ten percent of those phantom wake-ups lasted ten seconds or longer—long enough for the device to capture snippets of your private conversation. Under real-world conditions, some units averaged a false trigger every five hours. Each one chips away at trust. Enough of them, and the device winds up in a drawer.
Now a startup barely out of stealth mode thinks it's cracked the code. Ishiki Labs, part of Y Combinator's Winter 2026 batch, is building what amounts to social awareness for AI: models that know not just what to say, but crucially, when to shut up.
The Problem Is Bigger Than It Looks
The multimodal AI market, if you believe the projections, is on a tear. Grand View Research pegs growth from $1.74 billion in 2024 to nearly $11 billion by 2030—a 37 percent compound annual growth rate that assumes, perhaps optimistically, that the industry solves its fundamental interaction problems. Right now? Those problems remain stubbornly unsolved.
OpenAI's GPT-4o brought genuine real-time multimodal chops, with voice latencies between 232 and 320 milliseconds. Fast enough to feel conversational. Google's Project Astra and Gemini Live are trickling out to premium subscribers, native audio processing intact. Apple Intelligence is pushing on-device multimodal inference to millions of devices. Amazon, meanwhile, has reportedly pushed its LLM-powered Alexa overhaul into 2025 after early testers found it stiff, slow, and prone to breaking existing smart-home routines—problems severe enough that the company quietly shelved its "Let's Chat" preview.
The through line in all this: these systems respond quickly. They just don't know whether they should respond at all.
Meeting AI from Google and Zoom handles post-hoc summarization competently enough. But ask them to participate naturally in a live conversation—to know when "What do you think?" is directed at them versus the person sitting across the table—and you'll get blank stares or, worse, confident interjections at precisely the wrong moment. Meta's Ray-Ban smart glasses picked up live AI features in 2024, including continuous translation and visual search. User forums, though, overflow with complaints about "Hey Meta" unreliability. The assistant sleeps through legitimate requests, then springs to life when you're just talking to a friend.
The cautionary tales keep piling up. Humane's AI Pin shut down in February 2025 after withering reviews and a quiet acquisition by HP, leaving early adopters scrambling for refunds. Rabbit's R1 stumbled over latency and usability complaints. The lesson, increasingly unavoidable: always-on assistants need to master the basics before attempting parlor tricks.
Why This Is So Hard to Solve
Addressee detection sits at a gnarly intersection of multiple hard problems, each difficult on its own.
You need acoustic analysis sophisticated enough to determine whether someone's facing you, speaking in your direction, or just venting near you. Visual cues—gaze direction, facial orientation, the subtle tilt of a head that signals engagement. Prosodic features that distinguish a genuine question from a rhetorical one. Conversational state tracking who said what, when, and to whom. String those together in real time, at sub-second latencies, on battery-constrained devices, and you start to see why most companies treat this as a "later" problem.
Academic research demonstrates measurable, if incremental, progress. Apple researchers presented work at the 2024 NeurIPS AFM workshop showing that feeding LLMs both ASR uncertainty and conversational context for follow-up queries cut false alarms by 20 to 40 percent, at a fixed ten percent false-reject rate, compared to modeling follow-ups in isolation. Google's 2022 research on intended query detection for continued conversation claimed a 22 percent equal error rate improvement and shaved 600 milliseconds off latency using an end-to-end streaming model versus an independent detector.
Amazon's Conversation Mode for Echo Show devices—what they call device-directed speech detection, or DVAD—combined visual and acoustic cues in a multimodal dance. The approach slashed false rejects by 83 percent compared to vision-only methods, while also cutting phantom wake-ups. Google's "Look and Talk" on Nest Hub Max uses gaze tracking and Face Match to start listening without a wake word, all processed locally on the device.
But integrating these techniques into production systems at scale, with all the messy edge cases real users throw at you, remains brutally hard. False triggers have consequences beyond user annoyance. A 2020 catalog documented over 1,000 false wake phrases across major assistants—mundane utterances like "a city" triggering "Hey Siri," or "and sent a" waking Alexa. Some were darkly comic. Most were just frustrating.
The infrastructure layer is evolving to help. Deepgram's Flux STT includes integrated end-of-turn detection, configurable on the fly, achieving roughly 260 millisecond detection of when someone's actually done speaking. OpenAI's Realtime API bakes in voice activity detection and interruption handling. Still, these are primitives. Developers need to wire them together intelligently.
And then there's regulation, which isn't waiting for the technology to mature. The EU AI Act's timeline put prohibited practices into effect this February, with most high-risk obligations hitting by August 2026. Illinois's BIPA—the biometric privacy law with actual teeth—covers voiceprints and has spawned litigation across multiple industries. The FCC floated AI disclosure rules for robocalls last August. Companies building always-listening assistants need airtight consent flows and what one attorney I spoke with called "conservative abstention strategies"—fancy language for: when in doubt, stay quiet.
Enter Ishiki Labs

Amit Yadav and Robert Xu don't look like your typical AI startup founders, mostly because they've spent years inside the machinery they're now trying to fix.
Yadav logged time at Meta AI working on LLaMA speech and multimodal models, plus a video assistant for smart glasses within Reality Labs—the division tasked with building Meta's AR future. He holds a PhD from Purdue and has published north of 20 papers at venues like CVPR, NeurIPS, and ICASSP. Real credentials, not just Medium posts. Xu worked on orchestration for Meta's Orion AR glasses and built research infrastructure at Citadel Securities, where milliseconds and reliability aren't nice-to-haves.
Their product, called Fern, targets what they frame as AI that "knows when to stay silent and when to respond" to real-time video and audio streams. The company's research page showcases two models: fern-0.1-base for basic addressee detection, and fern-0.1-multistep for procedural guidance—think step-by-step instructions that know when you're ready for the next step. Yadav's public posts hammer home the core insight: "models don't understand timing; don't know when to stay silent." He positions Fern with "latency on par with Gemini," which, if true, would put them in the 200-400 millisecond range. Competitive, though hardly unheard of.
Demo videos show the system distinguishing between a colleague in a meeting asking the AI a question versus discussing something with a teammate. Another clip shows proactive suggestions during drawing and coffee-making tasks—the system offering help when it senses hesitation, staying quiet when you're in flow. The company launched a developer API supporting real-time WebRTC streaming and sub-second latency, plus a data capture iOS app for structured video and audio collection with guided recording workflows. That last bit suggests they're serious about training data quality, perhaps the most underrated competitive advantage in this space.
Ishiki's near-term focus, per Yadav, is "realtime socially aware AI for meetings." That positions them adjacent to, but distinct from, existing meeting copilots. Google Meet's "Take notes for me" and Zoom AI Companion excel at summaries and action items—the retrospective work. They don't attempt live participation with addressee awareness, mostly because getting that wrong in a client meeting would be catastrophic. The company raised a $500,000 convertible note with YC as an investor, according to CB Insights data, though official company sources haven't confirmed specifics.
Elsewhere in the voice AI infrastructure stack, companies are racing to provide the building blocks for better addressee behavior. Hume AI's EVI 2 and 3 offer speech-to-speech models with emotion detection baked in. Cartesia Sonic delivers 135 millisecond text-to-speech latency. ElevenLabs launched agent capabilities and an Expressive mode. Vapi provides a full voice agent platform. Retell AI claims to automate 75 to 80 percent of support calls for one US insurer and processes over 40 million monthly minutes, with annual recurring revenue exceeding $35 million as of December 2025, per company disclosures.
These infrastructure plays benefit from the broader shift toward full-duplex interaction, barge-in capabilities, and lower latencies across the board. But they still largely leave addressee logic to developers—a primitives problem masquerading as a product problem.
The Smart Glasses Test Case

If you want to understand why addressee detection matters, strap an AI to your face for a day.
Barron's estimates Meta will ship over seven million Ray-Ban smart glasses units in 2025, leading the nascent category. Multiple industry analyses sourced to IDC point toward double-digit million annual shipments by mid-decade. When you're wearing an AI all day, every day, it needs to understand social context or it becomes intrusive noise. The kind of noise you eventually stop wearing.
The on-device AI trend accelerates this requirement. Apple Intelligence with Private Cloud Compute, the Foundation Models API, Copilot+ PCs with NPUs delivering higher TOPS for local inference—all of it makes privacy-preserving, always-on assistants more culturally acceptable. But only if they demonstrate restraint. An assistant that listens constantly but speaks rarely is a feature. One that chatters unprompted is surveillance with a personality.
Regulatory timelines reinforce the need for what industry insiders are calling "conservative abstention strategies"—knowing when not to activate. The EU AI Act's staggered compliance deadlines mean general-purpose AI provider obligations hit this August, with most high-risk requirements following in August 2026. Biometric consent frameworks like Illinois's BIPA create genuine litigation risk for companies processing voiceprints without careful consent handling. The FCC's proposed disclosure rules for AI in calls add yet another layer.
The voice assistant market, per Astute Analytica projections, could hit nearly $60 billion by 2033. Realizing that potential, though, requires solving the addressedness problem at scale, not just in lab conditions. The industry seems to be converging on multimodal fusion as the answer: combining audio, video, and conversational state to infer intent. Apple's research on using LLMs with ASR uncertainty shows one path. Amazon's DVAD demonstrates another. Google's gaze-based approach offers a third.
Different implementations, same core bet: that human communication is inherently multimodal, and single-channel solutions will always fall short.
What Comes Next

Ishiki Labs enters this landscape with a thesis that feels both obvious and radical: that addressee detection deserves first-class status, not afterthought engineering. Whether their Fern API gains traction among developers building meeting assistants, smart glasses apps, or ambient agents will hinge on execution—can they actually deliver on those latency promises at scale? Do their models generalize across accents, environments, and the chaotic acoustic conditions of real life?
The company is part of YC's Winter 2026 batch, which holds Demo Day on March 24, 2026. By then, we'll know whether their core insight—that timing matters as much as content—resonates with developers tired of duct-taping together addressee logic from primitives.
For founders building voice AI products, the strategic calculus is sharpening. Latency matters, obviously. Accuracy matters. But knowing when to speak and when to stay silent might matter most of all.
The multimodal AI boom has given us systems that can see, hear, and respond faster than ever before. The next competitive frontier, perhaps the defining one, is teaching them to read the room. Because the alternative—assistants that talk over you, wake at phantom triggers, and sleep through actual requests—isn't just bad UX.
It's existential risk for the category.
