Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Fintech iconFintechMay 11, 2026

Haun Ventures Closes $1B Fund Betting on Crypto, AI Agents

Haun Ventures Closes $1B Fund Betting on Crypto, AI Agents
Venture CapitalCrypto Trading+3
SaaS iconSaaSMay 11, 2026

AI-Powered Micro-Startups Hit $300K ARR With 90% Margins, No VC

AI-Powered Micro-Startups Hit $300K ARR With 90% Margins, No VC
AiB2b Saas+3

Founders Mentioned

Parth Radia

Keyframe Labs

saas icon
SaaS

Kaahan Radia

Keyframe Labs

saas icon
SaaS

Parth Radia

Keyframe Labs

saas icon
SaaS

Kaahan Radia

Keyframe Labs

saas icon
SaaS
SaaS iconSaaS
May 11, 2026
Ai AgentsVideo GenerationEnterprise AiB2b Saas

AI Agents Go Visual: The Race to Build Real-Time Video Calls

As enterprise AI shifts from text to face-to-face interaction, startups and tech giants compete to deliver photorealistic video agents at scale—and at pennies per minute.

AI Agents Go Visual: The Race to Build Real-Time Video Calls

The executive dialed in expecting a familiar Zoom grid. What greeted her instead: a face that blinked on cue, nodded in sync with her questions, held eye contact without the telltale robotic stare. The conversation felt natural enough to be unsettling. She was speaking to an AI agent, not a colleague—and at six cents per minute, her company could theoretically deploy hundreds more.

This is where enterprise software has landed in 2026. Conversational AI, long confined to text boxes and disembodied voices, is now insisting on showing its face. The industry's latest wager—a massive one, judging by the capital and engineering talent pouring in—is that video interaction, even when the face is synthetic, unlocks something fundamentally different from chat or audio alone. The technology is shipping, the pricing structure is approaching parity with offshore labor, and the scramble to scale photorealistic agents has moved from lab demonstrations to production rollouts.

Whether this represents a genuine shift in how businesses handle customer interaction, or an expensive detour on the way to something else entirely, remains an open question.

A Market in Motion

The numbers suggest enterprise AI has crossed a threshold. McKinsey's State of AI report from late 2025 found 88% of organizations using AI in at least one business function—though it's worth noting this reflects 2025 data. Gallup reported in April 2026 that half of U.S. employees use AI at work, with 28% using it daily or weekly. But the activity still centers on text prompts and voice queries. Video represents uncharted territory—and the players building it span garage startups to trillion-dollar platforms betting their roadmaps on the premise that faces matter.

Google launched Gemini 3.1 Flash Live on March 26, 2026, a production API designed for real-time voice and vision agents capable of processing live video streams and handling phone calls. Microsoft updated its Copilot Studio documentation less than two weeks earlier to include real-time voice agents, though the service currently operates only in North America. Zoom announced AI avatars for meetings on March 10, pairing the feature with deepfake-detection tools—a nod to the trust issues that shadow any synthetic video product. The message from the big platforms: this is no longer R&D.

The startup ecosystem is moving at comparable speed. D-ID released V4 "Expressive Visual Agents" on March 16, emphasizing real-time digital humans connected to large language models for enterprise deployments. Kaltura went generally available with Agentic Avatars on March 12, then followed up with its Avatar Video Production Studio on May 7. DeepBrain AI launched real-time interactive avatars in late April. HeyGen migrated users to its new LiveAvatar platform by the end of March, built on WebRTC for low-latency streaming.

And then there's Keyframe Labs. A two-person team from Y Combinator's Spring 2026 batch—founders Parth Radia and Kaahan Radia—offering photorealistic video calls at $0.06 per minute with latency hovering around 500 milliseconds. The pricing undercuts many voice-only agent stacks, which raises questions about margins, infrastructure choices, and how long those rates can hold.

The Economics Start to Work

Digital illustration for article section "The Economics Start to Work" in "AI Agents Go Visual: The Race to Build Real-Time Video Calls" - A perfectly symmetrical, minimalist conceptual composition representing the economics of video agent...

Three forces are converging to push video agents from demo to deployment, though not without friction.

The cost structure has become defensible, if not exactly cheap. Voice-only AI agents already run between $0.05 and $0.20 per minute depending on provider and setup. Adding photorealistic video at similar rates shifts the business case. Keyframe Labs markets its service at $0.06 per minute; several voice agent platforms cite comparable figures. Twilio's telephony layer alone costs around $0.013 to $0.014 per minute for U.S. outbound calls. Factor in speech recognition, LLM inference, text-to-speech, and video rendering, and you're approaching parity with offshore human agents in specific use cases.

Gartner predicted in January 2026 that conversational AI could trim roughly $80 billion from global contact-center labor costs in 2026. But the same research warned that by 2030, GenAI cost-per-resolution might actually exceed some offshore human costs. The arbitrage window, in other words, may be narrower than the hype cycle suggests. Perhaps significantly narrower.

The technical stack has matured enough to support production deployments. WebRTC has emerged as the standard for real-time avatar streaming, showing up in products from HeyGen to Azure Speech Avatar to LiveKit's agent framework. OpenAI's Realtime API, which went generally available in August 2025, established patterns for voice-first agents that others are now adapting for video. Research papers like "Avatar Forcing" in January and "Hallo-Live" in April demonstrate latencies between 500ms and 940ms with improved facial expressiveness. NVIDIA open-sourced its Audio2Face animation model in September 2025, giving developers a foundation for driving 3D avatars in real time. The infrastructure layer—LiveKit, Agora, Twilio—has aligned around treating avatars as first-class participants in calls and conference rooms.

Enterprise expectations have shifted alongside the technology. TechRadar framed 2026 as "the year enterprise AI finally gets to work" in an April 3, 2026 publication, citing Gartner research predicting that nearly half of enterprise applications will include task-specific AI agents within twelve months. Axios called it AI's "show me the money" year. The pressure to demonstrate return on investment is forcing companies beyond pilot programs. Video agents offer a tangible differentiator—Keyframe Labs' founders argue that "face-to-face interaction triggers responses that text and voice simply can't." The psychology behind that claim is still debated, but early adopters are willing to test the premise at scale.

Different Bets, Different Strategies

Digital illustration for article section "Different Bets, Different Strategies" in "AI Agents Go Visual: The Race to Build Real-Time Video Calls" - A conceptual, minimalist still life featuring a beautifully crafted, tactile modular block suspended...

The approaches vary in ways that reveal competing theories about where the market is headed.

Keyframe Labs is optimizing for developer velocity. The company offers a no-code widget and low-code embeds—documentation promises users can "drop a snippet into your application to go live in minutes." Its Persona-1 and Persona-1-Live models claim faster-than-real-time performance on RTX 4090, L40S, and H100 GPUs, with real-time factors as low as 0.0625 on high-end hardware. An early blog post from November 2025 showed pricing examples with Persona-1 at $0.025 per minute compared to $0.050 per minute for competitors, though current rates have drifted higher. The pitch is straightforward: photoreal, emotionally responsive, scalable—and fast enough to embed in customer support, sales, or onboarding without rearchitecting your stack.

The established avatar vendors are taking enterprise-first paths. D-ID's V4 emphasizes LLM connectivity and expressiveness. Kaltura's Agentic Avatars, generally available since March, bundle conversational capabilities with a production studio for creating avatar assets. HeyGen's LiveAvatar platform integrates with LiveKit and replaces its older "Interactive Avatar" product—a signal that the company is betting on open standards over proprietary infrastructure. DeepBrain AI positions its AI Studios product for customer experience at scale. The common thread: Fortune 500 buyers who need SLAs, compliance certifications, and white-glove support.

The hyperscalers are embedding avatars into platforms where distribution is guaranteed. Google's Gemini Live API supports multimodal real-time agents out of the box. Microsoft's real-time voice agents in Copilot Studio are designed for enterprise workflows with Azure hosting. Zoom is adding avatars directly to its meeting platform—users will see AI representatives in the grid alongside human colleagues, with detection tools flagging synthetic participants. The big-tech play is about collapsing adoption curves. If avatars become a standard feature in Copilot or Zoom, the timeline shrinks from years to months.

Infrastructure providers are enabling all of the above, and then some. LiveKit's agent SDK now supports avatars as secondary participants with synchronized audio and video tracks. Agora announced partnerships in February to enable "Physical AI" agents. Livepeer is promoting open networks for real-time AI video pipelines. The abstraction layers are multiplying. The barrier to building a custom video agent stack is dropping—not quite to zero, but close enough that developers are stitching together APIs and launching products in weeks.

Reality Check on Three Fronts

Digital illustration for article section "Reality Check on Three Fronts" in "AI Agents Go Visual: The Race to Build Real-Time Video Calls" - A minimalist and conceptual still life featuring a sleek, modern video production slate resting alon...

The video agent market is about to collide with reality in ways that will separate serious products from well-funded experiments.

Regulation is arriving, ready or not. The EU AI Act's Article 50 transparency obligations take effect on August 2, 2026, requiring clear labeling of deepfakes and synthetic content. A draft Code of Practice on marking and labeling is expected by June, with final requirements due in the fall. In the U.S., the FCC declared AI-generated voices in robocalls illegal in February 2024 under the Telephone Consumer Protection Act. The FTC finalized its impersonation rule around the same time. State laws like Tennessee's ELVIS Act, effective July 2024, protect voice and image likeness from unauthorized AI cloning. Platform policies are hardening: YouTube began requiring disclosure of realistic synthetic media in March 2024. The C2PA launched Content Credentials 2.3 in February 2026 with live video provenance support, though labels can still be stripped on re-encoding.

The regulatory patchwork will force product teams to build disclosure flows, audio chimes, and metadata tagging into their UX from the start, not as an afterthought. That adds engineering time, design complexity, and user friction—costs that don't show up in the per-minute pricing.

The cost structure is more complicated than headline rates suggest. Gartner's January 2026 warning that GenAI cost-per-resolution could exceed offshore human agent costs by 2030 highlights a gap between per-minute pricing and per-outcome economics. A $0.06-per-minute agent that takes five minutes to resolve an issue costs $0.30. A human agent in the Philippines might resolve the same issue for less, with fewer follow-up tickets and better customer satisfaction scores. Engineering blogs cite latency budgets, frame gating, interruption handling, and tool-calling pauses that add friction users notice. The "demo versus production" gap remains wide. Early adopters are discovering that self-assembled stacks—combining speech recognition, real-time LLMs, text-to-speech, and video generation—can push all-in costs closer to $0.15 per agent-minute once you factor in guardrails, error handling, and the infrastructure needed to keep it running at scale.

Trust is the constraint no one has solved. Zoom is shipping deepfake detection alongside its avatars for good reason. Axios reported a spike in deepfake scams in early 2026, including romance scams using live video calls. The FTC and Better Business Bureau have highlighted surges in voice-clone fraud. If users cannot reliably distinguish between real and synthetic faces on video calls, the erosion of trust will outpace any technical gains. The industry's answer—C2PA watermarking, detection algorithms, mandated disclosures—may work in enterprise settings with clear contexts. The consumer internet will be messier. Product leaders building video agents should plan for user skepticism, not just technical scale.

But opportunities remain, and they're substantial. The convergence on $0.05 to $0.06 per minute for real-time video agents suggests a commoditization dynamic that will favor developers who can integrate quickly and iterate on use cases. The shift from "humans watching video" to "AI understanding video," as framed in recent research on AI-oriented RTC frameworks, opens adjacent markets: training, compliance review, spatial computing companions. Gartner predicts up to 40% of enterprise applications will include task-specific agents by the end of 2026. Photorealistic video could become table stakes in customer-facing roles, or it could prove to be a feature users tolerate rather than prefer.

The race is on. The pricing is converging. The platforms are live, the startups are shipping, and the enterprise buyers are testing at scale. Whether video agents become the next wave of business software or a well-funded detour depends on factors that have little to do with the quality of the rendering: trust, regulatory compliance, and the ability to prove value beyond the initial novelty.

For now, the $0.06-per-minute agents are ready to take your call. Whether you'll want to take theirs is another question entirely.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Haun Ventures Closes $1B Fund Betting on Crypto, AI Agents
  • AI-Powered Micro-Startups Hit $300K ARR With 90% Margins, No VC
  • YC-Backed Superlog Launches AI Agent That Auto-Fixes Production Bugs
  • YC-Backed CellType Lands Pharma Deal for AI-Driven Drug Discovery
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.