Beknazar Abdikamalov remembers the moment he realized building a decent voice agent had become almost comically complicated. Dozens of speech-to-text engines, each with different latencies and error rates. Language models that varied wildly in turn-taking ability. Text-to-speech voices that could handle currency formatting or completely butcher it. "You're basically running a science experiment every time you want to change providers," he said in an online forum introducing his new company.
That frustration led to Speko, a four-person San Francisco venture that participated in Y Combinator and launched quietly this past August. The startup sells itself as an automated matchmaker for voice infrastructure: hand it your language and use case, and its router picks the speech-to-text, language model, and text-to-speech combination most likely to perform well based on continuously updated public benchmarks. Developers pay a 5% markup on whatever the underlying provider charges, or $0.09 per minute if they want Speko to handle the entire stack.
The company went live in late August with what amounts to a bet that the proliferation of voice models has outpaced most teams' ability to evaluate them. For customer service bots, medical transcription tools, or narration engines, choosing the wrong cascade of models can mean the difference between smooth conversations and expensive failures. Speko's pitch: measurement-driven selection that routes each call through the best-performing combination for a given task.
Orchestration as infrastructure
Speko Router functions as a hosted layer between voice applications and the two dozen or so model providers it tracks. A developer supplies one API key, specifies whether they're building for customer service or transcription or something else, and the router consults rankings published at benchmarks.speko.ai. Those rankings compare word error rates, latency percentiles, and cost per minute or per million characters.
The system runs on anycast infrastructure via AWS Global Accelerator, according to company documentation. Health checks ping each provider; if one fails, the router skips it automatically. Transports include HTTP, server-sent events, and WebSocket. The documentation is explicit about one constraint: "Router never changes the billing source during a request," meaning mid-call provider swaps don't happen.
Alongside the managed router, Speko released Gateway, an MIT-licensed open-source runtime that lets teams run providers directly while optionally feeding routing choices back to Speko's benchmark engine. It ships with native hooks for LiveKit and Pipecat, two frameworks popular among voice-agent builders who prefer to manage their own keys and infrastructure.
The approach sidesteps the end-to-end managed agent model that competitors like Retell AI offer, where pricing ranges from $0.07 to $0.31 per minute depending on which models a customer selects. Speko instead positions itself as the layer that decides which models to use, not the agent runtime itself.
Public benchmarks as moat

Walk through Speko's leaderboards and you see granular comparisons that most developers would need days to run themselves. The speech-to-text board from late August listed Universal-3.5 Pro at a 2.0% word error rate and $0.0075 per minute, GPT-4o Transcribe at 2.3% WER and $0.0060 per minute, Velma 2 at 4.4% WER and $0.0010 per minute. Text-to-speech rankings use Elo scores, median latency, and cost per million characters to compare voices from Google, ElevenLabs, Deepgram, Cartesia, and Speechify.
Early blog posts detail edge-case testing: how models handle code-switching between languages mid-sentence, whether they correctly read "-$12.50" aloud, which ones stumble on turn-taking cues in natural conversation. "We don't train or sell models ourselves, that's precisely how we keep our rankings impartial," Abdikamalov wrote when the company launched on Hacker News.
Continuous updates to the benchmark data form the backbone of Speko's value proposition. The company makes everything public, a transparency play aimed at developers skeptical of black-box routing claims.
Threading a market needle

The competitive landscape splits along a few fault lines. OpenRouter added dedicated audio APIs for speech and transcription earlier this year, but positions itself as a multi-modality request router rather than a live-session orchestrator. LiveKit provides runtime infrastructure with swappable model providers but leaves benchmarking and selection to customers.
Abdikamalov drew the distinction in a forum exchange: "A voice agent is not a batch request... a gateway that terminates at http can route the requests inside a call; it can't route the call." The argument hinges on real-time session management, where latency and failover matter more than they do for one-off API hits.
Whether that distinction holds up as models converge in performance remains an open question. Component pricing on Speko's site ranges from $0.0010 per minute for budget STT engines to $12.00 per million output tokens for high-end language models. New signups get $100 in credit; bring-your-own-keys workflows through Gateway remain free.
Abdikamalov previously co-founded Hupo as CTO and worked at Amazon before starting Speko, according to the company's Y Combinator profile. The team of four operates out of San Francisco. LinkedIn shows a founding year of 2025 and a listed headcount between two and ten employees, though that range likely reflects the platform's bracketing rather than meaningful ambiguity.
The Hacker News launch post drew 118 upvotes and scattered comments, including one from a user claiming their company had already adopted the router. Router and managed infrastructure remain in public preview without service-level guarantees; enterprise customers can negotiate custom terms. The OpenAPI 3.1 specification lives at relay.speko.dev/openapi.json, though documentation makes clear the service isn't OpenAI-compatible and requires native integration.
For a market where model options multiply faster than most engineering teams can benchmark them, Speko's wager is straightforward: someone has to do the measuring, and developers will pay a modest premium to offload that work. Whether 5% proves modest enough is a question the next few months will likely answer.
