Three days. That's how long a startup barely out of its founding phase held the top position on the embedding world's newest benchmark, outscoring Google, OpenAI, and a field of well-funded competitors. When Octen-8B posted a Mean(Task) score of 0.8045 on the RTEB leaderboard in mid-January, it seemed almost improbable—a company founded in late 2025 had edged out Voyage (0.7812), Gemini (0.7602), and OpenAI's text-embedding-3-large (0.6567).
Then Voyage reclaimed the crown. The cycle resumed.
But the brief upset revealed something more consequential than leaderboard bragging rights: the infrastructure race powering AI search has accelerated to a point where technical agility matters more than incumbency. A well-executed approach, open-weight foundations, and domain-specific fine-tuning can now produce competitive results in months rather than years.
Whether that speed translates to sustainable advantage is another question entirely.
When Benchmarks Stop Playing Nice
RTEB—the Retrieval Embedding Benchmark—arrived last October with a mandate to fix a problem the industry had been politely ignoring. Embedding models were getting suspiciously good at the tests they'd been trained on. Memorization, not generalization, was driving scores higher.
Hugging Face and the MTEB team designed RTEB as a corrective. The benchmark combines open datasets with private ones, spanning legal documents, financial filings, healthcare records, code repositories, and more than twenty languages. Its primary metric, NDCG@10 (Normalized Discounted Cumulative Gain at position 10), measures how well a model surfaces relevant results in the top ten positions—a proxy for real-world utility.
The private component was intentional. By withholding portions of the evaluation set, RTEB aimed to force genuine performance rather than overfitting.
It also introduced a new wrinkle. As Octen's team noted in a blog post following their brief stint atop the leaderboard, the hybrid design "reduces overfitting but raises fairness questions about uneven access to private test data." Early or privileged access to hidden test sets could, in theory, inflate scores without meaningful technical progress. A GitHub issue tracking the concern remains open, with no resolution proposed.
Still, within weeks of its launch, RTEB became the proving ground. Hugging Face launched a dedicated leaderboard space in early January, and model providers scrambled to optimize.
Built for a Shifting Market
Octen's founder, Kuan Zou, left Alibaba Cloud in late 2025 with a thesis that sounded almost obvious in hindsight. Traditional search was ceding ground to AI agents and chatbots. Gartner had forecast in early 2024 that search engine volume would drop 25 percent by 2026—a prediction that appears to be tracking. Enterprises, meanwhile, were rebuilding their retrieval stacks around vector databases and embedding models.
Zou's pitch: "Search Infrastructure for AI."
Octen-Embedding-8B, the model that briefly topped RTEB, is a LoRA-finetuned variant of Qwen's Qwen3-Embedding-8B foundation. It produces 4,096-dimensional embeddings with support for context windows exceeding 32,000 tokens across multiple languages. The technical recipe—detailed in the company's January blog post—leaned on domain-specific synthetic data, hard negative mining, multi-positive utilization, and cross-device negative sharing. Perhaps most notably, Octen employed what it calls "fusion of domain adapters," essentially teaching the model to handle legal, financial, and healthcare retrieval simultaneously without sacrificing vertical-specific performance.
The model card lists three size variants (0.6B, 4B, 8B parameters), query and document input types, and API integration capped at 4,096 dimensions.
By March, Octen's broader pitch had crystallized around a proprietary distributed search engine claiming 99-millisecond average response times, minute-level index freshness, and capacity for over one million queries per second. A latency comparison table on the company's homepage shows P50 response times of 62 milliseconds in us-west-1—roughly four times faster than competitors like Tavily, Exa, and Perplexity on the SealQA Hard benchmark.
Bold claims, admittedly. Whether they hold under enterprise load remains to be tested.
The Leaderboard Keeps Moving

Octen's moment at the top lasted perhaps three days before Voyage AI announced voyage-4-large in mid-January, reclaiming first place. Voyage's model introduced a "shared embedding space" feature designed to let customers swap models without reindexing their entire vector database—a pragmatic touch that speaks to the operational headaches enterprises actually face.
The company also released voyage-4-nano as open weights and expanded distribution across MongoDB Atlas, Google Cloud, AWS, and Azure. A clear play for ubiquity.
Google, for its part, transitioned from text-embedding-004 to Gemini Embedding in late 2025 and early 2026, with community posts suggesting cutover timelines in January and February. By March, the company unveiled Gemini Embedding 2—a natively multimodal model in public preview.
Cohere had moved earlier. Embed Multimodal v4 launched in April 2025 with 128,000-token context, Matryoshka embeddings (allowing dimensions of 256, 512, 1024, or 1536), and unified text-plus-image representations. NVIDIA contributed Llama-Embed-Nemotron-8B, which reported state-of-the-art results on MMTEB multilingual retrieval last October. Jina released embeddings-v4 in June 2025, targeting multimodal document retrieval.
Stanford's AI Index 2025, published in early March, captured the velocity. Voyage-3-m-exp scored 74.03 on MTEB in early 2025, with NV-Embed-v2 trailing at 72.31. The report also flagged long-context retrieval challenges via benchmarks like RULER and HELMET, noting that gains on standard tests don't necessarily translate to reliability when documents exceed 100,000 tokens.
The leaderboard moves weekly. Sometimes daily.
Infrastructure Finally Catches Up

Model proliferation drove infrastructure evolution—or perhaps forced it. AWS launched vector search in ElastiCache last October, touting microsecond latency and high recall. Elastic integrated Jina embeddings and rerankers into its Open Inference API by late February. MongoDB expanded AI capabilities in early 2026, adding direct embedding and reranking APIs alongside Voyage model integrations.
Qdrant announced a $50 million Series B in March. Milvus released version 2.6 last June, optimizing for billion-scale search at lower cost, with a roadmap targeting version 3.0 and a "Vector Lake" architecture by late this year. Weaviate published HIPAA-compliant vector search patterns in June 2025, addressing healthcare's particularly thorny audit and privacy requirements.
By mid-2026, a standard technical stack had emerged: a high-capacity dense embedding model (typically 4B to 8B parameters), a strong cross-encoder reranker (such as BAAI's bge-reranker-v2-m3, widely adopted throughout 2025), and hybrid sparse-plus-dense search to balance recall and precision. Matryoshka embeddings—which let developers trade off between 256 and 1,536 dimensions depending on latency constraints—became table stakes. So did reindex-minimizing strategies like Voyage's shared embedding space.
Platform distribution accelerated adoption. Leading embedding and reranking models are now available as managed services across major clouds and database providers, reducing integration overhead. MongoDB Atlas, Elastic, and Cloudflare Workers AI all expanded their model catalogs in 2025 and early 2026, making it easier to swap in new embeddings without rewriting infrastructure.
One could argue the infrastructure finally caught up to the ambition. Or perhaps it's still catching up.
What the Leaderboards Miss

RTEB's arrival also highlighted an uncomfortable truth: fragmentation. Vertical benchmarks proliferated in 2025 and 2026. FinMTEB for finance (February 2025). PatenTEB for patent text (October 2025). MMTEB for cross-language retrieval (February 2025). Models that dominate one benchmark often stumble on others.
General-purpose leaderboards, it turns out, can't capture domain-specific performance. SuiteEval, proposed in February, aims to standardize IR evaluation workflows, but the underlying tension remains unresolved.
There's also the fairness issue Octen flagged. Private datasets mean uneven access. A model provider with early or privileged access to RTEB's hidden test sets could engineer higher scores without genuine technical progress. The trade-off is real: reduce overfitting risk, increase opacity risk.
Regulatory pressures are reshaping the landscape too. The EU AI Act entered force in August 2024, with major provisions applying in August 2026 and full enforcement ramping through 2027. High-risk AI systems—including some enterprise search and retrieval applications—will need to demonstrate provenance, auditability, and bias mitigation.
Copyright litigation adds another dimension. The New York Times v. OpenAI/Microsoft case, allowed to proceed last March, underscores how training data sourcing influences vendor credibility. Research on privacy-preserving vector retrieval—homomorphic encryption, embedding-space transforms—picked up in 2025, signaling that compliance and security will soon be baseline expectations rather than differentiators.
Where the Real Work Happens
Octen's brief hold on RTEB first place matters less than what it represents. A company founded in 2025 built a competitive embedding model in months by combining open-weight foundations with targeted fine-tuning and domain-specific data engineering. The leaderboard will keep shifting—it already has—but the barrier to entry has collapsed.
For technical founders and engineering leaders, the implications are straightforward. Embedding models are no longer a moat. The differentiation happens in retrieval architecture: how you generate hard negatives, how you fuse domain adapters, how you handle reranking at scale, how you keep indexes fresh as data changes by the minute.
The infrastructure layer is where the durable work lives. Vector databases. Hybrid search. Long-context reliability.
Investors tracking the embedding and vector database ecosystem should note the velocity. Enterprise customers are swapping models quarterly, sometimes monthly, as new benchmarks and capabilities emerge. The companies building reindex-free architectures, multi-tenant isolation, and compliance-ready retrieval pipelines are solving the problems that won't disappear in six months.
Octen's RTEB win was a moment.
Whether the company can turn that moment into infrastructure that endures—well, that's a longer story. One that's still being written.
