Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechMarch 12, 2026

Rasa Legal Raises $5M to Democratize Criminal Record Clearance

Rasa Legal Raises $5M to Democratize Criminal Record Clearance
Legal TechAutomation+2
SaaS iconSaaSMarch 12, 2026

Scanner Raises $22M Series A as AI Agents Reshape Security Operations

Scanner Raises $22M Series A as AI Agents Reshape Security Operations

Founders Mentioned

Sanchit Monga

RunAnywhere

saas icon
SaaS

Shubham Malhotra

RunAnywhere

saas icon
SaaS

Sanchit Monga

RunAnywhere

saas icon
SaaS

Shubham Malhotra

RunAnywhere

saas icon
SaaS
SaaS iconSaaS
March 12, 2026
YcAi InfrastructureModel OptimizationMacosOpen Source

RunAnywhere Claims Fastest AI Inference Engine for Apple Silicon

YC W26 startup launches MetalRT, claiming 2x faster AI inference than MLX and llama.cpp on Macs, plus open-source voice assistant reaching 658 tokens/second on M4 Max.

RunAnywhere Claims Fastest AI Inference Engine for Apple Silicon

Sanchit Monga and Shubham Malhotra want you to forget about cloud computing. At least when it comes to running AI on your MacBook.

Last week, their startup RunAnywhere—part of Y Combinator's winter batch—released MetalRT with a claim bold enough to raise eyebrows across the developer community: they've built the fastest AI inference engine for Apple Silicon. According to benchmarks the company published on March 3, their proprietary engine clocks 658 tokens per second on an M4 Max running the Qwen3-0.6B model. That's faster than Apple's own MLX framework. Faster than llama.cpp, the widely adopted open-source alternative. By margins that, if they hold up, matter.

The timing isn't accidental. As more developers push AI workloads to the edge—off cloud APIs, onto the laptops and phones people already own—speed has become the quiet bottleneck nobody wants to talk about. A chatbot that thinks for three seconds before responding feels broken. Voice assistants that lag feel like toys.

RunAnywhere's March 3 blog post laid out specific numbers: decode throughput between 1.10× and 1.19× faster than MLX using identical model files. Against llama.cpp, the advantage widened—1.35× to 2.14× faster, depending on which model you're running.

Not incremental gains. Real ones, if the benchmarks are accurate.

What the Numbers Actually Say

The company tested MetalRT against three established frameworks: Apple's MLX, llama.cpp, and Ollama. Hardware was consistent—an M4 Max with 64GB of RAM running a recent version of macOS. Models were 4-bit quantized variants: Qwen3-0.6B, Qwen3-4B, Llama-3.2-3B, and LFM2.5-1.2B. All tests used greedy decoding across five runs, with the company reporting the best result from each batch.

Standard practice, though it does leave questions about average performance and consistency across runs.

That peak figure—658 tokens per second—came from the smallest model tested, Qwen3-0.6B. For context, that's fast enough that generated text appears nearly instantaneous, like someone typing quickly in a chat window. The speedups over Ollama were more dramatic still, ranging from 1.41× to 2.40× depending on configuration.

Six days later, on March 9, RunAnywhere published speech benchmarks. MetalRT transcribed 70 seconds of audio in 101 milliseconds using a 4-bit Whisper Tiny model—714 times faster than real-time, by their count. Text-to-speech took 178 milliseconds to generate four words, which the company said was 2.8 times faster than MLX's audio implementation.

Impressive numbers. Self-published numbers.

A Voice Assistant You Can Actually Try

Digital illustration for article section "A Voice Assistant You Can Actually Try" in "RunAnywhere Claims Fastest AI Inference Engine for Apple Silicon" - A clean, minimalist 3D conceptual illustration representing an on-device macOS voice assistant, feat...

Alongside MetalRT, the company released RCLI—an open-source, on-device voice assistant for macOS that runs entirely from the command line. The setup is straightforward: install via Homebrew, run the configuration script, download about a gigabyte of models, and you're running a functional speech-to-text, LLM reasoning, and text-to-speech pipeline locally.

No cloud. No API keys. No waiting.

RCLI supports 38 macOS actions out of the box across categories including Productivity, Communication, Media, System, and Web—sending iMessages, controlling Spotify, creating reminders, taking screenshots. The documentation suggests sub-200 millisecond end-to-end voice latency on Apple Silicon, though that figure depends entirely on which models you choose. Smaller models respond faster; larger ones trade speed for nuance and accuracy.

The tool bundles streaming speech recognition using zipformer, offline transcription with Whisper and Parakeet options, multiple text-to-speech voices from Piper and Kokoro, and document retrieval using hybrid vector and BM25 search. The founders note that retrieval runs in roughly four milliseconds over 5,000-plus chunks—quick enough to feel instant in actual use.

RCLI ships under an MIT license, which is genuinely open source. The catch? The MetalRT engine powering it isn't. On M3 or later chips, RCLI automatically uses MetalRT when available. On older M1 or M2 machines, it falls back to llama.cpp and sherpa-onnx for inference. One of the cofounders assured Hacker News commenters that the system is "fully local—no data is collected," though the company hasn't published a formal privacy policy to back that claim.

Two Founders, One SEC Filing

RunAnywhere is the work of Monga and Malhotra, both engineers. According to an SEC Form D filing dated October 16, 2025, the Delaware C-corp had raised $10,000 of a planned $1 million SAFE offering as of that date. The company lists between two and 10 employees and operates out of San Francisco and Redwood City, California—the familiar geography of startups trying to make it.

Beyond RCLI, RunAnywhere is building a broader on-device AI platform with SDKs for Swift, Kotlin, React Native, Flutter, and WebAssembly. The documentation emphasizes Metal acceleration on Apple devices, ONNX backends for speech tasks, and a control plane for over-the-air model updates and policy management. It's infrastructure, in other words, for developers who want to ship AI features without the recurring costs or latency penalties of cloud inference.

The company launched on Hacker News on March 11, 2026. Within 24 hours, the post had drawn 235 points and 147 comments—a respectable showing for a technical product. The discussion surfaced both enthusiasm and skepticism in roughly equal measure. Some praised the performance claims and the fact that RCLI ships as working code rather than a demo video. Others raised questions about reproducibility, pointing out that the benchmarks are self-published and haven't been independently verified.

Fair point.

The Landscape They're Entering

Digital illustration for article section "The Landscape They're Entering" in "RunAnywhere Claims Fastest AI Inference Engine for Apple Silicon" - A contemporary 3D illustration depicting a new entity entering a competitive, fragmented landscape, ...

MetalRT is stepping into a competitive, fragmented field. Apple's MLX framework has become something of a de facto standard, backed by Apple's machine learning research team and tightly integrated with the Metal GPU API. It's open source, actively maintained, and supports a growing ecosystem of tools like mlx-lm for language models and mlx-audio for speech tasks.

Llama.cpp and Ollama, meanwhile, have become go-to solutions for running quantized models locally. Llama.cpp uses Metal backends for GPU acceleration on Macs; Ollama wraps it with a friendlier server interface. Both have seen performance improvements in recent releases—Ollama's 0.17 update specifically targeted Apple Silicon optimizations—though community forums remain full of reports about regressions and inconsistent behavior.

There are also newer efforts. A preprint on vLLM-MLX, published in January, described a native Apple Silicon adaptation of the vLLM serving engine that achieved 525 tokens per second on an M4 Max in certain configurations. The numbers aren't directly comparable to MetalRT's—different models, different testing parameters—but they illustrate an active optimization race with multiple teams chasing similar goals.

RunAnywhere's claim of being the fastest is, for now, unverified outside the company's own testing. The methodology looks reasonable—greedy decoding, multiple runs, consistent hardware—but reproducibility will determine whether these numbers hold up when others try to replicate them.

What's Available Today

Digital illustration for article section "What's Available Today" in "RunAnywhere Claims Fastest AI Inference Engine for Apple Silicon" - A clean, minimal 3D conceptual illustration of a sleek, modern laptop resting on a soft, uncluttered...

For teams building on-device AI applications, MetalRT offers something tangible. The RCLI demo runs on any Mac with M3 or later silicon and installs in a few terminal commands. The supported model list includes multiple Qwen variants, LFM options, and several speech models. It's not vaporware. Developers can download it and run it this afternoon.

The broader SDK platform is still seeking design partners, according to RunAnywhere's Y Combinator launch page. Documentation covers installation via Swift Package Manager for iOS and macOS projects, with examples for Android, React Native, and Flutter integrations. The control plane features—model updates pushed over-the-air, usage analytics, routing policies for hybrid cloud fallback—suggest the founders are thinking beyond raw speed toward the operational concerns that matter in production deployments.

Whether MetalRT's performance advantage proves durable is another question entirely. Inference optimization is a moving target. MLX and llama.cpp aren't standing still; both continue to evolve, and Apple has considerable resources to throw at the problem if it chooses.

But for now, RunAnywhere has put specific numbers on the board and shipped working code. For developers frustrated by slow local inference or spiraling cloud costs, that's enough to merit a closer look. Perhaps more than the founders initially expected, given how quickly the Hacker News thread filled up.

The real test comes next: when independent developers start running their own benchmarks and building their own applications on top of MetalRT. That's when we'll learn whether this is genuinely the fastest inference engine for Apple Silicon, or just the first to make that claim publicly.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Rasa Legal Raises $5M to Democratize Criminal Record Clearance
  • Scanner Raises $22M Series A as AI Agents Reshape Security Operations
  • Kled AI Raises $5.5M to License Hollywood Content for AI Training
  • OpenCFO Raises $2M Seed to Build AI-Native Finance OS for CFOs
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.