Sanchit Monga and Shubham Malhotra want you to forget about cloud computing. At least when it comes to running AI on your MacBook.
Last week, their startup RunAnywhere—part of Y Combinator's winter batch—released MetalRT with a claim bold enough to raise eyebrows across the developer community: they've built the fastest AI inference engine for Apple Silicon. According to benchmarks the company published on March 3, their proprietary engine clocks 658 tokens per second on an M4 Max running the Qwen3-0.6B model. That's faster than Apple's own MLX framework. Faster than llama.cpp, the widely adopted open-source alternative. By margins that, if they hold up, matter.
The timing isn't accidental. As more developers push AI workloads to the edge—off cloud APIs, onto the laptops and phones people already own—speed has become the quiet bottleneck nobody wants to talk about. A chatbot that thinks for three seconds before responding feels broken. Voice assistants that lag feel like toys.
RunAnywhere's March 3 blog post laid out specific numbers: decode throughput between 1.10× and 1.19× faster than MLX using identical model files. Against llama.cpp, the advantage widened—1.35× to 2.14× faster, depending on which model you're running.
Not incremental gains. Real ones, if the benchmarks are accurate.
What the Numbers Actually Say
The company tested MetalRT against three established frameworks: Apple's MLX, llama.cpp, and Ollama. Hardware was consistent—an M4 Max with 64GB of RAM running a recent version of macOS. Models were 4-bit quantized variants: Qwen3-0.6B, Qwen3-4B, Llama-3.2-3B, and LFM2.5-1.2B. All tests used greedy decoding across five runs, with the company reporting the best result from each batch.
Standard practice, though it does leave questions about average performance and consistency across runs.
That peak figure—658 tokens per second—came from the smallest model tested, Qwen3-0.6B. For context, that's fast enough that generated text appears nearly instantaneous, like someone typing quickly in a chat window. The speedups over Ollama were more dramatic still, ranging from 1.41× to 2.40× depending on configuration.
Six days later, on March 9, RunAnywhere published speech benchmarks. MetalRT transcribed 70 seconds of audio in 101 milliseconds using a 4-bit Whisper Tiny model—714 times faster than real-time, by their count. Text-to-speech took 178 milliseconds to generate four words, which the company said was 2.8 times faster than MLX's audio implementation.
Impressive numbers. Self-published numbers.
A Voice Assistant You Can Actually Try

Alongside MetalRT, the company released RCLI—an open-source, on-device voice assistant for macOS that runs entirely from the command line. The setup is straightforward: install via Homebrew, run the configuration script, download about a gigabyte of models, and you're running a functional speech-to-text, LLM reasoning, and text-to-speech pipeline locally.
No cloud. No API keys. No waiting.
RCLI supports 38 macOS actions out of the box across categories including Productivity, Communication, Media, System, and Web—sending iMessages, controlling Spotify, creating reminders, taking screenshots. The documentation suggests sub-200 millisecond end-to-end voice latency on Apple Silicon, though that figure depends entirely on which models you choose. Smaller models respond faster; larger ones trade speed for nuance and accuracy.
The tool bundles streaming speech recognition using zipformer, offline transcription with Whisper and Parakeet options, multiple text-to-speech voices from Piper and Kokoro, and document retrieval using hybrid vector and BM25 search. The founders note that retrieval runs in roughly four milliseconds over 5,000-plus chunks—quick enough to feel instant in actual use.
RCLI ships under an MIT license, which is genuinely open source. The catch? The MetalRT engine powering it isn't. On M3 or later chips, RCLI automatically uses MetalRT when available. On older M1 or M2 machines, it falls back to llama.cpp and sherpa-onnx for inference. One of the cofounders assured Hacker News commenters that the system is "fully local—no data is collected," though the company hasn't published a formal privacy policy to back that claim.
Two Founders, One SEC Filing
RunAnywhere is the work of Monga and Malhotra, both engineers. According to an SEC Form D filing dated October 16, 2025, the Delaware C-corp had raised $10,000 of a planned $1 million SAFE offering as of that date. The company lists between two and 10 employees and operates out of San Francisco and Redwood City, California—the familiar geography of startups trying to make it.
Beyond RCLI, RunAnywhere is building a broader on-device AI platform with SDKs for Swift, Kotlin, React Native, Flutter, and WebAssembly. The documentation emphasizes Metal acceleration on Apple devices, ONNX backends for speech tasks, and a control plane for over-the-air model updates and policy management. It's infrastructure, in other words, for developers who want to ship AI features without the recurring costs or latency penalties of cloud inference.
The company launched on Hacker News on March 11, 2026. Within 24 hours, the post had drawn 235 points and 147 comments—a respectable showing for a technical product. The discussion surfaced both enthusiasm and skepticism in roughly equal measure. Some praised the performance claims and the fact that RCLI ships as working code rather than a demo video. Others raised questions about reproducibility, pointing out that the benchmarks are self-published and haven't been independently verified.
Fair point.
The Landscape They're Entering

MetalRT is stepping into a competitive, fragmented field. Apple's MLX framework has become something of a de facto standard, backed by Apple's machine learning research team and tightly integrated with the Metal GPU API. It's open source, actively maintained, and supports a growing ecosystem of tools like mlx-lm for language models and mlx-audio for speech tasks.
Llama.cpp and Ollama, meanwhile, have become go-to solutions for running quantized models locally. Llama.cpp uses Metal backends for GPU acceleration on Macs; Ollama wraps it with a friendlier server interface. Both have seen performance improvements in recent releases—Ollama's 0.17 update specifically targeted Apple Silicon optimizations—though community forums remain full of reports about regressions and inconsistent behavior.
There are also newer efforts. A preprint on vLLM-MLX, published in January, described a native Apple Silicon adaptation of the vLLM serving engine that achieved 525 tokens per second on an M4 Max in certain configurations. The numbers aren't directly comparable to MetalRT's—different models, different testing parameters—but they illustrate an active optimization race with multiple teams chasing similar goals.
RunAnywhere's claim of being the fastest is, for now, unverified outside the company's own testing. The methodology looks reasonable—greedy decoding, multiple runs, consistent hardware—but reproducibility will determine whether these numbers hold up when others try to replicate them.
What's Available Today

For teams building on-device AI applications, MetalRT offers something tangible. The RCLI demo runs on any Mac with M3 or later silicon and installs in a few terminal commands. The supported model list includes multiple Qwen variants, LFM options, and several speech models. It's not vaporware. Developers can download it and run it this afternoon.
The broader SDK platform is still seeking design partners, according to RunAnywhere's Y Combinator launch page. Documentation covers installation via Swift Package Manager for iOS and macOS projects, with examples for Android, React Native, and Flutter integrations. The control plane features—model updates pushed over-the-air, usage analytics, routing policies for hybrid cloud fallback—suggest the founders are thinking beyond raw speed toward the operational concerns that matter in production deployments.
Whether MetalRT's performance advantage proves durable is another question entirely. Inference optimization is a moving target. MLX and llama.cpp aren't standing still; both continue to evolve, and Apple has considerable resources to throw at the problem if it chooses.
But for now, RunAnywhere has put specific numbers on the board and shipped working code. For developers frustrated by slow local inference or spiraling cloud costs, that's enough to merit a closer look. Perhaps more than the founders initially expected, given how quickly the Hacker News thread filled up.
The real test comes next: when independent developers start running their own benchmarks and building their own applications on top of MetalRT. That's when we'll learn whether this is genuinely the fastest inference engine for Apple Silicon, or just the first to make that claim publicly.
