The developer community has a saying: if the numbers sound too good to be true, someone probably forgot to carry a zero. So when Andrey Mikhaylov claimed his open-source runtime could squeeze a 26-billion-parameter AI model into just 2 gigabytes of RAM—on an entry-level MacBook Air, no less—the reaction was predictable. Skepticism, tinged with curiosity.
Yet TurboFieldfare, published on GitHub on July 31, 2026, appears to do exactly that. By keeping only a skeletal "shared core" and key-value cache in memory while streaming everything else from solid-state storage on the fly, Mikhaylov reports hitting 5 to 6 tokens per second on hardware that shouldn't even load the model in the first place. An 8GB M2 Air, to be precise—the kind of machine college students buy and consultants carry through airport security.
The timing feels deliberate, though perhaps more than Mikhaylov intended. Memory prices are climbing again. PC shipments are forecast to drop 11 percent in 2026, squeezed by component shortages and a market correction. And on August 2, 2026, the European Union's AI Act kicks in with obligations for general-purpose AI providers—compliance burdens that make on-device inference suddenly look less like a technical curiosity and more like strategic necessity.
For engineering leaders trying to figure out where to run large language models, the calculation is shifting. It's no longer just cloud versus edge. It's: what happens when the old tradeoffs—speed, cost, control—stop making sense?
The Hardware Bind
The infrastructure math has been unforgiving. Take Google's Gemma 4 26B, a mixture-of-experts model released under Apache 2.0 license back in April. It activates roughly 3.8 billion parameters per token. Load it the conventional way, even quantized down to 4 bits, and you're looking at 14 to 16 gigabytes of memory on Apple Silicon. Base-config laptops need not apply.
Yet demand for local AI keeps accelerating—Counterpoint Research projects that GenAI-capable smartphones will represent 45 percent of global shipments this year. Google rolled out its AI Edge Gallery in the spring, demonstrating that smaller Gemma variants could run entirely offline on recent iPhones at speeds above 50 tokens per second. Ollama, the open-source developer tool that's become something of a cultural touchstone in the local LLM movement, closed a $65 million Series B on July 9, 2026 while serving nearly 9 million monthly developers.
The hardware ecosystem, meanwhile, is creaking. IDC's outlooks through mid-year paint a picture of memory constraints depressing shipments while pushing the product mix upmarket toward "AI PCs" with 24GB or more of unified memory. Enterprises are quietly shifting production inference workloads to private cloud, according to Broadcom's June outlook—driven partly by regulatory anxiety, partly by total-cost-of-ownership arithmetic that doesn't pencil out on hyperscaler pricing.
Something had to give.
The Optimization Wave

What's changed, really, is the stack. Over the past year, a cascade of optimization techniques—some from research labs, others bubbling up from open-source communities—has rewritten the ceiling on what consumer hardware can handle.
Google Research's TurboQuant, published in late March, demonstrated KV-cache compression down to 3 to 3.5 bits effective, claiming 6× memory reduction and up to 8× speedups. Open-TQ-Metal, which followed in April, showed 48× attention speedups at 128,000-token context windows by operating directly in compressed space—cutting KV memory from 40GB to 12.5GB. BaseRT, a native Metal runtime published in July, reported up to 6.4× prefill improvements over llama.cpp in some configurations on M5 Pro chips.
These aren't lab curiosities. The vllm-mlx paper from January documented 21 to 87 percent higher throughput than llama.cpp across model sizes from 600 million to 30 billion parameters on an M4 Max, with continuous batching multiplying throughput by 4.3×. Llama.cpp itself, despite occasional Metal-related regressions that frustrated users, has absorbed mixture-of-experts support, multi-token prediction, and a growing ecosystem of forks. BeeLlama v0.3.1, released in May, integrated speculative decoding and TurboQuant compression ahead of the upstream merge.
But TurboFieldfare takes a different approach entirely. Instead of loading the entire model and compressing it, Mikhaylov's engine exploits the architecture of mixture-of-experts models themselves. It keeps only the shared transformer layers—roughly 1.35 gigabytes—and an FP16 key-value cache resident in memory. An 8-bit router selects the top eight experts per layer. A bounded 16-slot LRU expert cache handles frequently-used experts; cache misses trigger parallel read calls that fetch the needed weights into Metal-visible buffers directly from a custom ".gturbo" file layout, about 14.3GB on disk.
The whole thing, in other words, treats the SSD like very slow RAM.
When the Community Started Counting

The claimed numbers—31 to 35 tokens per second on a 24GB M5 Pro, and 5.1 to 6.3 on an 8GB M2 Air—generated immediate discussion when the repository went live on July 31. Some of it was enthusiastic. Much of it was skeptical.
A back-of-the-envelope calculation on Reddit's r/LocalLLaMA suggested trouble: 4 billion active weights at roughly 4 bits would require more than 10 gigabytes per second of sustained disk reads at 6 tokens per second. That exceeds the sequential read speeds of many consumer SSDs, particularly the base M2 Air's 256GB single-NAND configuration, which independent reviewers clock at 1.4 to 2.3GB/s. The physics didn't seem to work.
Mikhaylov's design counters with caching and overlap. The 16-slot expert cache, combined with the router's locality—similar prompts tend to reuse the same experts—means the engine isn't actually streaming 4 billion parameters per token in practice. How much it's streaming, exactly, depends on workload characteristics that vary by the second. Independent replication on base-model SSDs remains, as of this writing, pending.
The broader concept isn't new. An earlier community demonstration ran a 397-billion-parameter MoE model on an M1 Ultra with 64GB total memory, peaking at just 14GB RAM usage by keeping 20 resident experts and paging the rest from SSD. The throughput—1.59 tokens per second—was glacial, but it proved viability. FlexGen, published back in March 2023, pioneered weight and KV offloading from GPU to CPU and SSD, achieving 1 token per second on OPT-175B with a 16GB GPU through aggressive batching.
What distinguishes TurboFieldfare, if the numbers hold, is the claim that interactive latency is possible on 8GB consumer laptops. If verified, it shifts the bottleneck from memory capacity to I/O locality and kernel efficiency. A different set of engineering tradeoffs, certainly—but ones that open the door to 26-billion-class models on hardware that retails for $999.
The Strategic Calculation

There's a larger question lurking here, one that extends beyond benchmarks and disk bandwidth. On-device inference addresses several converging pressures at once.
The EU AI Act's general-purpose AI obligations, which entered application on August 2, 2026, push companies toward data minimization and residency. On-device stacks reduce data-in-transit risk, even if governance for tool-calling and agentic behaviors remains an unsolved problem. Cloud inference costs continue climbing; IDC and Omdia analysts, quoted in mid-year press coverage, suggest AI PCs represent a partial answer, with a potential inflection point in 2028 as memory costs normalize and agentic workflows mature.
For founders and CTOs evaluating infrastructure options, the trade space has widened considerably. Smaller models like Gemma 4 E2B and E4B already deliver usable offline experiences on phones through LiteRT-LM. Mainstream engines—Ollama, LM Studio—have absorbed MLX and speculative decoding, with users reporting 20 to 25 tokens per second on Mac mini M4 configurations with 24GB. Community posts document Gemma 4 E4B running on a Raspberry Pi 5 with 8GB and swap, albeit at 0.3 to 0.5 tokens per second. Slow, but functional.
TurboFieldfare, if it survives scrutiny, suggests that even 26-billion MoE models could run on constrained hardware at interactive speeds. But the caveats matter, perhaps more than the headline number.
The engine is model-specific—Gemma 4 26B-A4B text-only—and macOS-26/Metal-4-only. SSD endurance and energy costs from continuous weight paging are real concerns. Academic work warns that MoE weight offload can increase energy per token substantially compared to HBM or DRAM. And reproducibility depends on variables that shift month to month: disk capacity, macOS and Metal versions, unified-memory pressure. Llama.cpp users have reported a roughly 3.8× generation regression on Apple M4 between specific builds on Gemma 4 26B, a reminder that these stacks are fragile.
Still, the trajectory seems clear. The combination of open weights under permissive licenses—Google cited 400 million Gemma downloads as of early April—aggressive compression techniques, and community-driven optimization has turned what seemed like luxury hardware requirements into accessible engineering challenges. Whether TurboFieldfare's specific approach generalizes or fades into the noise of experimentation, the principle stands: mixture-of-experts architectures and SSD paging offer a path to run frontier-class models on consumer devices, if you're willing to trade memory for disk bandwidth and accept the complexity that comes with it.
For organizations navigating the AI pricing paradox—spiraling cloud costs on one side, regulatory and operational dependencies on the other—these developments matter. The future of inference may not be purely cloud or purely edge, but hybrid by design and necessity.
And sometimes that future arrives in a GitHub repository with a README and some benchmarks that make you wonder whether the impossible just became merely difficult. Or whether someone, in fact, forgot to carry a zero.
