A San Francisco startup has compressed an AI agent into a package smaller than most smartphone photos.
Cactus Compute released Needle 2, a 45-million-parameter language model that fits into a 14-megabyte binary and runs tool-calling agents on devices ranging from budget Android phones to microcontrollers, using just 28 megabytes of memory. The Y Combinator-backed company says the model solves a stubborn problem in edge computing: most mobile-sized language models still consume hundreds of megabytes and deplete batteries within hours.
Needle 2 runs entirely on-device, eliminating the server costs and network delays that plague cloud-dependent AI applications. The model has already found its way into production hardware. Eric Migicovsky, who founded the Pebble smartwatch, now ships Needle inside his Index 01 AI wearable. "We run Cactus Needle locally in the app," Migicovsky said. "The model's footprint is tiny and the performance never lets us down."
The technical achievement centers on what the model doesn't try to do. Unlike general-purpose language models that attempt conversation and content generation, Needle 2 outputs only function calls. When a user makes an off-topic request, the model returns an empty call rather than hallucinating a response. A built-in retrieval mechanism narrows candidate tools per turn, while grammar-constrained decoding enforces JSON schema rules, according to the model's Hugging Face documentation.
Cactus Compute trained Needle 2 on 115 billion tokens during pre-training, followed by 38 billion tokens of post-training data that included compact reasoning traces. The model uses a 256-token sliding window with tools pinned as key-value sinks, which keeps memory consumption around 28 megabytes even during extended interactions.
Trading Precision for Portability
The company's benchmarks reveal both the promise and limitations of extreme compression. On Google's Mobile Actions benchmark, which tests 961 device-control scenarios, Needle 2 scored 63.7% accuracy using ordered strict exact matching. That trails the 69.1% achieved by Liquid AI's LFM2.5-230M model running at full 16-bit precision, though it edges out Google's own FunctionGemma-270M at 64.0%. Apple's on-device foundation model scored 57.6%.

Needle 2 showed stronger performance on Seal-Tools benchmarks, which test tool-calling capabilities in both familiar and novel domains. The model hit 32.6% on in-domain tests, ahead of LFM2.5-230M at 26.9% and FunctionGemma at 16.3%. On out-of-domain cases, Needle 2 reached 28.7% compared to 17.0% for LFM2.5 and 15.6% for FunctionGemma.
Cactus Compute acknowledges two asymmetries that complicate direct comparison. Needle 2 runs at 2-bit quantization while the baseline models operated at 16-bit precision, and the model was optimized specifically for device control rather than general language understanding.
The speed gains from compression prove substantial in real-world hardware. A Raspberry Pi 5 decodes Needle 2 at roughly 500 tokens per second, according to the company. VR headsets like Meta's Quest 3S and Apple's Vision Pro range from 400 to 1,500 tokens per second. Budget Samsung A-series phones manage 300 to 700 tokens per second.

Architecture Departures
Needle 2 abandons the standard transformer architecture in favor of what Cactus Compute calls a Simple Attention Network. The design replaces the typical feedforward layer with a Hadamard MLP and incorporates grouped-query attention alongside "engram" key-value memory constructed from hashed n-gram tables. Henry Ndubuaku and co-authors detailed the approach in a paper submitted to arXiv in July.
The model's efficiency derives partly from arithmetic choices invisible to end users. Needle 2 requires about 70 megaFLOPs of computation per token, using 35 million "matmul-active" parameters. That compares to 87 megaFLOPs for a conventional 43-million-parameter transformer and 540 megaFLOPs for Google's FunctionGemma-270M. Apple's on-device model, by Cactus Compute's accounting, runs at roughly 6,000 megaFLOPs per token.
The inference engine adapts to available hardware at startup, selecting optimized kernels for instruction sets including SDOT, NEON, AVX2, RISC-V, and WebAssembly SIMD. The company ships static C++ binaries and libraries for ARM64, x86-64, ARMv7, and RISC-V processors, plus iOS, watchOS, and tvOS frameworks, Android builds, and a WebAssembly version that developers can test in a browser sandbox.
Open Access, Commercial Traction
The model weights carry an Apache 2.0 license and are available through Hugging Face. Python tooling installs via pip, and developers can fine-tune Needle 2 using LoRA adapters or define custom tools through Python decorators before exporting specialized binaries. The Python components use an MIT license, according to the project's GitHub repository.
Cactus Compute has not publicly disclosed any seed funding amount. The startup's first model, Needle v1, generated considerable attention when it appeared on Hacker News in May, drawing 776 upvotes and more than 200 comments. Whether the second iteration gains similar traction may depend less on benchmark scores than on how many hardware makers follow Migicovsky's lead in shipping the model inside consumer devices.
