There's a particular kind of hubris that comes with announcing a breakthrough on March 31st.
Babak Hassibi, a Caltech professor turned startup CEO, picked that date to unveil PrismML's answer to what he calls "one of the most stubborn problems" facing artificial intelligence: making sophisticated language models small enough to run on phones, laptops, and other devices that lack the bottomless memory and power of cloud servers. The company emerged from stealth that day claiming to have built the first commercially viable 1-bit large language models—a compression technique so aggressive it squeezes an 8-billion-parameter model into just over a gigabyte without, supposedly, the catastrophic performance loss that usually accompanies such dramatic size reduction.
The timing aside, the code is real. The models are downloadable. And the technical community is paying attention, if with the kind of wary curiosity reserved for claims that sound too good.
Smaller than should be possible
PrismML has released three models under an Apache 2.0 license: Bonsai 8B, 4B, and 1.7B. All are available now on Hugging Face. The company has reported raising $16.25 million through SAFE and seed rounds led by Khosla Ventures and Cerberus Ventures, with additional compute support from Google and Caltech—a detail that matters, because training these models from scratch required significant TPU resources.
What makes them unusual isn't just their size. It's the method.
Traditional quantization takes an existing model trained at full precision and compresses it after the fact, accepting some degradation in exchange for efficiency. PrismML's approach is different—perhaps more radical. The company trained these models from scratch using binary weights: every parameter set to either plus or minus one, with no middle ground. That includes the embeddings, the attention mechanisms, the language modeling head. The entire architecture.
Well, almost.
Here's where the marketing meets the mathematics. To make the binary system work, PrismML stores one FP16 scale parameter for every 128 weights, which brings the effective precision to roughly 1.125 bits per weight. It's a clever workaround, and it's also the reason the models function at all. Truly 1-bit systems, with no scaling factors, tend to collapse into incoherence at this scale.
The Bonsai 8B model, according to its Hugging Face card, occupies 1.15 gigabytes of weight memory—fourteen times smaller than the FP16 equivalent. The 4B variant: 0.57 GB. The smallest, 1.7B, sits at roughly 0.24 GB. For context, that's small enough to fit comfortably on older smartphones with tight storage constraints.
The performance question
Size is one thing. Utility is another.

Hassibi is positioning these models as competitive with full-precision alternatives despite the extreme compression. On a battery of six standard benchmarks—IFEval, GSM8K, HumanEval+, BFCL, MuSR, and MMLU-Redux—Bonsai 8B scores an average of 70.5. On MMLU-Redux alone, it hits 65.7, a number that immediately sparked debate on Reddit's r/LocalLLaMA forum when the launch went live.
The speed claims are equally striking, if difficult to verify independently. PrismML says the 8B model runs 6.2 times faster than FP16 baselines on an RTX 4090 GPU, while consuming one-fifth the energy per token generated. The 4B variant shows a 4.2x speedup on the same hardware. For mobile chips, the company claims the 4B model reaches 132 tokens per second on an M4 Pro, while the 1.7B processes at 130 tokens per second on a recent-generation iPhone Pro Max.
Those numbers come with asterisks the company itself acknowledges, to its credit. Native hardware optimized for 1-bit operations doesn't exist yet. Some mobile power estimates are projections rather than direct measurements. And the benchmarks compare against full-precision models that, as PrismML notes dryly, "continue to advance." In other words: the goalposts are moving.
Independent verification hasn't surfaced yet. It's early.
The infrastructure catch
Downloading the models turns out to be the easy part.
To actually run them, developers need PrismML's custom forks of llama.cpp and Apple's MLX frameworks. The models use a grouped quantization approach with 128 weights per scale factor, which standard llama.cpp doesn't support. One developer on Reddit spent the better part of an evening trying to load Bonsai-4B.gguf before discovering this particular dependency, a detail buried in documentation that didn't make the initial announcement.
PrismML has published a demo repository on GitHub with setup scripts and a whitepaper outlining the technical approach. For those working in Python or Swift on Apple devices, MLX support is available through the company's forks. The startup lists Locally AI as a deployment partner for iPhone integration, though details on broader mobile SDKs remain thin.
This is the kind of friction that determines whether a research project becomes production infrastructure or remains a curiosity.
Not the first to try
PrismML isn't alone in chasing extreme quantization. Microsoft's BitNet b1.58, a ternary model using three weight levels—minus one, zero, and plus one—at a smaller 2-billion-parameter scale, demonstrated near-parity with full-precision models and kicked off a wave of academic interest in specialized hardware architectures designed for low-precision computation.
The distinction between ternary and binary weights might sound academic, but it has implications for chip design and memory access patterns. PrismML's binary approach with grouped scales represents a different bet: that you can scale to larger parameter counts without the third weight level. Whether that thesis holds—especially as full-precision models continue their relentless improvement—is the question that will decide if this is a genuine breakthrough or simply an impressive incremental step.

The industry has seen plenty of both.
What it means for practitioners
For CTOs evaluating edge AI deployments, PrismML's models offer something testable: the hypothesis that you can run capable language models on constrained hardware without leaning on cloud infrastructure. The Apache 2.0 license removes the usual licensing friction. The open weights mean organizations can benchmark the claims themselves before committing engineering resources.
PrismML has coined a metric it calls "intelligence density"—negative log of one minus score, divided by model size—which it calculates as 10.8 times higher than full-precision Qwen3-8B. It's a provocative framing, suggesting that compression delivers more capability per byte even when raw performance lags behind state-of-the-art. For applications where latency, privacy, or offline operation matter more than absolute accuracy, that tradeoff might make sense.
Maybe.
The company is hiring across Pasadena and San Francisco, though employee count remains undisclosed. The founding team includes Sahin Lale, Omead Pooladzandi, and Reza Sadri as VP of Strategy alongside Hassibi. The Caltech connection runs deep—compute resources, institutional backing, a research-first sensibility that comes through in the technical documentation.
Now comes the hard part
The models are available. The forks are public. The whitepaper is live on GitHub.
What's missing is the independent replication that transforms vendor claims into validated benchmarks, and the upstreaming work that turns research code into something production systems can depend on. The March 31st launch date still raises eyebrows—though the repositories contain functioning code and the models produce actual outputs, not vaporware.
Whether they produce reliable outputs at enterprise scale is what engineering teams will now get to discover for themselves. That process, more than any press release, will determine whether PrismML has cracked AI's size problem or simply found another interesting way to compress it.

