When Apple quietly toggled a switch last January—enabling Apple Intelligence by default in iOS 18.3—most users probably didn't notice. No fanfare. No special event. Just a software update that suddenly made on-device AI the default posture for hundreds of millions of iPhones.
But for developers, it was something closer to a wake-up call. The move signaled that on-device AI had crossed an invisible threshold: from experiment to expectation, from nice-to-have to table stakes. And it exposed a problem that had been festering beneath the hype: how do you actually ship AI features that work reliably across a fragmented ocean of devices—each with different chips, wildly different memory constraints, and runtime behaviors that range from elegant to maddening?
Enter RunAnywhere, a startup just emerged from Y Combinator's Winter 2026 batch. Their pitch lacks the usual startup gloss. "Edge AI is inevitable," the founders wrote in their launch post, "but every device behaves differently, runtimes vary, and performance collapses under memory and power constraints." It's not the kind of messaging that wins TechCrunch headlines or lights up Twitter. Boring infrastructure work, really.
But maybe that's the point.
The Hardware Scramble
The on-device AI market isn't growing—it's exploding. Grand View Research pegs the sector at $10.8 billion in 2025, with projections to hit $75.5 billion by 2033. That's a 27.8% compound annual growth rate, the kind of number that makes venture capitalists sit up straight. North America currently commands 34.5% of the market, though the real action is happening in device proliferation.
Consider the PC space alone. AI-capable machines shipped 48 million units in 2024—roughly 18% of the PC market, per Canalys data. That figure is expected to more than double this year, reaching 100 million units and capturing 40% market share. Microsoft's Copilot+ PC specification set a new floor: at least 40 TOPS (trillion operations per second) of neural processing unit performance. Qualcomm's Snapdragon X clears that bar at 45 TOPS. Intel and AMD are scrambling to catch up.
Smartphones are moving even faster, though. Counterpoint Research forecasts that GenAI-capable phones will surge from 11% market share in 2024 to 43% by 2027—more than 550 million units annually, with an installed base topping a billion devices. Google has baked Gemini Nano into Android 15's core accessibility features, emphasizing offline and privacy-first operation. The message is clear: AI is going local.
Beneath all this, the chip layer is fragmenting in ways that would give any developer a headache. NVIDIA's Jetson Orin family alone spans from 67 TOPS on the Nano variant (running at 7–25W) to 275 TOPS on the AGX model. Hailo launched low-power GenAI accelerators targeting edge deployments. ABI Research projects RISC-V edge AI chipsets alone will reach 129 million shipments by 2030.
It's a hardware land grab. And land grabs, historically, create messy boundaries.
The Runtime Chaos
Meta's ExecuTorch reached 1.0 general availability status in partnership with Arm, offering what they call a unified PyTorch runtime across billions of Arm devices. The company detailed in July how ExecuTorch now powers on-device machine learning features across Instagram, WhatsApp, Messenger, and Facebook. It sounds seamless in a press release.
The reality is messier. Android deprecated its NNAPI (Neural Networks API) NDK as of Android 15, pushing developers toward vendor extensions, HAL implementations, and third-party runtimes like ExecuTorch and ONNX Runtime Mobile. Apple has its own Core ML stack. Browser vendors shipped WebGPU across Chrome, Edge, Firefox, and Safari, enabling GPU compute directly in browsers—but WebNN, which promises to route workloads intelligently to CPU, GPU, or NPU, remains under development with uneven adoption.
Translation: there's no standard way to deploy AI features across devices. Every platform has its own runtime, its own operator support, its own quirks. When an unsupported operator forces a model to fall back from GPU to CPU, latency spikes and memory usage balloons. Performance collapses. Users notice.
This is the infrastructure gap that Shubham Malhotra and Sanchit Monga, RunAnywhere's founders, are betting on. Their approach is pragmatic to the point of being almost dull: a single SDK plus a control plane for routing, policies, and observability across fragmented edge environments.
The company's open-source SDKs—supporting Swift, Kotlin, React Native, Flutter, and Web—have attracted roughly 8,700 GitHub stars. The unified API covers large language models, speech-to-text, text-to-speech, and voice activity detection, all privacy-by-default and offline-capable. Under the hood, RunAnywhere uses plugin backends: GGUF via llama.cpp, Core ML, MLX, and TensorFlow Lite. The control plane adds policy-based hybrid routing (prefer local, fall back to cloud), analytics, observability, and over-the-air model delivery with versioning.
It's an abstraction layer designed to hide the chaos. Whether it can actually deliver on that promise at scale remains to be seen.
Why Now? Economics, Privacy, Physics

Three forces are converging to make on-device inference not just viable but, arguably, necessary.
Start with cost. A UC Riverside study linked to Qualcomm found that on-device inference delivers roughly 90% power reduction versus cloud-based generative models. The Edge AI & Vision Alliance reported potential savings of up to 95% energy, 88% carbon, and 96% water when shifting inference from cloud to phones—though these figures come with scope caveats. Still, the directional pressure is unmistakable: running models locally is cheaper at scale.
Then there's privacy, which has shifted from nice-to-have to regulatory requirement. Apple's Craig Federighi framed the company's approach bluntly at WWDC: "You shouldn't have to hand over all the details of your life to be warehoused and analyzed in someone's AI cloud." Apple's Private Cloud Compute architecture combines on-device processing with a hybrid model where cloud components can't store or access user data.
Qualcomm CEO Cristiano Amon went further, calling AI "the new UI" while emphasizing that "AI at the edge" delivers immediacy, privacy, and cost benefits. It's marketing language, sure. But it's also directionally accurate.
The regulatory environment is tightening around these issues. The EU AI Act entered force on August 1, 2024, with staged obligations for general-purpose AI and high-risk systems rolling out through 2026–2027. The FDA, Health Canada, and the UK's MHRA have published guidance on Predetermined Change Control Plans for ML-enabled medical devices, requiring monitoring, rollback capabilities, and evidence trails—features that depend on robust edge infrastructure.
In the U.S., the Biden administration's Executive Order 14110 was rescinded in January 2025, but NIST's AI Risk Management Framework continues as a voluntary baseline across agencies and contractors.
Then there's simple physics. Memory shortages—DRAM and HBM (high-bandwidth memory)—are driving up costs and lead times. IDC warned in 2026 that skyrocketing RAM pricing could shrink the PC market by up to 9%. SK hynix projects the HBM market alone will expand at roughly 30% annually, reaching close to $98 billion by 2030.
Developers are responding by leaning harder on quantization, pruning, and distillation techniques just to fit models within mobile memory and power envelopes. It's not elegant, but it works. Sometimes.
a16z's "Big Ideas in Tech 2025" report flagged "small, on-device AI" as expected to dominate in usage due to economics, latency, and privacy. Their follow-up "Notes on AI Apps in 2026" argued for multi-model architectures and specialized edge experiences. The venture firm's thesis: hybrid is the future, but the pendulum is swinging toward local-first.
Whether that thesis holds depends largely on whether the infrastructure can keep pace.
What Success Looks Like (Maybe)
RunAnywhere's documentation includes memory guidance for model sizes and quantization on Android, acknowledging the realities developers actually face: a 7-billion-parameter LLM quantized to 4-bit might fit a flagship phone but choke a mid-tier device. The company's control plane aims to handle these edge cases programmatically—detecting available memory, routing to cloud when local execution would fail, collecting telemetry to optimize future rollouts.
The pattern isn't entirely new. NVIDIA's Fleet Command offers remote management for Jetson devices in industrial settings, handling end-to-end deployment and monitoring at the edge. One example: industrial vision quality assurance for beverage cans, where edge inference avoids bandwidth bottlenecks and latency spikes. AWS IoT Greengrass provides similar capabilities with SageMaker Edge Manager, offering OTA updates and on-device components for IoT fleets. Balena and Mender serve device fleets with delta updates and canary rollouts. StreamDeploy focuses on containerized OTA for edge AI devices and robots.
What RunAnywhere is attempting, though, is to bring that fleet-management approach down to consumer devices—phones, tablets, PCs, web browsers. It's a category that hasn't had a clear infrastructure layer, partly because the hardware and software stacks have been too fragmented and too fast-moving. Whether there's room for a third-party abstraction layer, or whether the platforms will eventually subsume this functionality themselves, is an open question.
Toyota's partnership with Invisible AI offers a glimpse of what's possible when edge AI infrastructure works properly. The companies deployed Jetson-based cameras for factory-floor vision analytics, processing data on-device to preserve privacy and minimize bandwidth. NVIDIA's own benchmarks from Sima Labs show Jetson Orin handling retail and surveillance workloads with sub-millisecond latency in DeepStream pipelines—proof that edge hardware can meet production demands when properly deployed.
Research is beginning to formalize hybrid patterns. Papers like AppealNet, PrivateLoRA, and PRISM outline "edge-first, cloud-assist" architectures where devices attempt local inference and offload selectively based on difficulty, confidence thresholds, or privacy requirements. A 2025 arXiv paper explored sustainability-aware LLM routing across edge clusters (Jetson Orin NX and Ada GPUs), finding that intelligent placement could reduce energy footprints without sacrificing performance.
The Tooling Arms Race

The deployment gap is widening faster than the tooling can close it. As of early 2025, developers confront a fragmented landscape: Core ML on iOS; deprecated NNAPI on Android; vendor SDKs from Qualcomm and MediaTek; ExecuTorch for PyTorch models; ONNX Runtime Mobile for cross-platform export; and early WebNN for browser-based NPU access.
Each runtime has different operator coverage, quantization support, and fallback behavior. A new 2025 runtime framework called Parallax is attempting to solve the fallback problem by parallelizing operations, but it's early days.
Quantization has become table stakes. Apple's Core ML documentation now includes detailed guidance on activation quantization and performance gains on the A17 Pro and M4 chips. Techniques like AWQ, OWQ, and SASQ push 4-bit and INT4 kernels for LLMs, squeezing models into mobile constraints. Tools like llama.cpp with GGUF quantization have become the de facto standard for local LLM deployment, offering cross-backend support across CPU, Metal, CUDA, ROCm, and Vulkan.
Competitive activity is heating up at the deployment tooling layer. RunLocal, recently highlighted in industry coverage, automates the porting of PyTorch models to QNN and TensorRT, claiming 30% better optimization and 70% faster development cycles. Aina, set to launch in Q4 2025, promises one-click deployment, AI-powered tuning, and canary rollbacks for edge ML. Advantech integrated Edge Impulse tooling—now part of Qualcomm—into its platforms to streamline edge AI deployment.
The browser, curiously, is emerging as a dark horse distribution channel. WebGPU's ubiquity means developers can ship GPU-accelerated ML inference to any modern device without native code. Early adopters are experimenting with local embeddings, on-device search, and privacy-preserving personalization—all running in JavaScript. WebNN's eventual maturity could unlock NPU access from web apps, potentially collapsing the gap between native and web performance.
Whether RunAnywhere can stay ahead of these platform shifts—Android's NNAPI deprecation, Jetson's evolving TensorRT-LLM support, WebNN's uneven rollout—and build an enterprise-grade control plane that differentiates from DIY solutions will determine if "boring infrastructure" translates to defensible value.
The Maintenance Nightmare Ahead
For founders building in this space, a playbook is crystallizing: use portable export formats (ONNX, ExecuTorch) to reduce fragmentation costs. Build hybrid routing policies to balance cost, latency, and privacy. Optimize models aggressively with quantization and selective compression. Consider browser distribution for reach where feasible. Instrument everything for fleet-scale observability.
Regulatory pressures will shape what's shippable. The EU AI Act's staged rollout through 2027 will force companies deploying high-risk AI systems to maintain audit trails, implement risk management, and demonstrate data minimization—easier with on-device processing, but not trivial. FDA's PCCP guidance means medical device makers using edge ML must prove they can monitor, update, and roll back models safely. ISO 26262 discussions around functional safety in ADAS systems with AI are ongoing, per a 2025 SAE technical paper.
RunAnywhere's traction—8,700 stars and Y Combinator backing—suggests developers are genuinely hungry for this kind of abstraction. The infrastructure gap is real. The pain is real.
But here's the uncomfortable truth: the on-device AI market isn't waiting for infrastructure to catch up anymore. It's already here. Hundreds of millions of devices are shipping with AI capabilities. Developers are deploying features today, cobbling together solutions from fragmented runtimes, praying they scale.
The question isn't whether someone will build the infrastructure layer. The question is whether it arrives in time to prevent the next wave of edge deployments from becoming a maintenance nightmare—or whether developers will simply accept that nightmare as the cost of doing business.
That's not a bet most would make willingly. But sometimes, in the rush to ship, we make it anyway.
