Every mobile developer knows the feeling. You've built a beautiful AI feature that works perfectly on your latest flagship iPhone, reasonably well on last year's Samsung Galaxy, and then crawls—or worse, crashes outright—on anything older. You're wrestling with three different inference frameworks, manually tuning model files for disparate NPU architectures, wondering if that Android device your QA team just tested even supports half the ML operations your app requires.
The fragmentation isn't just annoying. For many, it's existential.
This is the problem RunAnywhere, a Y Combinator-backed startup, emerged to solve. The startup's arrival, announced in January, comes at a moment when the entire on-device AI landscape is simultaneously exploding and splintering—creating both the opportunity and, perhaps more urgently, the need for infrastructure that can tame the chaos.
The Growth Story, By the Numbers
The figures tell a straightforward tale of expansion. Grand View Research projected in March 2024 that the global edge AI market would balloon from $24.91 billion in 2025 to $118.69 billion by 2033—a compound annual growth rate of 21.7%. Mordor Intelligence, reporting in February 2025, placed the edge AI hardware market at $25.08 billion for 2025, climbing to $30.74 billion in 2026, with smartphones commanding a 46.68% share.
These aren't just projections riding a hype cycle. By the end of last year, Gartner estimated that AI-capable PCs had reached 31% of worldwide shipments. The firm projected that figure would hit 55% in 2026, roughly 143 million units. Canalys data from early 2025 suggested similar numbers, with AI-capable PCs reaching 30% penetration that year and 50% by 2026—figures largely aligned with Gartner's forecast.
The mobile story runs parallel, though messier in execution. Google's ML Kit GenAI Prompt API entered alpha in October, bringing Gemini Nano to Android devices through the AICore framework. Chrome's Prompt API, available in Chrome 138 and later, runs Gemini Nano locally in the browser. Apple has been shipping its M4 chip with a 38 TOPS Neural Engine since May 2024, and in March announced the M5 family with an upgraded 16-core Neural Engine optimized for Apple Intelligence workloads. Qualcomm's Snapdragon X2 Elite and Plus chips for laptops deliver around 80 TOPS of INT8 performance, according to press materials from late last year and early this year. In March, Qualcomm unveiled the Snapdragon Wear Elite for wearables, claiming support for on-device models up to 2 billion parameters.
Microsoft's AI Economy Institute reported in January that global generative AI adoption had reached approximately 16.3% of the world's population in the second half of 2025. Perhaps more telling: at Google I/O last year, the company noted that 7 million developers were building with Gemini, and Vertex AI usage had grown 40x year-over-year.
The infrastructure is arriving. The demand is real. But beneath those rosy projections lies something gnarlier.
The Fragmentation Nightmare
The edge AI ecosystem is fractured across hardware, runtime frameworks, and deployment models in ways that make cross-platform development genuinely painful. Apple runs Core ML and MLX on its own silicon, with reports from early March suggesting the company may unveil a unified "Core AI" framework at WWDC this year. Google offers AICore and ML Kit for Android, plus the Chrome Prompt API for browsers. Meta released ExecuTorch 1.0 as generally available in July 2025, optimizing it across Meta's own apps and partnering with Arm in November to extend KleidiAI optimizations across different power envelopes. Intel pushes OpenVINO GenAI—the 2026.0 release in March added mixture-of-experts support, improved NPU coverage, and enhanced compression. NVIDIA maintains TensorRT and TensorRT-LLM for its Jetson platform.
Then there are the cross-platform runtimes: ONNX Runtime Mobile, MLC LLM, WebLLM, llama.cpp. Each has its advocates and its gaps. A developer targeting iOS, Android, and web simultaneously must navigate at least three distinct stacks. Often more. Each runtime optimizes differently for hardware it knows about and degrades unpredictably on hardware it doesn't.
This creates a testing matrix that grows exponentially. An app shipping on both iOS and Android must account for Apple's Neural Engine generations, Qualcomm's Hexagon DSP and Snapdragon NPU variants, Samsung's Exynos chips, MediaTek's Dimensity processors, and whatever else sneaks into the device mix. Device manufacturers layer their own modifications atop Android's ML frameworks. Performance benchmarks from one device rarely translate cleanly to another, even within the same chipset family.
It's a mess, frankly.
RunAnywhere, founded by Sanchit and Shubham, positions itself as an abstraction layer over this fragmentation. The company's pitch, articulated at launch in January, centers on a single SDK supporting on-device LLM inference, speech-to-text, text-to-speech, and voice activity detection across Swift, Kotlin, Flutter, and React Native. Version 0.17.5 of the SDK, released in February, delivered cross-platform updates for all four frameworks. The open-source repository on GitHub shows active development through early this year, with developer posts on Reddit and Flutter forums referencing alpha and beta programs.
More intriguing, though, is the enterprise control plane. RunAnywhere's platform includes tooling for over-the-air model rollouts, fleet-wide policies, and analytics—capabilities critical at scale but rarely bundled into developer SDKs. The startup describes itself on its Y Combinator profile as "the default way to run on-device AI at scale."
That framing matters. It isn't just a library. It's infrastructure for managing AI across fragmented hardware at production volume.
Hardware Convergence, Software Divergence

While runtimes diverge, hardware is converging on a shared set of capabilities, even if implementations differ wildly. The trend is toward higher TOPS counts, lower power consumption, and tighter integration between CPU, GPU, and NPU.
NVIDIA's Jetson platform exemplifies the trajectory. JetPack 6.2, documented through early this year, introduced "Super" modes for the Orin NX and Orin Nano modules. The Orin NX gained a 70% boost in TOPS; both modules saw bandwidth improvements. Developer kits and production modules now support heavier workloads at the edge. LiteVLA-Edge, a quantized vision-language-action model detailed in a March arXiv paper, runs entirely on Jetson Orin-class hardware. At CES in January, Innoviz showcased its SMARTer platform combining scalable LiDAR with Jetson Orin Nano for real-time edge processing.
Hailo, an Israeli edge AI accelerator startup that raised $85 million in August last year, launched the Hailo-10H in July—a 40 TOPS M.2 accelerator targeting generative AI workloads. HP incorporated the Hailo-10H into point-of-sale systems and workstations. At CES, Hailo demonstrated deployments across security, retail, robotics, and healthcare, including partnerships with Truen, Vicon, and Advantech for Evolv's threat detection systems.
SiMa.ai, which also closed an $85 million round in August 2025, announced an automotive AI blueprint with Synopsys in January, positioning its MLSoC platform for automotive and industrial use cases. Qualcomm expanded the Snapdragon roadmap in late last year and early this year, with Alex Katouzian, the company's senior vice president for mobile, compute, and XR, emphasizing NPU-accelerated agentic AI as the near-term direction for mobile and PC platforms.
Academic research is keeping pace, too. A March arXiv paper detailed methods for efficient reasoning on the edge under mobile constraints. Another from December introduced Parallax, a runtime parallelization technique achieving up to 46% latency reduction versus state-of-the-art approaches. On-device continual learning for industrial anomaly detection appeared in a December paper. CFIS-YOLO, published last April, achieved 135 frames per second on the BM1684X edge processor with just 17.3% of the original model's power consumption and a 0.5 percentage point drop in mean average precision.
The hardware is ready. The question—and it's not a trivial one—is whether the software ecosystem can keep up.
The Platform Hedges
The major platforms are hedging their bets with hybrid approaches, splitting the difference between on-device privacy and cloud-scale compute. Apple's Private Cloud Compute, announced in June 2024 and entering production in October last year, represents a new standard for privacy in cloud-augmented AI. When a task exceeds on-device capacity, requests route to attested, stateless cloud nodes running the same model families. The system preserves privacy guarantees while enabling heavier workloads. Apple made server code available for public security review and emphasized that requests are ephemeral—no data persists after inference completes.
Google's strategy spans on-device, in-browser, and cloud tiers. The AICore framework and ML Kit GenAI Prompt API bring Gemini Nano to Android devices that meet hardware thresholds. The Chrome Prompt API enables local LLM inference in browsers on devices with sufficient resources. Both systems fail over to cloud endpoints when local capacity is insufficient or unavailable. Android Authority reported through last year that early-access partners for Gemini Nano included select flagship devices, but OEM adoption has expanded this year as more hardware meets the requirements.
Meta's ExecuTorch 1.0, released in July last year, powers on-device inference across Instagram, Facebook, WhatsApp, and Meta's hardware ecosystem. The November partnership with Arm deepened optimization efforts, targeting power envelopes from milliwatts to megawatts. Meta frames ExecuTorch as production-grade infrastructure for running PyTorch models at the edge—a deliberate contrast to frameworks that prioritize research flexibility over deployment stability.
Intel's OpenVINO GenAI releases through early this year added GenAI pipeline support, GGUF import paths, and improved NPU plugin performance. The 2026.0 release in March introduced mixture-of-experts support and smarter compression. Intel even previewed an ExecuTorch backend in OpenVINO 2025.1, signaling an intent to interoperate rather than compete head-on with Meta's runtime.
These aren't rival standards so much as parallel bets. Each platform owns part of the stack—Apple controls iOS hardware and OS, Google shapes Android and Chrome, Meta influences social app distribution—and each is optimizing for the constraints it knows best.
Developers building across platforms must still navigate the seams. And those seams aren't getting any narrower.
The Infrastructure Gap

What's missing is the connective tissue—the layer that handles model management, versioning, rollout orchestration, and fleet observability at scale. Most runtimes assume developers will handle these themselves.
NVIDIA's Fleet Command offers managed edge AI deployment for GPU-based systems, but it's designed for industrial and enterprise IoT deployments, not mobile apps. AWS IoT Greengrass and Azure IoT Edge provide edge compute frameworks, but Azure's tooling faced discontinuations and transitions last year, and Greengrass leans heavily toward AWS's broader IoT ecosystem. Edge Impulse has built an end-to-end platform for embedded and tiny ML with over-the-air updates and a studio environment, but its focus remains sensors and microcontrollers, not smartphones and tablets running large language models.
DeGirum released a hardware-agnostic edge SDK and AI Hub last year, with support for accelerators like Hailo and Axelera. It's closer to what mobile developers need, but adoption remains nascent. The gap persists: a unified developer experience for deploying, updating, and monitoring AI models across heterogeneous mobile fleets at enterprise scale.
This is where RunAnywhere's control plane becomes interesting—and where the founders seem to be placing their bet. The platform doesn't just abstract runtime differences; it layers on policy management, OTA rollout controls, and analytics across the device fleet. A February blog post from RunAnywhere demonstrated an on-device browser agent built with the SDK, showing how local models can power autonomous interactions without server round trips.
The challenge, as with any abstraction layer, is whether it can keep pace with underlying platform changes. Apple may unveil Core AI at WWDC this year. Google and Qualcomm are iterating on AICore and Snapdragon features quarterly. Meta ships ExecuTorch updates tied to its app release cycles. A third-party SDK must either track these changes religiously or risk becoming a lowest-common-denominator layer that misses platform-specific optimizations.
That's a treadmill few startups manage to stay on indefinitely.
Regulation Accelerates the Shift

Regulation is accelerating the shift to on-device AI, perhaps unintentionally. The EU's AI Act entered into force in December 2024. The majority of its provisions become applicable in August, with certain high-risk embedded systems covered by Annex II deferred until August 2027. The European Commission has signaled guidance delays into early this year, creating uncertainty around compliance timelines, but the direction is clear: high-risk AI systems face transparency, documentation, and governance requirements that are easier to satisfy when models run locally and data never leaves the device.
The EU's Cyber Resilience Act, which took effect in December 2024, imposes incident and vulnerability reporting requirements starting in September, with full cybersecurity obligations—including software bill of materials mandates, conformity assessments, and vulnerability disclosure policies—kicking in by December 2027. On-device processing reduces the attack surface for certain classes of data exfiltration and cross-border transfer risks, though the CRA's lifecycle security requirements still apply to the software and firmware stack.
In the United States, state-level privacy laws are proliferating in a way that makes compliance increasingly complex. Colorado's right-to-cure provision for privacy violations sunset on December 31 last year, meaning enforcement began immediately this year. Connecticut lowered its applicability threshold mid-year. Utah introduced a new right to correct data in July. Colorado's AI Act, effective June 30, places obligations on deployers and developers of high-risk AI systems to prevent algorithmic discrimination in specified domains.
Keeping data on-device simplifies compliance with some provisions but doesn't eliminate the need for impact assessments and transparency measures. The regulatory burden creates an incentive—perhaps even a commercial imperative—for platforms that can demonstrate where models run, what data they touch, and how updates are controlled. A control plane that logs model versions, rollout percentages, and device-level telemetry becomes a compliance artifact as much as an operational tool.
RunAnywhere's enterprise focus—policies, rollouts, analytics—maps directly to the documentation requirements emerging from Brussels and U.S. state capitals. Whether that proves to be strategic foresight or merely good timing remains to be seen.
The Bifurcation Ahead
The on-device AI market will likely bifurcate, and the split is already visible. Platform-native developers—those building exclusively for iOS or Android—will increasingly rely on Apple's and Google's first-party frameworks. Core ML, AICore, and their successors offer the tightest hardware integration and the fastest access to new chip features. As Apple potentially consolidates its ML stack under a unified Core AI umbrella and Google extends Gemini Nano availability across more devices, the platform-native path becomes more attractive. Why fight the platform when you can ride it?
Cross-platform developers face a harder choice. The testing matrix for on-device AI is brutal, and it's growing. Gartner's projection of 55% AI PC penetration this year means the laptop and desktop install base is also fragmenting across NPU capabilities. Developers shipping AI features to Windows, macOS, iOS, Android, and web must either maintain separate codebases per platform or adopt an abstraction layer that trades some performance for portability.
That's the wedge for infrastructure startups like RunAnywhere. If the SDK can deliver consistent performance across platforms and the control plane can simplify fleet management, the value proposition is straightforward. Enterprises running AI in production apps care less about squeezing every last TOPS from the hardware and more about predictable behavior, manageable rollouts, and visibility into what's happening across millions of devices.
The risk is commoditization. If OpenVINO, ONNX Runtime, and MLC LLM converge on a common set of backends and APIs, the differentiation narrows to tooling and support. If Apple, Google, and Meta coordinate on interoperability—unlikely, admittedly, but not impossible—third-party abstraction layers lose their reason to exist.
Regulation could tip the scales. As compliance requirements grow, platforms that bundle policy controls, audit trails, and lifecycle management will have an edge over raw inference libraries. The YC W26 batch is betting on that thesis, RunAnywhere among them. Whether it proves out depends on how fast the fragmentation problem worsens and how quickly the major platforms respond.
For now, the edge remains a mess. The hardware is maturing, yes. The models are shrinking. The use cases are multiplying. But the infrastructure to manage it all at scale—especially across platforms—remains a work in progress.
That gap won't stay open forever. It rarely does. The question is who closes it first, and whether there's room for more than one winner.
