At a technical conference last November, Microsoft engineers unveiled something that seemed to violate the known laws of cloud computing: a fully isolated virtual machine that executed a task in less than a millisecond. Not 100 milliseconds, not 10—but 0.9 milliseconds, according to details the company published this February.
This wasn't a container sharing a kernel with its neighbors. It wasn't a lightweight WebAssembly sandbox. It was a complete VM with its own kernel boundary, the kind of hardware isolation that traditionally came with a performance penalty measured in hundreds of milliseconds.
For years, the serverless world has lived with a compromise that felt almost physical. Fast startup times? Sure, but you'd have to accept containers that share the underlying operating system kernel. Real isolation with separate VMs? Absolutely—just be prepared to wait. Cloudflare solved this by avoiding the problem altogether, using V8 JavaScript isolates that fire up in under 5 milliseconds. AWS Lambda leaned into the overhead, pre-warming containers so customers wouldn't notice the lag. The industry found its lanes and stuck to them.
Now those lanes are merging. Microsoft's Hyperlight project, AWS Firecracker's evolving snapshot capabilities, memory-forking techniques pioneered by companies like CodeSandbox and Fly.io—together, they're collapsing what seemed like an immutable gap between security and speed. Hardware-grade isolation boundaries are approaching startup times that were, until very recently, the exclusive domain of process-level tricks.
The Stakes Behind the Speed
Serverless isn't a science experiment anymore. Market researchers disagree on the exact numbers—Precedence Research pegs the global market reaching $92.22 billion by 2034, while Fact.MR projects growth from $28 billion this year to $92.6 billion by 2035—but the trajectory tells the real story. Serverless has evolved from infrastructure novelty to infrastructure default.
The security architecture, though, hasn't quite kept up. Standard containers still share the host kernel, and that shared boundary has shown cracks. Throughout 2024 and into 2025, multiple container-escape vulnerabilities emerged in runc, the runtime that underpins most container deployments. CVE-2025-31133, CVE-2025-52565, CVE-2025-52881—disclosed last November—underscored what security researchers have been saying quietly for years. TechRadar's coverage noted that Docker containers "may not be as secure as they like," a bit of understatement that carries weight coming from the mainstream tech press.
Cloud providers have been reading the room. Google's updated documentation for its Kubernetes Engine now increasingly points customers toward gVisor sandboxing when running untrusted code. In February, Trend Micro published a compliance rule formalizing the recommendation to use GKE Sandbox with gVisor. At KubeCon North America last fall, Google showcased an "Agent Sandbox" design built primarily on gVisor, a signal that isolation architecture is moving from operational footnote to first-class API concern.
Traditional VMs offered one answer—real isolation—but at a latency cost that often didn't pencil out. AWS Firecracker, the microVM technology behind Lambda and Fargate, advertises boot times of 125 milliseconds or less, using under 5 MiB of memory per VM. Impressive, certainly, for a full virtual machine. But still two orders of magnitude slower than isolate-based platforms.
Cloud Hypervisor, backed by Intel, Microsoft, Arm, and AMD, clocks in around 200 milliseconds according to a February analysis from Northflank. Kata Containers, which wraps Firecracker and other hypervisors into an OCI-compatible runtime, typically lands in the 150-300 millisecond range in community benchmarks.
That gap created a practical ceiling. If your serverless function runs for 50 milliseconds, burning 150 milliseconds just to boot the VM doesn't make economic sense. Developers drifted toward shared-kernel solutions and quietly accepted the security trade-off. You can almost hear the rationalization: probably fine, right?
The Technical Breakthrough—Or Breakthroughs

Three innovations converged to shatter the latency floor, and understanding them requires stepping briefly into the technical weeds. Memory forking via copy-on-write. Userfaultfd-based page management. Snapshot/restore optimizations that skip the boot sequence entirely.
Memory forking represents the conceptual leap. Instead of booting a fresh VM for every request, you maintain a "parent" VM in a known good state and fork it. The child inherits the parent's memory pages through copy-on-write semantics—physical memory only gets duplicated when the child actually writes to a page. Linux's userfaultfd mechanism makes this possible in userspace, with the kernel's UFFD-WP (write-protect) feature handling page-level tracking.
CodeSandbox, a developer sandbox platform, documented this evolution in real time. Their initial approach, described in a July 2022 blog post, cloned VMs using snapshots with copy-on-write disk images. Cloning took roughly two seconds. Snapshot writing consumed about a second per gigabyte of memory—workable, but not fast.
By September 2023, they'd rebuilt the system using memfd for VM RAM and userfaultfd to intercept page faults. The filesystem disappeared from the critical path entirely. Memory stayed shared until something wrote to it, at which point UFFD triggered a fault and the hypervisor copied only the affected page. The difference between copying gigabytes of disk versus copying individual memory pages on demand? Substantial, it turns out.
The kernel support matured in parallel. Throughout 2024 and 2025, Linux underwent ongoing UFFD-WP correctness fixes. CVE-2025-21696, disclosed and patched early this year, addressed a corner case in the write-protect logic. ARM64 enablement patches landed during the same window, broadening the architectures where userspace memory forking works reliably. These aren't headline-grabbing changes, but they represent the unglamorous engineering that makes breakthrough performance possible.
Snapshot/restore, the second pillar, sidesteps kernel boot and userspace initialization completely. Firecracker's documentation describes snapshot support that became generally available in recent releases, including on-demand memory loading with a UFFD handler. You restore directly to a pre-initialized state rather than waiting for the kernel to boot and init to run.
NumaVM published benchmarks in March showing Firecracker restore running 6.4 times faster than a full boot in their configuration. Academic work has pushed further—the Sabre project, presented at OSDI 2024, introduced hardware-assisted compression of microVM snapshots. A September 2025 preprint called "Spice" proposed OS co-design claiming near-warm cold-restores from disk, with the authors arguing that kernel limitations, not storage speed, had become the real bottleneck.
Security concerns accelerated adoption, perhaps more than the cloud vendors expected. The November container-escape advisories affected production infrastructure at scale, not just hobbyist deployments. Organizations running multi-tenant SaaS platforms began re-evaluating shared-kernel architectures with fresh urgency. MarketsandMarkets projected the confidential computing market—including trusted execution environments and confidential VMs—would grow from $5.3 billion in 2023 to $59.4 billion by 2028. Even discounting analyst optimism, the direction seems clear.
Who's Building What

Microsoft's Hyperlight represents the most aggressive bet on sub-millisecond VM execution. Open-sourced as a Rust library in late 2024, the project provides hardware-protected per-function micro-VMs running on KVM or Hyper-V. The February 11 blog post detailed how they hit that 0.9-millisecond warm execution time: a pre-warmed pool of micro-VMs with minimal guest state, serving HTTP requests with average latency under a millisecond.
Cold starts tell a different story. An earlier November 2024 Hyperlight post claimed 1-2 milliseconds for cold micro-VM spawn, and by late March, Microsoft had added WebAssembly support with similar performance targets. These are vendor-provided numbers, not independent benchmarks, but they establish a performance goal that was previously exclusive to isolate-based approaches.
AWS continues iterating on Firecracker, though the company remains characteristically tight-lipped about internal performance metrics. Public documentation still cites the 125-millisecond boot claim, but an internal re:Invent deck from 2019—surfaced by conference attendees—showed "microVM restore to snapshot point" in single-digit milliseconds during controlled demos. Those were ideal conditions, perhaps not representative of production workloads with real I/O patterns, but they indicate where the theoretical ceiling sits.
Fly.io has commercialized snapshot/restore under the "Machines" product name. Their documentation, updated last July, describes suspend/resume from snapshots taking "a few hundred milliseconds" compared to over two seconds for a cold boot. This January and February, Fly.io launched "Sprites," marketed as stateful Firecracker-based sandboxes with instant creation, checkpoint/restore, and object-storage persistence. Community coverage from mid-February highlighted the AI agent workflow use case: the ability to fork a running VM mid-execution to explore multiple decision branches, then commit the successful one.
Freestyle's product documentation describes exactly this pattern. Their VM forking feature, documented as stable in materials from March, allows branching a running workload "almost instantly" via memory copy-on-write. One VM becomes two with shared memory pages, diverging only as each executes its own path. Freestyle claims sub-second provisioning—under 800 milliseconds—from pre-booted snapshots.
CodeSandbox's evolution illustrates the engineering journey from "pretty fast" to "fast enough to rethink architecture." They went from two-second cloning with copy-on-write disks to memory-based CoW cloning that eliminated filesystem serialization. By August 2024, they were working on low-latency memory decompression to further scale snapshot/restore operations. Their use case is developer sandboxes rather than production serverless, but the techniques transfer directly.
E2B, open-sourced last May, targets AI code execution sandboxes built on Firecracker. The positioning is explicit: stronger isolation than containers for untrusted AI-generated code. Google's Agent Sandbox design, showcased at KubeCon, uses gVisor but acknowledges a pluggable isolation backend model where Kata Containers or Firecracker could slot in.
The pattern holds across platforms. As AI agents generate and execute more code—autonomously, without human review—security boundaries become load-bearing infrastructure rather than optional defense-in-depth.
Not every approach requires VMs, of course. Cloudflare Workers achieves sub-10-millisecond cold starts using V8 isolates, and an October 2025 InfoQ piece on their "Shard and Conquer" optimization reported 99.99% warm start rates. For many edge workloads, that's sufficient. But isolates share the V8 process and have limited syscall access. When the workload needs full Linux system calls or true kernel isolation, microVMs remain the only option.
What Comes Next

The technology roadmap is coming into focus, though the timeline remains uncertain. First, snapshot/restore with memory copy-on-write will likely become standard in serverless platforms and multi-tenant SaaS architectures. The kernel primitives—UFFD, UFFD-WP—are maturing, and multiple production systems have validated the approach. Independent benchmarking is needed, particularly for end-to-end request latency across I/O-heavy versus CPU-bound workloads, but the direction feels set.
Second, the distinction between "VM-based" and "isolate-based" serverless will blur, maybe faster than the industry expects. Hyperlight's demonstration that micro-VMs can match isolate latencies in warm scenarios removes the performance justification for weaker isolation. Cold paths remain slower for VMs, but techniques like Spice's OS co-design and Sabre's hardware-assisted snapshot compression suggest that gap will continue closing.
Third, AI agent sandboxes represent a major demand driver—perhaps the major one. The ability to fork a running VM state, explore multiple execution branches, checkpoint progress, and restore from failures aligns perfectly with agentic workflows. Fly.io's Sprites and Freestyle's VM forking are early commercial bets on this pattern. Expect API standardization here. Google's Agent Sandbox hints at a pluggable model where the isolation backend becomes a configuration detail rather than a platform lock-in.
Fourth, confidential computing will merge with this stack. Market projections vary widely, but all point to double-digit annual growth. AMD SEV-SNP and Intel TDX provide hardware-level memory encryption and attestation. Combining that with sub-millisecond VM startup creates a platform for privacy-sensitive workloads that was previously impractical. A medical AI agent analyzing patient data in an encrypted VM that boots in under two milliseconds and forks instantly for parallel analysis? Technically feasible now, where it wasn't eighteen months ago.
The kernel remains a bottleneck worth watching. Monitor Linux releases for UFFD-WP changes and broader architecture support. Correctness fixes like CVE-2025-21696 indicate the feature is moving from experimental to production-ready, but edge cases remain. They always do.
Hardware acceleration will matter, probably more than software optimization. Sabre's snapshot compression at OSDI 2024 demonstrated that offloading snapshot operations to hardware can cut latency meaningfully. Future CPU generations may include dedicated snapshot/restore instructions or memory-forking accelerators, similar to how AES-NI transformed encryption performance from "expensive" to "basically free."
Developers should prepare for a world where every serverless function could run in its own hardware-isolated VM with minimal latency overhead. That changes the security calculus for multi-tenant platforms. It also changes cost structures, since memory overhead per VM matters when you're spinning up thousands per second. Firecracker's under-5-MiB footprint sets a baseline, but tighter memory management will become a competitive differentiator.
The race to sub-millisecond VMs isn't really about bragging rights, though there's certainly some of that. It's about removing the last technical barrier to secure-by-default serverless computing. Microsoft hit 0.9 milliseconds in a warm pool. AWS demonstrates single-digit millisecond restores in controlled settings. The pieces exist. What remains is packaging them into platforms that just work—where isolation is the default and performance is no longer the excuse for cutting corners.
That's the finish line everyone's running toward. Whether it arrives in 2026 or 2028, the direction is no longer in doubt.
