The robots keep getting smarter, but they're still tethered to the cloud.
That tension—between increasingly capable AI models and the real-world constraints of running them on hardware that moves, breaks connectivity, or operates in places where sending data back to some data center simply isn't an option—is what Bill Jiao and Guanming Wang decided to solve. Their two-person startup, General Instinct, emerged from Y Combinator with a singular pitch: take frontier AI models and make them run offline, on resource-starved devices, without compromise.
Around May 19, they launched Instinct Edge. Not a model. Not a robot. A deployment layer.
The promise is deceptively simple. Hand the platform a large vision-language model, specify your target hardware—a drone's onboard computer, maybe, or a Jetson module bolted to a factory robot—and set a latency budget. Instinct Edge compresses, optimizes, and packages everything into a binary that fits and runs without ever phoning home.
"Give us your model and hardware specs," the founders wrote in their Launch YC post around May 19. "We'll give you a runtime that works."
When the Cloud Isn't an Option
Most cutting-edge AI assumes you have bandwidth to spare and compute on tap. Edge devices don't live in that world.
A robot navigating a warehouse can't afford the 200-millisecond round-trip to query GPT-4's API every time it needs to recognize an object. An agricultural drone surveying crops in rural farmland might not have connectivity at all. Industrial inspection systems in oil refineries or power plants often face air-gapped security requirements that prohibit cloud calls entirely.
Then there's cost. Sending inference requests to cloud APIs at production scale gets expensive quickly, especially when you're processing video streams or running decisions dozens of times per second.
General Instinct's solution involves the unglamorous, grinding work of model compression—quantization, distillation, and hardware-specific optimization. They take models built for data centers and squeeze them down until they fit onto chips like NVIDIA's Jetson modules, Qualcomm Snapdragon processors, or Apple's Neural Engine. Custom CUDA kernels for NVIDIA hardware. Metal shaders for Apple silicon. ARM NEON instructions for mobile chips.
According to their early documentation, one production deployment runs a multimodal classifier on a Jetson Orin NX with a 111-millisecond cold start. Every inference stays under 150 milliseconds. Zero cloud dependency.
That might not sound transformative until you've tried to make a robot react in real time.
The Founders: Patents on Tardigrades, Time at DeepMind

Jiao and Wang are running this out of San Francisco as a true two-person operation, which raises obvious questions about scalability. But their backgrounds suggest they've at least thought through the technical depth required.
Jiao previously worked on Siemens' first multimodal foundation model and—somewhat improbably—holds two patents on tardigrade proteins dating back to his high school years. (The connection to edge AI deployment is unclear, but it signals a certain kind of technical ambition early on.)
Wang spent time at DeepMind and has cultivated a 30,000-follower audience on RedNote, the Chinese social platform. That presence may prove useful if General Instinct pursues international robotics markets where hardware vendors and integrators rely on Chinese platforms for developer engagement.
Before Instinct Edge, the team shipped an earlier version called Instinct v0 around March, billed as "the first agent that sees the physical world." That iteration focused on hybrid inference—running partly on-device, partly in the cloud depending on connectivity. Instinct Edge represents a sharper, more committed bet: fully offline, with guaranteed latency characteristics baked in.
A Crowded Moment for Edge AI

They're not alone in seeing this opportunity. The first quarter of 2026 has brought a wave of similar infrastructure plays, perhaps driven by growing skepticism about cloud dependency and rising awareness of edge AI's practical constraints.
NVIDIA released TensorRT Edge-LLM in January, an open-source framework targeting its Jetson and automotive platforms. RunAnywhere, another YC company from the Winter batch, launched its own production-grade on-device AI platform in March. Meta's ExecuTorch framework has quietly been running inference across the company's mobile apps for months.
At CES in January, Qualcomm and Edge Impulse announced an on-premises appliance designed explicitly for air-gapped deployments—no internet, no cloud, no exceptions. Legion Intelligence introduced Centurion, aimed at military and defense applications where connectivity is unreliable by design. SoundHound demoed a "completely on the edge" multimodal assistant for vehicles at NVIDIA's GTC conference in March, avoiding cloud calls even for conversational AI.
The common thread: real-world AI applications increasingly demand local execution, either for latency, privacy, cost, or reliability reasons. Sometimes all four at once.
General Instinct's angle is infrastructure for builders. Their Launch YC post reads like an invitation to a specific kind of engineer—teams "running large vision models on edge and having problems fitting onto hardware" can reach [email protected]. It's not a plug-and-play consumer product. It's a tool for robotics companies, IoT developers, and hardware teams already deep in the weeds.
The Two-Person Question

Whether Jiao and Wang can support the level of customization that different hardware platforms, use cases, and performance requirements demand is an open question. Edge optimization isn't one-size-fits-all. A drone running on a Snapdragon chip needs different kernel optimizations than a factory robot on Jetson. Automotive applications have different thermal and power constraints than mobile devices.
The 111-millisecond benchmark suggests they've cracked at least part of the puzzle. But translating that into a platform that works reliably across dozens of hardware configurations, with customers who need hand-holding through deployment, is a different challenge entirely.
Still, maybe that's the point. YC companies often start narrow—solve one hard problem exceptionally well, then figure out how to scale the solution. General Instinct seems to be betting that getting edge AI deployment right is hard enough, and valuable enough, that starting small makes sense.
For now, the founders are testing their thesis in the most practical way possible: engaging directly in niche Reddit communities like r/JetsonNano and r/computervision, asking developers what breaks when they try to deploy large models on constrained hardware. The feedback loop is tight. The problems are specific.
And if the robots are going to get truly autonomous, someone has to make sure they can think without asking the cloud for permission first.
