A two-person startup out of Y Combinator claims to adapt industrial robots using demonstration videos filmed on a smartphone, processing the footage as context without retraining.
Shiraz AI, based in San Francisco, says it is deploying what it calls "robot in-context learning" at manufacturing facilities where production lines shift constantly. The approach skips the data collection, on-site engineering visits, and overnight retraining cycles that typically accompany robotic automation. Instead, operators record a demonstration video on their phones, and the robot attempts to replicate the task using a foundation model that processes the footage at execution time.
"A new task shows up on the floor? You just film it once on your phone and the robot picks it up," CEO Navid Aghasadeghi explained in the company's Y Combinator introduction. "No sitting around collecting data. No engineer flying out to your site. No overnight fine-tuning. You show it, and the robot learns."
It's an ambitious claim in a field where real-world deployment challenges have proven persistent. But Aghasadeghi and co-founder Mojtaba Mozaffar bring pedigrees from Boston Dynamics and Amazon Robotics, respectively, and they've started landing early customer deployments.
The company targets high-mix and contract manufacturers where rigid automation falls apart each time a new part arrives. Traditional robotic systems demand extensive programming or task-specific training data, making them impractical for factories running dozens of product variations through the same workstation. Shiraz AI aims to collapse those changeover costs by treating demonstration videos as runtime input rather than triggering model retraining.
Scope and limitations
Shiraz AI has been careful to define boundaries. The technology handles new variations within task families the model has already learned, not entirely novel physical skills from a single example. Demo videos show pick-and-place operations on packages with different shapes and orientations. That's a narrower promise than some headlines suggest, though still potentially useful for the targeted manufacturing environments.
The company describes its approach as feeding demonstration footage directly into the foundation model during task execution. No specialized gloves, no lengthy data collection runs, no weight updates for supported variations. Whether this holds up under the messy realities of factory floors remains to be seen, but the architectural choice sidesteps one of robotics' persistent headaches: the gap between lab performance and real-world deployment.
The team behind the pitch

Aghasadeghi spent years at Boston Dynamics, contributing to the commercialization of Spot and leading manipulation projects. Earlier, he worked at the now-defunct Rethink Robotics and earned a PhD in electrical and computer engineering from the University of Illinois Urbana-Champaign. A profile from several years back confirmed his time at Rethink, which tried and failed to crack the collaborative robot market before shutting down in 2018.
Mozaffar, the CTO, ran whole-body control and bimanual manipulation efforts at Amazon Robotics. He also held a research faculty position at Northwestern University, where he completed a PhD in mechanical engineering and robotics. Both founders announced their Y Combinator participation on LinkedIn, though details about funding amounts and customer names remain undisclosed.
A crowded field
Shiraz AI enters a robotics landscape already thick with startups chasing adaptable automation. Skild AI has touted its S1 foundation model, claiming strong performance on long-horizon tasks from single video demonstrations and noting deployments with Foxconn. Universal Robots and Scale AI launched an imitation-learning system earlier in the year. Intrinsic, the Alphabet-backed robotics unit, ships Flowstate for machine-tending applications. Companies like Augmentus, Wandelbots, and Rapid Robotics offer faster programming tools or template-based setups for similar use cases.
Academic labs have accelerated work in this direction as well. A paper titled "HOST: Robots Acquire Manipulation Skills in Seconds from a Single Human Video" appeared on arXiv in mid-2026, and Rhoda AI published research on its "Direct Video-Action" model around the same time. The underlying idea is gaining traction: can we teach robots by showing instead of coding?
What's next

The company says robots are operational at customer sites but won't name them yet. Y Combinator's standard deal typically provides $500,000 total, though Shiraz AI hasn't disclosed its specific terms. The startup lists a team size of two and is hiring roboticists in San Francisco.
Whether single-demo learning scales beyond controlled scenarios into the chaos of real manufacturing remains the open question. Shiraz AI has the technical backgrounds and early traction to make a serious run at it. But the graveyard of robotics startups that oversold flexibility is large, and factory operators have heard bold pitches before. The phones are out, the videos are rolling. Now comes the harder part: proving it works when the production line is running and the pressure is on.
