Hebbian Robotics released HFlow on August 19, 2026, an open-source SDK designed to help robotics teams scale data quality pipelines from thousands to millions of hours without building custom infrastructure. The Y Combinator-backed startup published the toolkit on GitHub under an Apache 2.0 license, shipping version 0.2.0 to PyPI the following day.
The release addresses what has become a persistent bottleneck in robotics development: teams now collect sensor data faster than they can process it. "We built HFlow because robotics teams are collecting more hours of video data than they can handle," Kingston Kuan, co-founder, wrote in a Product Hunt post on August 26. Brandon Ong, the company's CEO, framed the problem more bluntly. HFlow, he said, is for teams "that have the ambition to process a million hours of Physical AI data and are starting today."
That ambition isn't purely theoretical. Hebbian points to an August engineering writeup from Dyna Robotics that reported a throughput jump from 1,400 episode-hours per week to 44,000—a scale that suggests data volume has begun to outpace available tooling across the industry.
How It Works
HFlow treats MCAP files, the default ROS 2 bag format since the Iron release, as canonical "episodes." It runs quality checks without requiring teams to train a model first. Engineers write Python functions to transform and inspect data, and HFlow generates Apache Airflow 3 DAGs to orchestrate the pipeline behind the scenes.
The architecture separates evidence from verdicts. Checks record measurements, intervals, and tags; curation applies pass/fail policies later. That separation matters. A Voxel51 audit published two weeks before HFlow's launch showed that smoothness metrics scored defective robot demos with an AUROC of 0.244—worse than random chance. Decoupling detection from judgment, in other words, may prevent teams from accidentally filtering out the data they need.
The pipeline unfolds in stages. Collection lands raw MCAP files. Ingestion runs user-defined transforms and toggleable gates. Curation writes a Parquet catalog queryable via DuckDB. HFlow implements H.264 GOP-length presets aligned to read patterns and stamps provenance metadata, including schema version and robot software version when available, according to the project's architecture documentation on GitHub.
A local Docker Compose runtime ships with the open-source release, though teams can plug in their own Airflow 3 cluster if they prefer. For development, an app.test() method runs pipelines in-process without Docker.
The SDK requires Python 3.11 or later. It includes optional extras for S3, Google Cloud Storage, and Azure blob backends. Hebbian pushed three PyPI updates in the first week: 0.2.0 on August 20, 0.2.1 on August 25, and 0.2.2 on August 26.
The Founders

Brandon Ong and Kingston Kuan founded Hebbian Robotics as a two-person team. Kuan previously built infrastructure at Jane Street and Verkada, according to the startup's Y Combinator profile. The company participated in a recent Y Combinator batch and joined NVIDIA Inception last month, Ong said in a LinkedIn post. Based in San Francisco, the startup describes itself as a public benefit corporation.
"After many conversations with data teams, we decided the path to building the best version of HFlow is open source," Ong wrote in the August 19 blog post announcing the release.
Competitive Landscape
Hebbian enters a market with both horizontal data platforms and robotics-specific tooling already in play. Foxglove markets what it calls an "agentic data platform for Physical AI" with cloud storage and MCAP support. Rerun ships open-source multimodal visualization that reads MCAP files. Scale AI's Nucleus offers data curation with embeddings search and a "Smart Sample" feature, though it isn't robotics-specific. Roboflow manages computer vision datasets but doesn't focus on episode-centric MCAP pipelines.
HFlow's architecture document cites design decisions from Dyna's million-hour pipeline, including storage gains versus per-frame JPEG and read performance from topic-group chunking. The open-source release marks features as "implemented," "simplified," "deferred," or "out of scope" relative to production-scale architecture. Batch scheduling uses first-fit-decreasing. Arbitrary-step replay is deferred. Distributed corpus cache sits out of scope for now.
The Open Source Robotics Alliance recently launched a Physical AI Special Interest Group to standardize messages and data-management pipelines in ROS. NVIDIA released what it described as "a major collection of open-source agent tools and skills for Physical AI" over the summer, according to company press releases.
What's Next
The project is marked pre-v1, with "core lifecycle working end to end," per the GitHub README. Hebbian also maintains Pareto, a Rust tool for curating LeRobot datasets with clustering, duplicate detection, and text search. The architecture documentation mentions a possible hosted control plane but notes it's "not a pre-v1 release commitment."
Teams can install HFlow with uv add hflow and run a quickstart script included in the repository. Whether the toolkit will gain traction beyond early adopters depends, perhaps more than anything, on whether Hebbian can sustain the kind of rapid iteration its first-week release cadence suggests.
