The numbers alone explain why biotech executives lose sleep. Turning raw clinical trial data into an FDA submission package can take a pharmaceutical company anywhere from three to nine months—sometimes longer if the statisticians and programmers spot inconsistencies in the mapping specs.
It's a gauntlet by design. Biostatisticians spend weeks translating electronic data capture outputs into CDISC standards, the industry's lingua franca for regulatory submissions. Programmers generate analysis datasets and produce hundreds of tables cross-referenced to statistical analysis plans. Quality assurance teams validate every derivation, every footnote, every decimal place. Only then can the medical writers draft the documents regulators will actually read.
Astraea, a San Francisco startup that emerged from Y Combinator this spring with just two founders, thinks it can collapse that entire timeline into a matter of days.
The company launched its platform publicly in mid-May, pitching what it calls "the first end-to-end AI system" for clinical biometrics—the statistical programming and documentation workflows that sit between locked databases and submission-ready packages. If the technology works as advertised, it represents a meaningful acceleration in an industry where speed often determines which drugs reach patients first. If it doesn't, well, Astraea will join the long list of software vendors that underestimated how unforgiving pharmaceutical compliance can be.
The Automation Play
Astraea isn't a clinical data management system, and it's not medical writing software. The founders describe it as something closer to an execution layer—a platform that ingests study artifacts (protocols, case report forms, statistical analysis plans) and orchestrates the production of everything regulators expect to see.
That means CDISC-compliant SDTM and ADaM datasets. Tables, listings, and figures wired directly to the statistical analysis plan. Define-XML files, annotated case report forms, data reviewer's guides. The platform even runs Pinnacle 21 validation automatically, the industry-standard conformance check that catches formatting errors before the FDA does.
According to materials on the company's website, the system generates production-grade R and SAS code with "cell-level parity" to reference outputs—a technical detail that matters considerably if you're a biostatistician who knows how a single misaligned decimal can delay a submission by weeks. The platform also maintains data lineage and traceability logs designed to satisfy 21 CFR Part 11 requirements for electronic records, the compliance framework governing how pharmaceutical companies handle digital data.
Rather than relying on a single large language model, Astraea uses what it calls a multi-agent architecture. Four specialized "engines" handle different parts of the workflow: standards mapping, compliance orchestration, evidence extraction from study documents, and statistical output generation. Whether that architectural choice translates to more reliable outputs is an open question. The technology is too new for independent validation.
Who's Buying In

The company claims on its website to have landed a "Top-10 Pharma Partner," though it hasn't disclosed which one or provided independent verification. (Pharma companies, for their part, rarely announce experimental software deals until the tools have proven themselves in production.)
Astraea is initially targeting oncology and rare disease Phase II and III trials—therapeutic areas where timelines matter acutely and where smaller biotech firms often lack deep in-house biometrics expertise. The marketing claims are aggressive: 30–50% faster biometrics cycles, 95% faster reporting workflows, 99.8% validated outputs. Those figures come from the company's internal benchmarks rather than independent studies, a caveat worth noting. Still, they track with broader industry projections. McKinsey estimated in late 2025 that agentic AI could lift clinical development productivity by 35–45% over five years, a forecast that has circulated widely among pharma executives in 2026.
Astraea is hardly alone in chasing this opportunity. The competitive landscape for CDISC automation has filled out quickly. TrialNexus markets a seven-agent workflow covering SDTM mapping through Define-XML generation. Clymb Clinical announced its Data Mapper tool in late May, emphasizing AI-driven mapping specifications. CDISC Pro, KlinAI, Clinaform—they all promise some version of the same pitch: take manual, error-prone processes and make them push-button, traceable, fast.
What Astraea claims as differentiation is its scope. In the company's Y Combinator launch post from mid-May, CEO Joshua Wang wrote that Astraea is "the first to actually completely automate workflows across SDTM, ADaM, TFL generation, Pinnacle 21 checks, and specification creation." Competitors might contest that framing. But the ambition to handle the entire pipeline—protocol ingestion through submission packaging—represents a bet that pharma buyers want fewer tools in their stack, not more. Whether that's correct is another question.
The Founders

Joshua Wang and Sanmay Sarada both studied computer science and math at Stanford. Wang previously built multi-agent AI systems and worked on enterprise AI applications; Sarada focused on clinical and hospital data infrastructure in regulated environments. They founded Astraea earlier this year and joined Y Combinator's Spring batch, working with partner David Lieb.
Their core thesis is that the clinical development stack needs an execution layer that understands regulatory standards natively—not as an afterthought, but as the structural logic governing every step from data mapping forward. "Astraea is the bridge between protocol intent and regulator-ready evidence," the company's platform page explains.
That framing aligns, perhaps conveniently, with how the FDA has been talking about standardized data. The agency has consistently pushed sponsors to submit in SDTM and ADaM formats, and recent FDA commentary—including reports this April on real-time review pilots—suggests the regulator is exploring how AI might accelerate its own evaluation processes. Astraea's positioning, which emphasizes alignment with "FDA's human-centric, risk-based, standards-based AI posture," reads like a calculated nod to that regulatory moment.
The Risk Ahead

Astraea is operating on a request-demo model. No public pricing, no freemium tier. The website emphasizes enterprise compliance: HIPAA alignment, GDPR readiness, GxP validation, governed orchestration with role-based approvals. Those are table stakes for selling into pharma, but they also underscore the compliance risk the founders are assuming. If Astraea's outputs fail a Pinnacle 21 check—or worse, introduce mapping errors that surface during FDA review—the consequences extend well beyond customer churn.
The broader question is scalability. SDTM and ADaM workflows are relatively structured; that's why they're amenable to automation in the first place. But clinical trials generate messy, unstructured data: adverse event narratives, imaging reports, biomarker assays that resist tidy categorization. Astraea's evidence synthesis engine claims to extract key information from protocols and clinical study reports, but how far that capability extends into real-world messiness will determine whether the platform becomes a full biometrics replacement or a high-value assistant.
For now, the company is betting that pharma's operational pain is acute enough to justify taking a risk on a Y Combinator startup with a two-person team. The pitch is appealingly simple. What if you could go from database lock to FDA submission in days instead of months?
If Astraea can deliver on that—and if its outputs hold up under regulatory scrutiny—the total addressable market is every drug company trying to move faster. The regulatory pathway has always been slow by design, built on layers of verification meant to protect patients. The gamble Astraea is making is that the design itself can change.
