Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechJune 23, 2026

AI Scientists Compress Decade of Chip Materials R&D Into Months

AI Scientists Compress Decade of Chip Materials R&D Into Months
YcAi Agents+3
SaaS iconSaaSJune 23, 2026

Verified AI Training Data: Can Formal Methods Cure Hallucinations?

Verified AI Training Data: Can Formal Methods Cure Hallucinations?
Training DataLarge Language Models+2

Founders Mentioned

Harsha Vardhan Khurdula

Interfaze

saas icon
SaaS

Harsha Vardhan Khurdula

Interfaze

saas icon
SaaS
SaaS iconSaaS
June 23, 2026
YcDocument IntelligenceComputer VisionEnterprise AiB2b Saas

Interfaze Merges Specialized Models With Transformers for Deterministic AI

YC-backed startup's hybrid architecture targets reliability gap in document AI, offering OCR and structured outputs that prioritize consistency over flexibility.

Interfaze Merges Specialized Models With Transformers for Deterministic AI

The problem reveals itself at scale, usually around invoice number 3,847 or somewhere in the middle of a mortgage pipeline churning through 2,000 packets a day. The same AI model that breezed through test documents starts returning different results on retry. An OCR system confident on Tuesday turns tentative by Thursday, or worse, confident but wrong. JSON validates perfectly but populates with the wrong vendor name, the wrong total, the wrong borrower.

For developers building automation into the guts of enterprise workflows—insurance claims, medical records, the endless churn of paperwork that keeps the economy moving—this isn't academic. It breaks things.

A five-person team based in San Francisco thinks the industry has been solving the wrong problem. Interfaze, which went through Y Combinator's Spring 2026 cohort according to the YC directory, argues the issue isn't model quality. It's architecture. Their thesis: deterministic tasks demand deterministic design, which means ditching general-purpose foundation models for something more surgical. Specialized CNN and DNN encoders feeding directly into a transformer decoder, with "partial activation" that fires only what a given task requires.

Co-founder Harsha Vardhan Khurdula describes it as "tackling determinism for developer-facing tasks" in a LinkedIn post from May announcing the YC acceptance. Whether that actually works at scale is still an open question, but the problem they're chasing is real enough.

The Consistency Tax

Here's the thing about foundation models: they optimize for flexibility. OpenAI, Anthropic, Google—they're all building systems meant to handle thousands of different use cases. That breadth comes with variance baked in.

Interfaze's architecture paper, posted to arXiv in February, proposes something different. OCR gets a vision encoder. Speech-to-text gets an audio encoder. Web scraping gets a custom browser engine. All of it routes into a shared transformer with task-specific adapters. The model activates only the path it needs for the job at hand, returning structured output alongside what the team calls "precontext"—bounding boxes, confidence scores, the metadata breadcrumbs that let developers actually debug what went wrong.

Structured output modes have become table stakes in the past year. OpenAI introduced server-side constrained decoding in August 2024, guaranteeing that outputs would validate against JSON schemas. Anthropic and Google followed suit. But there's a difference between "valid" and "correct," and that distinction matters when you're processing a few thousand documents before lunch.

A paper from this past May, titled "Constraint Tax," quantifies the tradeoff. Hard schema enforcement can degrade value-level accuracy, particularly in smaller models. Another study from February explored the latency costs of draft-conditioned constrained decoding. The research keeps piling up, all pointing to the same uncomfortable truth: forcing models into structured formats has consequences.

So Interfaze built their own benchmark. The Structured Output Benchmark, or SOB (yes, really), which they released April 28, measures value accuracy across text, image, and audio inputs. Did the model extract the right invoice total? Did it get the vendor name? Not just: did it return valid JSON? The paper appeared on arXiv the same day the company announced it on their blog, which is either admirably transparent or a reminder to treat vendor-generated benchmarks with appropriate skepticism.

According to the company's blog, their leaderboards compare Interfaze Beta—which launched May 11—against Gemini-3-Flash, Claude-Sonnet-4.6, GPT-5.4-Mini, and others across nine tasks. The site reports an OCRBench v2 score of 70.7, olmOCR at 85.7, VoxPopuli word error rate around 2.4 percent, and SOB value accuracy of 80.5. A third-party tracker called BenchLM lists the model as "current" but noted in June that there's "insufficient non-generated benchmark coverage to rank safely."

Translation: the benchmarks are theirs to prove.

Hybrid Architectures, Assembled Differently

The architecture itself is more interesting than the scores, at least for now. Native fusion, as described in the February paper: task-specific encoders pass representations into a shared embedding space, then through adapters into the decoder. For a fixed-schema OCR task, the system activates the vision path, processes the document, and returns both extracted text and bounding-box metadata in a single API call. Developers access it through an OpenAI-compatible endpoint—$1.50 per million input tokens, $3.50 for output—with a default rate limit of 50 requests per second.

Interfaze isn't alone in exploring hybrid designs, though their particular implementation differs. A May arXiv paper titled "Operationalizing Document AI" detailed separating GPU inference from CPU orchestration with asynchronous I/O. An October 2025 study demonstrated a hybrid OCR-LLM framework achieving sub-one-second latency on structured documents. Another November paper explored hybrid detection and generation for document layout analysis.

The pattern emerging across the industry: specialized components for deterministic subtasks, orchestrated by or alongside a language model. The question is how to compose them.

AWS case studies from this spring show the approach in production. Ricoh built what they're calling a GenAI IDP Accelerator, combining Textract, Bedrock foundation models, and serverless orchestration. The company claims customer onboarding dropped from weeks to days with a projected seven-fold capacity increase. Rocket Close reported a fifteen-fold processing speedup for mortgage documents—roughly 2,000 files daily, 75 pages each—using the same stack.

Google Cloud's case studies follow a similar playbook. Fifth Dimension runs what they describe as a "multi-model" strategy on Vertex AI, mixing Gemini and Claude. Staple AI advertises "near-100% accuracy" across 300-plus languages using Gemini Flash, Vision, and Translation APIs, though such claims typically deserve the usual skepticism reserved for marketing materials.

The incumbents, meanwhile, haven't been sitting still. AWS updated Textract's API documentation as recently as June 12, adding features for rotated text and subscripts. Azure Document Intelligence shipped updates in March including new skills and versioning adjustments. Google expanded Workspace and Vertex AI integrations through the spring. ABBYY launched Vantage 3.0 with GenAI integration in January. Mistral entered the space with OCR 3 in December 2025.

And then there's the open-source toolchains—Docling, Marker, Surya OCR—all widely used as PDF-to-Markdown preprocessors feeding into LLM post-processing. The components exist. How to compose them remains the question.

A Market in Transition, Maybe

Digital illustration for article section "A Market in Transition, Maybe" in "Interfaze Merges Specialized Models With Transformers for Deterministic AI" - A conceptual, minimalist representation of a market in transition, featuring a clean, abstract arran...

Industry analyst reports published this spring paint a picture of a market that hasn't quite figured itself out yet. The Everest Group's 2026 IDP PEAK Matrix, released across March and April, analyzed 32 providers. Forrester's Q2 Wave on Document Mining and Analytics Platforms followed in May. A Forrester blog post accompanying the Wave emphasized "use-case alignment over vendor logo" and noted the market's fragmentation—which is analyst-speak for "nobody's winning decisively yet."

Market sizing estimates vary wildly depending on methodology. Grand View Research projects the global intelligent document processing market will reach $12.35 billion by 2030, implying a 33.1 percent compound annual growth rate. Mordor Intelligence estimated $3.17 billion in 2026 with a more conservative 17.78 percent CAGR through 2031. 360iResearch offered the most modest forecast: $2.56 billion climbing to $5.26 billion by 2032, a 10.81 percent CAGR.

The discrepancies reflect scope differences more than consensus, but the directional bet is clear enough: enterprises are spending more on this.

Enterprise surveys suggest demand is real even if the exact numbers remain squishy. Deloitte's State of AI in the Enterprise survey, based on fieldwork from late 2025, found companies with at least 40 percent of AI projects in production were positioned to double that proportion within six months. Productivity and efficiency led realized benefits at 66 percent. PwC's Digital Trends in Operations survey reported that more than half of health services executives see faster processing and higher decision consistency from AI-assisted workflows; 94 percent of tech and telecom firms have implemented AI to some degree, with 40 percent scaling enterprise-wide.

An Adobe study from April found adoption surging across document-heavy functions, though legal reported the lowest uptake—make of that what you will—and governance gaps persist. Dun & Bradstreet's survey indicated 56 percent of respondents plan to increase AI investment over the next year, describing an inflection point from experimentation to ROI.

The workflows most commonly deployed in recent rollouts include the usual suspects: invoice and accounts payable processing, KYC and identity capture, mortgage and title packets, healthcare records, insurance claims. The architecture pattern emerging: move from single-model OCR to hybrid systems—OCR plus language model, or fused encoders—with schema validation and human-in-the-loop exception handling for the cases that don't fit the pattern.

Observability matters now in ways it didn't two years ago. Confidence scoring, bounding-box provenance, value-level accuracy benchmarks like SOB—these reflect production needs the early demos conveniently ignored.

Regulation, Scraping, and Other Complications

Deployment is getting more complicated legally. The EU AI Act's general applicability date arrives August 2, with obligations for general-purpose AI model providers already in force since last August. High-risk application requirements may shift under the Digital Omnibus proposal working through Brussels, but procurement teams are already citing ISO/IEC 42001 and the NIST AI Risk Management Framework in RFPs for regulated-sector use cases. The U.S. framework remains voluntary, though NIST is developing sector profiles for critical infrastructure through this year.

Then there's web scraping, which remains legally nuanced at best. The LinkedIn versus hiQ litigation concluded with contract and trespass claims prevailing, even though the original CFAA theory didn't hold. Public site scraping sits in a gray area defined by terms of service, anti-circumvention laws, and a patchwork of regional rulings. Deployers need careful legal review before scraping at scale—a reality that won't change just because the technology makes it easy.

Interfaze positions its custom web engine as part of the native architecture, returning metadata alongside scraped content through the same API. Whether that design sidesteps legal pitfalls depends entirely on how customers use it and which sites they target. The technology doesn't settle the compliance question; it just moves it downstream.

Building in Public, Sort Of

The team rebranded from JigsawStack on March 26, a month before announcing YC acceptance. The architecture paper lists three authors: Harsha Vardhan Khurdula, Vineet Agarwal, and Yoeven D Khemlani. The company spoke at the Agents & APIs SF Developer Meetup in late May, sharing a lineup with Firebase, Public.com, and Postman. The YC directory lists the team size as five, with Aaron Epstein as the primary partner. No funding amount has been disclosed beyond the standard YC investment.

Blog posts from late May track like a team iterating in public. A May 28 post announced OCR improvements, Gemini 3.5 benchmarks, speech-to-text word-level accuracy updates, and caching features. A May 27 post asked, "Does the ability to refuse make a model more intelligent?"—a positioning play on guardrails. The May 11 launch post detailed the architecture, benchmarks, partial model activation, and SDK instructions.

It reads like founders releasing features and refining their framing as they go, which is probably exactly what's happening.

The Road Gets Longer From Here

Digital illustration for article section "The Road Gets Longer From Here" in "Interfaze Merges Specialized Models With Transformers for Deterministic AI" - A sleek, continuous, minimalist path stretching forward into a soft, uncluttered horizon, symbolizin...

Foundation-model providers will keep advancing structured-output reliability and long-context PDF handling. OpenAI's 2024 structured outputs established a baseline; research from this year shows active optimization, though correctness gaps remain stubborn. Specialized vendors differentiate with metadata provenance, validation layers, and domain-specific encoders. Interfaze's fusion approach represents one answer. AWS GenAI IDP Accelerator case studies point to another. Open-source toolchains offer a third path for teams willing to assemble components themselves.

Analysts expect continued consolidation in IDP platforms and more embedded "agentic" capabilities—classification, splitting, extraction, validation orchestration instead of pure OCR. Forrester's Q2 Wave blog emphasized the breadth and fragmentation of the current market, suggesting winners will align tightly with specific use cases rather than trying to serve everyone. Which is perhaps another way of saying the market hasn't matured yet.

The determinism problem Interfaze targets isn't going away. Production workflows demand consistency. General-purpose models optimize for breadth. The gap between demo-quality outputs and production-grade reliability remains the central engineering challenge for document AI right now, in 2026 and likely beyond.

Interfaze's bet is that architecture matters more than scale, that fusing specialized encoders into a transformer decoder can deliver the consistency that general-purpose models can't promise. The benchmarks are theirs to defend, and third-party validation will take time. But the problem they're solving—the one that reveals itself around invoice 3,847—is real enough that enterprises keep writing checks to make it go away.

Whether a five-person team from San Francisco has cracked it remains to be seen. Then again, that's always been the bet with startups coming out of Y Combinator, hasn't it?

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • AI Scientists Compress Decade of Chip Materials R&D Into Months
  • Verified AI Training Data: Can Formal Methods Cure Hallucinations?
  • Lamina Labs Bets on Deterministic Video as Sora Exits the Market
  • AI Voice Agents Come to Legacy Call Centers Without the Rip-and-Replace
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.