Concentrix had a problem. Their invoice processing system, built on Microsoft's document intelligence stack, was hitting 96% extraction accuracy—respectable by industry standards, impressive even. But in the world of enterprise document automation, that remaining 4% represented thousands of invoices requiring human review, bottlenecks in payment cycles, and the kind of operational friction that CFOs notice.
By January this year, they'd clawed their way to 99%. Those three percentage points required what one engineer described as "a fully orchestrated pipeline"—traditional optical character recognition engines feeding structured data to large language models, with human oversight filling the gaps that neither system could bridge alone. Painstaking work. The kind that defines enterprise document AI today, where "good enough" rarely is.
Now a small startup called Interfaze, fresh out of Y Combinator's spring cohort, is arguing that the entire industry has been playing the game wrong.
According to Interfaze's stated approach, pure transformer models—the architectural foundation behind ChatGPT, Claude, and every other large language model currently capturing headlines—are fundamentally mismatched for deterministic extraction tasks. Instead of forcing LLMs to handle every step of the process, Interfaze claims to have fused specialized convolutional neural networks with transformers in a single model architecture. Let each component handle what it does best, the thinking goes. Stop asking transformers to be everything.
It's a technical argument with significant commercial implications. The intelligent document processing market varies wildly depending on whose projections you believe—somewhere between $2.8 billion and $3.9 billion projected for 2026, growing at compound annual rates that range from a modest 10.8% to an explosive 33.8%. Sprawling market, divergent methodologies. But the underlying demand signal? Unmistakable.
When Perfect Formatting Isn't Enough
OpenAI introduced structured outputs last August, promising 100% schema adherence in their evaluations. A month doesn't go by now without one of the major AI providers announcing improvements to JSON generation, schema compliance, or format reliability. Anthropic made Claude's structured outputs generally available earlier this year. Google enhanced Gemini's JSON Schema support around the same period. Every major player treats structured outputs as table stakes.
OpenAI's evaluations show 100% schema match—but that doesn't guarantee value accuracy. At least in theory.
Except—and this is where things get interesting—schema adherence isn't the same as value accuracy. A model can return a perfectly formatted JSON object with completely wrong numbers inside. For a know-your-customer pipeline or financial reconciliation workflow, that's not just an inconvenience. It's catastrophic.
The Structured Output Benchmark, released in late April, illustrated the gap with uncomfortable clarity. Across text, image, and audio inputs, tested models showed 83.0%, 67.2%, and 23.7% value accuracy respectively, even when the JSON structure came back perfectly formed. The wrapper looks right. The contents are garbage.
Field reports from production deployments catalog the failure modes in exhaustive detail: refusals when confidence drops below some invisible threshold, partial extractions that quietly omit fields, nested objects that hallucinate structure where none exists in the source document. An analysis published by Codexical in May documented these patterns, noting that constrained generation solves format but not correctness. Developers need both. They're discovering that transformers alone deliver neither reliably.
The document intelligence sector has responded with layered architectures—systems that route tasks through multiple specialized models before synthesizing results. UiPath's Document Understanding platform, updated continuously through spring, bridges traditional ML-based OCR with LLM reasoning. Microsoft announced Azure Content Understanding at Build, routing tasks through specialized models. ABBYY's Vantage 3.0, launched in January, threads generative AI capabilities into governance-focused workflows but maintains separate extraction engines. Even AWS customers like Ricoh are building multi-tenant IDP solutions that orchestrate Textract alongside Bedrock models.
These are all hybrid systems, cobbled together at the integration layer. What Interfaze claims to have done—and the claim remains largely untested in production—is merge the hybridization into the model itself.
Architecture as Ideology
Interfaze wasn't always Interfaze. Founded last year as JigsawStack, the company rebranded after getting into Y Combinator's spring batch. Five people, according to the accelerator's directory. Several months out from public beta. But already making an architectural statement that feels almost ideological in its conviction.
The team describes their approach as "model-level fusion," detailed in a paper accepted at IEEE CAI earlier this year. The architecture stacks specialized perception modules—OCR engines, object detectors, speech recognition systems, classifiers—beneath a transformer-based controller. The small models do the extraction work. The LLM orchestrates context and outputs structured results with provenance metadata: bounding boxes, confidence scores, block-type classifications.
"Pure transformer models are the wrong tool for deterministic tasks," they wrote in an April blog post, a position they've repeated across product communications with the consistency of a campaign slogan. The architecture achieves determinism not because transformers got better at precision tasks, but because transformers aren't doing those tasks anymore.
Pricing sits at $1.50 per million input tokens and $3.50 per million output tokens—roughly flash-tier economics with a 1 million token context window and 32,000 token maximum output. The model handles text, images, audio, files, and video. According to Interfaze's public leaderboard, their beta model is benchmarked against Gemini 3 and 3.5 Flash, Claude Sonnet 4.6, GPT-5.4 Mini, and Grok 4.3 across nine tasks spanning OCR, spatial reasoning, speech recognition, and structured extraction.
Marketing materials reference "98-99% structured output accuracy" and sub-five-second response times for specialized tasks. How those numbers hold in production remains to be seen—perhaps more than the founders expected, given the team size and timeline. What matters more is the architectural statement itself: that document AI requires task-specific small models orchestrated by, not replaced by, large language models.
Convergent Evolution

If Interfaze represents one end of the spectrum—fusion at the model layer—other players are converging on similar principles from different angles. Sometimes that happens in technology. Different teams, working independently, arrive at adjacent solutions to the same underlying problem.
Mistral released OCR 4 in late June, positioning it as structure-aware document AI with bounding boxes, block-type classification, confidence scoring, and support for 170 languages. Media coverage cited strong benchmark performance. The emphasis in Mistral's communications was deployment flexibility: self-hosting or API, integrated into a broader Document AI toolkit alongside search infrastructure.
Reducto took yet another approach with its Deep Extract agent, announced in early April. Rather than change the model architecture, they implemented iterative verify-and-reextract loops—an agentic pattern designed to close the value accuracy gap through multiple LLM passes. Slower, more expensive, but acknowledging the same underlying problem: single-pass transformer inference isn't reliable enough for high-stakes extraction.
IBM's Docling stack focuses upstream on document conversion. The Granite-Docling models transform messy PDFs into richly structured formats—DocTags, purpose-built for downstream retrieval and grounding—but stop short of final extraction. It's preprocessing, not inference. Still, it signals a broader consensus that document AI pipelines need multiple specialized stages rather than one model trying to do everything.
Traditional IDP vendors are retrofitting architectures too. Hyperscience, positioned as a Leader in Forrester's second-quarter Document Mining and Analytics Wave, released "Hypercell" this spring—intelligent routing that sends tasks to the right model based on document type and required output. UiPath's Document Understanding platform saw monthly updates through spring. Microsoft's Azure Document Intelligence integrated with Foundry at Build. AWS customers continue building on Textract, which saw documentation updates as recently as mid-June.
The open-source ecosystem matters here too. Chandra OCR 2 launched in March with benchmarks across 90-plus languages. Surya OCR 2 followed in May. PP-OCRv6 arrived in early June. These aren't one-off research releases; they're production-grade systems being deployed in enterprises that need on-premises processing or want to avoid cloud pricing. Licensing and throughput considerations vary, but the ecosystem is maturing rapidly. Faster than many anticipated.
Regulation Meets Reality
Timing, as they say, is everything. The EU AI Act's general obligations take effect in August. High-risk applications—and many document processing use cases fall squarely into this category—will face transparency, auditability, and human oversight requirements. Some sector-specific rules have been proposed for next year and beyond, but procurement teams are already asking vendors about compliance. Now, not later.
ISO/IEC 42001, the AI management systems standard, is gaining traction in enterprise RFPs. NIST's AI Risk Management Framework continues evolving, with recent concept notes addressing critical infrastructure profiles. Regulatory pressure creates demand for exactly what Interfaze's architecture promises: provenance, confidence scores, verifiable outputs.
You can't audit a black-box transformer decision, after all. But you can audit a CNN extraction with spatial coordinates and a confidence threshold. The architecture isn't just about accuracy; it's about building systems enterprises can actually govern. Systems they can defend in an audit.
At the same time, the broader AI infrastructure market is in hyperdrive. Gartner forecasted worldwide AI spending at $2.59 trillion this year, up 47% year-over-year in a report released in May. IDC's spring briefs describe an "AI supercycle" across infrastructure and semiconductors, driven by hyperscaler and sovereign AI buildouts. Document intelligence is a small slice of that spending. But it's a slice where accuracy failures have immediate, measurable costs.
WNS published a case study in June on underwriting document-to-decision transformation. Rocket Close deployed Textract and Bedrock for mortgage processing, documented in April. These aren't pilots anymore; these are production systems processing thousands of documents daily. When they fail, money is lost or compliance is breached. The market is mature enough now that "90% accurate" doesn't cut it.
That three-percentage-point gap? That's where the real money is.
What Happens Next

Expect tighter integration between document extraction and downstream retrieval or search systems. Mistral's pairing of OCR 4 with its Search Toolkit points toward a future where ingestion and retrieval share structure and metadata—where the extraction step produces not just text, but richly annotated documents that preserve spatial relationships and confidence levels for later reasoning steps.
Google's March updates embedded Gemini capabilities directly into Docs, Sheets, and Slides, normalizing ambient document intelligence in consumer workspaces. Enterprise versions of that ambient processing are coming. Or they're already here, quietly, in systems that don't announce themselves with press releases.
Tooling standardization via the Model Context Protocol should reduce the integration cost of agentic pipelines. OpenAI adopted MCP last spring; broader ecosystem adoption followed through the year. When tool calling and context protocols become universal, orchestrating heterogeneous model stacks—CNN for OCR, transformer for reasoning, classifier for validation—gets easier. Interfaze benefits from this trend, but so does every competitor trying to build governed, auditable pipelines.
The reliability gap between schema adherence and value accuracy will remain a focal point. Benchmarks released in April with thousands of test samples give teams a way to measure the problem. Frameworks proposed at PharmaSUG in March—agentic architectures with deterministic tool calls and constrained generation—offer design patterns for solving it. But solving it in production, at scale, across diverse document types and formats? Still an open problem.
Interfaze's bet is that architecture matters. That you can't prompt-engineer your way out of a fundamental mismatch between model strengths and task requirements. Whether their specific implementation wins market share is less important, perhaps, than the broader thesis: the document AI systems that succeed won't be pure LLMs. They'll be purpose-built hybrids that know when to transform and when to convolve, with provenance and confidence built in from the ground up.
That gap Concentrix closed between 96% and 99%—that's where the real work happens. And that's the gap pure transformers may never close on their own, no matter how many parameters you add or how much compute you throw at the problem. Sometimes the answer isn't a bigger model. Sometimes it's the right architecture for the job.
