In February, a startup called Taalas did something audacious. It etched Meta's Llama 3.1 8B model directly into silicon—not as firmware, not as software, but as physical circuitry baked into the metal layers of a chip. The result: 17,000 tokens per second per user, roughly ten times faster than GPU baselines, according to the company. The catch? The chip could run only that one model. Nothing else.
Two months later, AMD bought Taalas. The acquisition signaled that hardwiring AI models into application-specific integrated circuits had moved from fringe experiment to something the industry might actually need.
Nvidia still holds about 70 percent of the AI chip market, according to TrendForce data from late last year. But beneath that dominance, a different kind of chip economy has been taking shape. The hyperscalers have already deployed custom silicon: Microsoft shipped its Maia 200 accelerator on TSMC's N3 process in January, Meta expanded its MTIA roadmap in March, and AWS began delivering Trainium3 chips early this year, according to an April SEC filing. Jon Peddie Research counted 151 companies and more than 290 AI processor products in a report published this past July.
What distinguishes the newest wave is its willingness to abandon flexibility. Startups are etching entire models into mask layers, fixing compute graphs for specific workloads, or redesigning memory architectures from scratch. Taalas baked model weights into the hardware itself. Etched designed a chip called Sohu that runs only transformers. d-Matrix moved inference compute inside SRAM arrays to avoid the bottleneck of fetching data from external memory. Each approach trades the programmability that made GPUs dominant for gains measured in orders of magnitude: speed, power, cost. Each also raises the same uncomfortable question. When does specialization tip into obsolescence?
The Memory Wall and the Power Surge
AI inference has hit what engineers call the memory wall. During the decode phase of inference, the operation is bandwidth-bound: fetching model weights and the key-value cache from external DRAM consumes most of the energy and latency, according to a survey published on arXiv in August. High-bandwidth memory remains scarce. TrendForce reported in June that HBM wafer input share among top DRAM makers would climb from 18 percent last year to 30 percent by 2027, with contract prices surging further. Gartner wrote in May that the bottleneck had shifted from GPU die supply to HBM and advanced packaging.
Power constraints are the second driver, and they matter in ways that extend beyond data center budgets. The International Energy Agency reported in February that data centers and AI now rank among the key drivers of global electricity demand. A separate IEA analysis found that AI workloads create rapid power swings requiring onsite generation capacity 30 to 70 percent above peak demand. The U.S. Energy Information Administration highlighted AI clusters in the Southwest in its February short-term energy outlook, warning that faster-than-expected data center load could push up fossil fuel generation. McKinsey wrote in March about how these sharp, rapid power swings are reshaping industrial planning around a data center build-out valued at $7 trillion.
Cost is the third force. Omdia raised its semiconductor revenue growth forecast for 2026 to 94.1 percent in July, driven by AI DRAM and NAND amid persistent supply tightness. Axios reported, also in July, that custom chip design has become easier than securing HBM and packaging capacity. IBM wrote in a trend piece the same month that everyone is trying to control more of the hardware stack. McKinsey said in May that inference cost reductions would come from system-level innovations—low-bit formats, memory-centric architectures, workload disaggregation—rather than from next-generation GPUs alone.
Hardcore Silicon

Taalas called its approach "hardcore," and the company's CEO, Ljubisa Bajic, meant it. The HC1 chip hardwired Llama 3.1 8B at the metal layer, eliminating external weight fetch entirely. Bajic said in the February announcement that the chip delivered "nearly 10× faster than the current state of the art, while costing 20× less to build, and consuming 10× less power." Forbes analyst Karl Freund, writing in February, confirmed token rates between 14,357 and 17,000 per second and noted that Taalas claimed a two-month metal-only respin cycle for new models. Heise reported in March that HC1 "can only execute Llama 3.1 8B, not any other models" and relies on proprietary 3-bit quantization with 6-bit parameters.
Taalas disclosed in February that it had 24 employees and had spent $30 million of more than $200 million raised. The company planned a mid-sized reasoning model on HC1 for spring and a frontier LLM on its next-generation HC2 chip for late in the year, according to its blog. AMD acquired Taalas on August 6, according to official announcements. The Register described the move as AMD hedging its GPU roadmap with model-specific options for mature, high-volume models.
Etched pursued a different bet. The startup designed Sohu to run any transformer architecture but nothing else: no convolutional neural networks, no LSTMs, no vision encoders. Founder and CEO Gavin Uberti said in a July TechCrunch podcast that ChatGPT "showed folks that, yes, there is demand for half a million tokens a second for one of these big models." He drew a parallel to Bitcoin mining. "Whether or not you're a fan of crypto currency, the Bitcoin mining ASIC companies have been able to do quite well for themselves… the moment that the first ASICs came out, they were better than GPUs by an order of magnitude."
Etched raised funding at a reported $10.3 billion valuation in July, with participation from SK hynix and other investors, TechCrunch reported. The chips were not yet on the market at the time. Uberti told the podcast he believes Etched has "more than an 18-month head start" on hyperscalers. The company built its bet on three technical choices: low-voltage inference, prefill-decode disaggregation, and cluster-scale memory, according to a June transcript of his appearance on the Invest Like the Best podcast.
d-Matrix took a third path. Rather than burning weights into silicon, the company moved compute into SRAM. d-Matrix announced on June 9 that its Corsair inference platform had entered full production and was shipping to priority customers. Corsair was built on TSMC's N6 process with organic substrates and LPDDR5, sidestepping the HBM and CoWoS bottlenecks that have slowed other accelerators, the company said. A March blog post detailed a 3D DRAM test chip called Pavehawk designed to extend on-chip capacity for full-model inference without external memory trips.
SambaNova and Cerebras represent a fourth pattern. Both companies pair GPUs for prefill with specialized decode accelerators. SambaNova CEO Rodrigo Liang said in an April announcement with Intel that "agentic AI is moving into production—and the winning pattern we're seeing is GPUs to start the job… and SambaNova RDUs to finish it fast." Cerebras published a deep dive on ultrafast frontier inference at Hot Chips in August, emphasizing wafer-scale memory and disaggregated inference architectures.
Market Signals

Capital markets are pricing specialization as viable. Groq raised $650 million in growth funding on June 22 for its programmable LPU inference cloud. Etched's reported $10.3 billion valuation in July signaled investor confidence despite the company's pre-production status. AMD's acquisition of Taalas in August indicated that incumbents are hedging GPU roadmaps with model-specific options for high-volume workloads.
Broadcom's custom AI semiconductor revenue guidance climbed from roughly $4 billion in fiscal 2023 to between $24 billion and $28 billion in fiscal 2026, according to a July analysis by Nodvolt citing company earnings. Microsoft deployed Maia 200 into Azure for GPT-class inference in January. Meta partnered with Broadcom on its MTIA generations scaling through 2027, the company wrote in March. CB Insights reported last October that Anthropic planned to use up to one million Google TPUs and that Oracle was preparing large AMD deployments for the second half of this year.
Omdia's market forecast, published three weeks ago, breaks out AI processor share by vendor, HBM content, and use case. The firm's December trends report, now nine months old, predicted that AI data center demand would keep data-processing semiconductors above 50 percent of overall semiconductor revenue in 2026 and said a diversified accelerator ecosystem was emerging.
The Trade-Offs
Every approach carries risk, and some are more obvious than others. Model-specific ASICs like Taalas HC1 require a new mask spin when weights change, turning every model update into a multi-month hardware cycle. Heise noted in March that quantization limits and single-model lock-in constrain flexibility. Taalas acknowledged the trade-off. CEO Bajic told Forbes in February that "our debut model is clearly not on the leading edge, but we decided to release it as a beta service anyway—to let developers explore what becomes possible when LLM inference runs at sub-millisecond speed and near-zero cost."
Architecture-specific chips like Etched Sohu face different exposure. Transformers dominate today. The field moves fast, though. Reconfigurable FPGA approaches explored in research published last November and new quantization schemes like BitNet b1.58, published in February 2024 and April 2025, could shift the economics. A paper on mask-programmed compute-in-ROM toolchains called Ankhdjet, published in August, suggests academic interest in hardwired inference remains strong. Deployment at scale has yet to prove out.
Independent benchmarks remain scarce. Taalas reported 17,000 tokens per second and a tenfold power reduction in February. d-Matrix cited SRAM bandwidth advantages in June. Cerebras and SambaNova published disaggregated inference claims through August. All of these are vendor-reported figures. No neutral, peer-reviewed head-to-head results comparing these accelerators against Nvidia B200 or AMD MI325X on identical models, contexts, and quality settings have surfaced as of early September.
Regulatory dynamics add another layer. U.S. export rules updated in January now key off chip performance metrics: total processing performance and DRAM bandwidth. They include case-by-case licensing below certain thresholds for China and Macau, according to the Federal Register. Design choices that reduce HBM or place weights on-chip may interact with those thresholds, though end-use and end-user review still applies, per BIS guidance updated this year. The EU AI Act's transparency rules took effect August 2, with general-purpose AI obligations and compute thresholds still under guidance, according to official EU pages.
What Comes Next

The next twelve months will clarify whether hardwired inference is a niche or something approaching mainstream. Etched's timeline suggests chips should reach customers by late this year. AMD's integration of Taalas will show whether a GPU incumbent can productize model-specific silicon alongside its MI300 line. d-Matrix's production ramp will test whether memory-centric inference can win enterprise deployments without HBM. Cerebras and SambaNova will prove whether disaggregated architectures scale beyond early adopters.
Three forces will shape adoption, and all three have become harder to ignore. Power and site availability matter first. The IEA and EIA analyses from February underscore that infrastructure constraints now limit deployment as much as chip supply does. That reality favors more efficient inference per watt. Model stability matters second. If foundation model architectures and weights ossify around a few dominant checkpoints, the case for hardwiring strengthens. If rapid iteration continues, flexibility retains its premium. Supply-chain access matters third. Axios reported in July that design is easier than securing HBM and packaging capacity, giving an opening to architectures that bypass those bottlenecks.
For buyers, the decision tree is relatively straightforward. Production workloads on frozen models with extreme latency or cost requirements suit hardwired approaches. Research workloads, multi-model deployments, and rapid fine-tuning still demand GPUs or programmable alternatives. The middle ground will be fought over by architecture-specific chips and memory-centric designs competing on total cost of ownership, vendor risk, and software maturity.
Venture investors should watch for three signals: customer proof points beyond beta programs, independent benchmark results on real workloads, and ecosystem traction around specialized toolchains. The Bitcoin ASIC analogy Uberti cited is instructive, perhaps more than he intended. Mining ASICs crushed GPUs on a fixed algorithm. AI inference is less fixed. But the economics of specialization follow the same physics. If the workload stabilizes, custom silicon wins.
The question is whether transformers and frontier LLMs will stabilize long enough for hardwired bets to pay off, or whether the next architecture shift will arrive before the first chips ship at scale. Taalas built a chip that can run only one model. That decision looked either visionary or reckless in February. By August, AMD had bought the company. The verdict on hardwired inference may arrive sooner than anyone expected.
