Etched: The $5B Sohu Chip Betting the Transformer Never Dies
AI News 5 min read

Etched: The $5B Sohu Chip Betting the Transformer Never Dies

Etched, a startup building the transformer-only Sohu inference ASIC, has booked over $1 billion in contracts and reached a $5 billion valuation, with reports of new rounds valuing it up to $20 billion. Sohu hard-wires the transformer graph into silicon on TSMC N4P with 144GB HBM3E, and Etched claims an 8-chip server exceeds 500,000 Llama 70B tokens/sec. No independent benchmarks exist yet.

Sarah Chen
Sarah Chen
Jul 25, 2026

Nvidia spent a decade convincing the world that the GPU was the only way to run AI. Etched is betting $800 million that it was wrong — and this month, investors started agreeing at eye-watering numbers.

The startup came out of stealth in late June with a claim that would sound reckless from anyone else: it has already booked more than $1 billion in contracts for a chip that does exactly one thing. That chip is called Sohu, and it has now vaulted Etched to a $5 billion valuation — with reports this month that the company is in talks for rounds that could push it far higher.

The One-Trick Chip

Sohu is an ASIC — an application-specific integrated circuit — and its specialty is narrow on purpose. Where an Nvidia GPU is a general-purpose parallel computer that can run transformers, Sohu hard-wires the transformer computation graph directly into silicon: attention, linear projections, softmax, and layer normalization are etched into the hardware itself.

That is the whole bet. Nearly every model that matters today — GPT, Claude, Gemini, Llama — is a transformer. Etched's founders wagered in 2022 that this architecture would dominate long enough to justify building a chip that can run only transformers, and nothing else.

The tradeoff is stark. If a fundamentally new architecture dethrones the transformer, Sohu becomes an expensive paperweight. If it doesn't, Etched owns the fastest inference silicon on the planet.

The Numbers Etched Wants You to See

Etched's performance claims are the loudest in the industry. According to the company, a single server packing eight Sohu chips can sustain more than 500,000 tokens per second running Meta's Llama 70B.

For context, here is how Etched frames the same workload:

Server (8 accelerators) Llama 70B throughput
8× Sohu ~500,000 tokens/sec
8× Nvidia Blackwell B200 ~45,000 tokens/sec
8× Nvidia H100 ~23,000 tokens/sec

The chip is built on TSMC's N4P 4nm-generation process and pairs its compute die with 144GB of HBM3E memory. Etched credits the throughput to floating-point utilization above 90% on attention-heavy workloads, versus roughly 30% for general-purpose GPUs that waste cycles on control logic and scheduling.

The bottleneck in AI has quietly shifted. Training grabs headlines, but inference — actually serving answers to users — is now the biggest recurring cost for every AI company. That is exactly the pain Etched is aiming at.

A Cap Table That Reads Like an AI Hall of Fame

Founded in 2022 by Harvard dropouts and Thiel fellows Gavin Uberti (CEO) and Robert Wachen (president), Etched spent its early years nearly running out of cash. The founders have said that in 2023 every major investor they pitched passed, even after a 30-page memo arguing AI would eventually demand specialized chips.

The mood has reversed completely. Etched's $500 million round — which closed in December at the $5 billion post-money valuation and was led by Stripes — drew in Jane Street, Hudson River Trading, Two Sigma, and Ribbit Capital. The angel list is its own flex: Andrej Karpathy, Geoffrey Hinton, Fei-Fei Li, Arthur Mensch, and Scott Wu, with billionaires Stanley Druckenmiller and Peter Thiel on the cap table too.

That momentum has only accelerated. This month, multiple outlets reported Etched is in talks for back-to-back rounds — including a deal led by Sequoia — that could value the company at up to $20 billion, a roughly 4x jump in a matter of months.

The Crowded Race to Un-seat Nvidia

Etched is not alone in smelling blood. Cerebras delivered one of the year's breakout IPOs, AI chipmaker Groq confirmed a $650 million raise, and even OpenAI unveiled its first custom chip built by Broadcom. Amazon, Google, and Microsoft all design their own in-house silicon. Inference has become the single most contested frontier in AI infrastructure.

What sets Etched apart is how far it is willing to specialize. Everyone else hedges; Etched has removed the hedge entirely.

The Asterisk

Here is the part the funding announcements skip over: Sohu has not yet shipped in volume to external customers, and no independent third-party benchmarks exist. Every eye-popping number above is Etched's own. The systems — which the company calls "frontier inference clusters" and sells as bundled racks with custom software — are still in customer testing.

A billion dollars in contracts is a strong signal that serious buyers believe the pitch. It is not the same as a billion dollars of chips proven in production. Until an independent lab straps a Sohu server to a real workload, the 20x claims live in the space between engineering ambition and marketing.

The Bottom Line

Etched is the purest expression of a conviction sweeping AI infrastructure: that the transformer is here to stay, and that the future belongs to hardware built for it specifically rather than for everything generally. If that conviction holds, a $20 billion valuation may look cheap in hindsight. If a new architecture breaks through, Etched will have built the most expensive bet on a single idea in chip history. For now, the contracts are real, the silicon is fabricated, and the benchmarks are still waiting for a referee.

More in AI News

Project Perception: Microsoft's Cheaper Rival to Claude Mythos
AI News

Project Perception: Microsoft's Cheaper Rival to Claude Mythos

Microsoft is reportedly developing Project Perception, a multi-model AI security platform that routes vulnerability-scanning tasks across models from Microsoft, OpenAI, and Anthropic to reserve expensive frontier calls for high-value steps. Its pitch is matching Anthropic's Claude Mythos on capability while costing far less. Microsoft has not officially confirmed details, so the news should be treated as a credible report pending benchmarks.

By Sarah Chen · 5 min · Jul 21, 2026

Inkling: Mira Murati's Thinking Machines Ships Its First Open Model
AI News

Inkling: Mira Murati's Thinking Machines Ships Its First Open Model

Thinking Machines Lab, founded by ex-OpenAI CTO Mira Murati, released Inkling on July 15, 2026 — an open-weight mixture-of-experts model with 975B total parameters (41B active), trained on 45 trillion multimodal tokens. The company openly says it isn't the strongest model available; instead it's a customizable foundation enterprises fine-tune via the Tinker platform. The release doubles as an argument that owned, adaptable models beat rented one-size-fits-all APIs.

By Sarah Chen · 5 min · Jul 18, 2026

Kimi K3: Moonshot's 2.8T Open Model Nears the Frontier
AI News

Kimi K3: Moonshot's 2.8T Open Model Nears the Frontier

Moonshot AI released Kimi K3 on July 16, 2026, a 2.8-trillion-parameter open Mixture-of-Experts model that activates 16 of 896 experts, ships native vision and a 1M-token context, and leads benchmarks like SWE Marathon, BrowseComp, and OmniDocBench while trailing Fable 5 and GPT-5.6 Sol overall. Weights release July 27 under a Modified MIT license.

By Sarah Chen · 5 min · Jul 17, 2026