Decoding Discontinuity

Decoding Discontinuity

TurboQuant and the Memory Stock Sell-Off: Why the Panic Outpaced the Paper

Why this efficiency gain is ultimately bullish for the memory chokepoint and the entire inference economy.

Raphaëlle d'Ornano's avatar
Raphaëlle d'Ornano
Mar 31, 2026
∙ Paid
Photo by Alexander Mils via Unsplash

A Google blog post about compressing AI memory went viral. Within 48 hours, the market capitalization of memory semiconductors evaporated by over $100 billion. TurboQuant helps solve a core problem of the Agentic era: making long-context LLM inference efficient. But the algorithm compresses only the inference-time cache, not the model weights, training data, or storage.


The latest sign of the market’s struggle to price the Agentic Era arrived in the form of a Google blog post about AI memory compression. Within 48 hours, ~ $100 billion in semiconductor value evaporated, not because the technology changed the economics of the stack, but because investors misidentified where those economics now sit.

The tremor began when Google Research published a blog post on March 24, titled “TurboQuant: Redefining AI Efficiency with Extreme Compression.“ The Google post summarizes a series of papers actually published between mid-2024 and April 2025.

Nothing about the science was new. Only the packaging was new.

Google distilled three academic papers into a single, accessible narrative in a blog post and then a tweet: 6× less memory, 8× faster inference, zero accuracy loss. The tweet received close to 19 million views.

The hot take reactions were predictable: Improved memory efficiency would reduce overall hardware demand. TechCrunch referenced the HBO satire “Silicon Valley,” calling it the “real-life Pied Piper.” Cloudflare CEO Matthew Prince called it “Google’s DeepSeek moment,” noting that there is “so much more room to optimize AI inference for speed, memory usage, power consumption, and multi-tenant utilization.”

The comparison primed investors for fear. Algorithmic selling began a scorched-earth campaign across the memory infrastructure sector.

Micron fell 30% since its March 18 earnings report. SK Hynix dropped 6.2% in a single session. SanDisk - NAND flash storage, zero connection to inference-time cache compression - shed 18% in five days. NVIDIA, which actively builds quantization tools and whose Blackwell architecture is optimized for exactly the kind of low-precision computation TurboQuant enables, fell 6.6%.

Here is the problem with those hot takes: TurboQuant is significant. But not for the reasons widely assumed.

TurboQuant points to an important advance in a very specific way, one that may address a critical bottleneck. However, the sell-off conflated a narrow efficiency gain at one layer of the AI stack with a structural reduction in demand across the entire stack.

In that respect, the market reaction to TurboQuant is not just a misreading of a paper. It is a symptom of a deeper analytical gap in how investors are pricing the AI stack in the Agentic Era. The consensus still treats AI infrastructure as a monolithic trade. Everything rises together on bullish narratives; everything falls together on efficiency headlines.

But memory is not compute. The KV cache is not the hard drive. And Sandisk is not SK Hynix.

With the right lens, the TurboQuant tale will, in fact, deliver yet another twist: rather than being a looming threat value destroyer, it is more likely that overall demand will expand as cheaper, more efficient inference unlocks far greater scale, concurrency, and adoption of AI systems.

Share

Why Memory Stocks Fell After TurboQuant

To understand why the TurboQuant sell-off was wrong, you first need to understand what memory has become in the AI economy, and why it is not the “picks and shovels” metaphor we keep reaching for.

Picks and shovels are commodity inputs. Abundant, interchangeable, priced at marginal cost. Memory is the opposite of that. Memory is the single most concentrated, supply-constrained, highest-pricing-power layer in the entire AI infrastructure stack.

South Korean companies SK Hynix and Samsung control about 80% of the global HBM (high-bandwidth memory) supply, the specialized memory chips physically stacked onto every AI GPU. Micron, a US company and the third-largest producer, acknowledged in December that it meets only about half of its backlog and warned that the crunch will persist beyond 2026. New fabrication facilities take years to qualify. Goldman Sachs projects a 4.9% DRAM undersupply in 2026. That’s the most severe shortfall in more than fifteen years.

Even if you manufacture enough memory, you cannot assemble it without TSMC’s CoWoS advanced packaging, which physically bonds HBM to the GPU. TSMC has indicated CoWos capacity is fully utilized and will remain very tight into 2026. NVIDIA alone has secured 60% of total CoWoS output.

Both constraints are physical. Neither responds to software breakthroughs. Neither can be resolved by an algorithm.

This is why memory stocks ran 200 to 1,200% over the prior year. These three companies represent a form of structural scarcity at the binding constraint of a $4T capex cycle over five years per FT estimates. These companies are directing the bulk of their operating cash flow, supported increasingly by debt, to build AI infrastructure.

Every dollar of that spending flows through the memory bottleneck at some point.

Forecast capex by calendar year
Figure 1. Forecast capex by calendar year ($bn) (source: FT)

In any technology system, value concentrates at the layer that constrains throughput. Cloud compute, by contrast, is being built by five hyperscalers, plus CoreWeave, Lambda, sovereign clouds, and enterprise on-premises deployments.

While the cloud layer converges toward utility pricing and, in the worst case, leads to oversupply, memory does not. The scarcity is structural, the barriers are physics-based, and the timeline for new capacity is measured in years, not quarters.

Memory is not picks and shovels. It is closer to the new Magnificent Seven. Except that the concentration is even tighter. Two or three companies, not seven, sit at the single most constrained point in the most capital-intensive buildout in the history of technology.

That is the context the market forgot when the Google blog post went viral.

What TurboQuant Actually Compresses: The KV Cache

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Raphaëlle d'Ornano · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture