Editorial note, August 2026. This piece was published on 23 September 2025 as Part 1 of two and is preserved as written. Oracle's backlog and the model costs cited below reflect that date. Part 2 covers inference economics.
Oracle's stock surged 36% after disclosing a staggering $455B backlog that included an OpenAI agreement to buy $300B in computing power. The stock’s largest one-day jump in 30 years added $244 billion in market cap. This might feel like market exuberance, but in fact, it is recognition that AI has broken Moore's Law. Where compute once got cheaper every two years, AI now demands exponentially more compute for linear improvements, with GPT-5 training alone costing an estimated $500M and compute training cost projections of 10x for next SOTA models per industry experts. The result is a bifurcated market where training compute is oligopolistic, while inference compute becomes the recurring cost battleground. The entire $7 trillion infrastructure buildout rests on one fragile assumption: that OpenAI and a handful of labs keep getting funded. If that funding stops, Oracle's backlog evaporates, and the AI bubble bursts.
On September 10, 2025, Oracle Corporation became $244 billion more valuable in a single trading session.
The catalyst wasn't a breakthrough product or strategic acquisition. It was a number so large that seasoned analysts struggled to contextualize it: a $455 billion backlog for compute demand, up 359% year-over-year! The stock's 36% surge, its best day since 1992, came after OpenAI agreed to buy $300 billion in computing power from the company, a deal that validates what insiders have known for months: compute infrastructure has become AI's existential bottleneck.
As I wrote this summer (2025), compute scarcity is reshaping the future of AI. Almost two months after my analysis of compute constraints as a potential single point of failure, Oracle's historic surge, and Qwen3-Next's efficiency breakthrough demand a deeper examination of AI's extraordinary infrastructure investments.
The momentum shows no signs of slowing.
Last week, Bloomberg reported that Oracle is in talks with Meta for a potential $20 billion AI cloud computing deal, adding to existing multi-billion-dollar commitments from OpenAI, xAI, and Google. The numbers are so massive, even the CEOs appear to have trouble keeping track. Meta CEO Mark Zuckerberg, fielding an unexpected question at a recent dinner with President Trump, said he planned to spend "at least $600 billion through 2028 in the US,” though admitting later that he was subsequently caught by surprise, suggesting he just made up a random number. Meta would have to dramatically ramp up its AI spending to hit $600 billion in the next three years!
No matter. The message is the same. In the age of AI, access to compute infrastructure is existential.
Hyperscalers and AI labs are accelerating ahead, and also creating an interdependency that is not without systemic risks. What truly defines today's compute landscape is the intricate web of dependencies between AI labs, NVIDIA, and cloud providers. Yesterday’s announcements underscore this fragility. NVIDIA's $100B investment in OpenAI ties chip supply directly to lab funding, while its $6.3B guarantee for CoreWeave's unused capacity ensures overflow compute but highlights overbuild risks if demand softens.
OpenAI cannot function without Microsoft's Azure infrastructure and NVIDIA's H100 clusters. Anthropic depends on AWS's custom Trainium chips and compute capacity. Even supposedly independent players like xAI must partner with Oracle and CoreWeave while spending billions building their own infrastructure.
This interdependence creates a precarious equilibrium: If any part of this intricate balance of partnerships buckles, the size of the economic collapse could be massive.
And yet, none of them can afford to be cautious. Waiting means falling irrecoverably behind.
This is the discontinuity we must decode: Compute has become the bottleneck and the arbiter of who wins in AI. Not algorithms, not data, not even talent—though all matter—but raw, specialized, massively scaled computational power. The implications ripple through every layer of the stack, from the oligopolistic dynamics of training clusters to the knife-edge economics of inference at scale.
The End of Moore’s Law Era: The AI Revolution's Physics and Economics
The AI revolution operates under different physics and economics compared to traditional Moore's Law-driven progress.
For nearly six decades, Gordon Moore's observation governed the technology industry's cadence. Moore's Law, articulated in 1965 when Moore was director of research at Fairchild Semiconductor, predicted that the number of transistors on a microchip would double approximately every two years while the cost would halve. This became the industry's metronome, the assumption underlying every business plan, every venture investment, every strategic roadmap. Compute would get exponentially cheaper and more powerful as reliably as the seasons change.
Every major technology wave of the past half-century rode on Moore's Law's reliable cadence: costs would fall, capabilities would rise, and innovation would naturally follow. This slowed in recent years, with transistor density improvements slowing to perhaps 20-25% annually. While this was a far cry from Moore's predicted doubling, the overall capacity/price ratio continued to move just enough in the right direction to improve the underlying economics.
Today's AI revolution operates under fundamentally different physics and economics.
First, there is no guarantee that increased demand for compute can be met with increased supply. As AI compute demand has exploded by orders of magnitude, the rush to build the necessary infrastructure involved complex projects with timelines of years.
Next, the transformer paradigm, with its empirically validated scaling laws, has revealed that model performance scales predictably with compute, but at escalating costs, turning AI progress into a capital-intensive race rather than a predictable, cost-reducing evolution.
Operating costs are now at levels that would have seemed fantastical just five years ago:
GPT-5's training alone is estimated to have consumed over $500 million in compute resources
Frontier models from Anthropic and Google approach similar magnitudes
DeepMind's Gemini Ultra reportedly required $191 million for a single training run, with multiple runs needed for hyperparameter optimization
This goes deeper than raw costs or dependencies.
In the Moore's Law era, waiting two years meant getting the same capability at half the price. In the transformer era, waiting means falling irrecoverably behind.
OpenAI's GPT-4 required an estimated 10,000 times more compute than GPT-2, released just three years earlier. The next SOTA model is expected to require 10x more compute.
This isn't exponential improvement at declining costs.
It's exponential costs for linear improvements in capability.
The Chinchilla scaling laws, validated across dozens of models, show that optimal performance requires balanced scaling of both parameters and training data, which means compute requirements grow faster than model size alone would suggest.
To understand this discontinuity, we must examine two distinct but interrelated compute realities: training and inference.




