
Just before the launch of GPT-5, I published “The 11% Paradox - Why Orchestration Lock-In is Rewriting AI's Rules." As part of this longer article, I explored the topic of “Code as the Orchestration Wedge.” This concept prompted strong reader reaction and curiosity. GPT-5 was positioned as OpenAI's "smartest, fastest, most useful" model yet, with heavy emphasis on coding and agentic capabilities—directly amplifying the orchestration dynamics. Now that we are two weeks past the big GPT-5 reveal, I wanted to revisit the coding wedge to unpack it in more detail using the lessons learned from OpenAI’s launch.
The release of GPT-5 by OpenAI took place this month with great fanfare.
CEO Sam Altman had been hyping the release for weeks, calling it “a significant step along the path to AGI.” GPT-5 was formally unveiled during a one-hour-long livestream watched by millions. After almost two years of speculation about the development of this new model, users were beyond eager to test it.
The verdict has been...mixed.
Coming 162 days after the release of GPT-4.5 (reportedly a version intended to be GPT-5 but not achieving the necessary performance gains), developers immediately began dissecting its performance gains. The stats seemed promising: 74.9% on SWE-bench Verified (solving real-world software engineering problems), 94.6% on AIME 2025 (advanced mathematics), and reported improvements on health benchmarks. By the numbers, GPT-5 appeared to be bigger, faster, and smarter.
But community feedback has ranged from underwhelming to critical. Reddit users described the experience as disappointing, with some calling it a "massive downgrade." The technical community on Hacker News expressed skepticism about the claimed improvements. Even supportive reviews acknowledged it was primarily "quality of life features" and interface improvements rather than fundamental advances.
So, what’s the real story?
Neither. The true significance of GPT-5 isn't captured by a leaderboard or performance metrics. The new model represents OpenAI’s strategic push further into the orchestration layer, the invisible substrate that will determine who owns the AI economy.
While conventional wisdom holds that value in the AI stack accrues at the bottom with compute and the top with applications, this view is rapidly becoming obsolete. The most durable moat in the agentic era is being built in the middle at the orchestration rails that turn models into workflows and workflows into autonomous systems.
This is the new battleground. And so far, OpenAI has been outflanked by competitors such as Anthropic.
The launch of GPT-5 is one part of the company’s multi-front campaign to compete for this critical layer of the stack. The new model contains embedded orchestration intelligence, an important step toward owning more of the larger dedicated orchestration stack.
However, OpenAI recognized the need to move further up the stack and attempted to acquire Windsurf, which has developed an Integrated Development Environment (IDE) that enhances workflow coordination for software developers. The collapse of its multi-billion-dollar bid for Windsurf was a setback to these orchestration efforts.
While software development is just one of many tasks that LLMs can perform, it has emerged as the strategic high ground, a competitive wedge that every major player is competing to control. That’s because the model that captures the coder ecosystem is positioned to define the infrastructure layer of the agentic era—and with it, the next decade of enterprise AI architecture.
In the wake of the GPT-5 launch, I want to examine why the coding wedge has become so critical and its implications for competing in the Agentic Era.
I will also release my pre-launch analysis of OpenAI's reinforcement learning strategy tomorrow, in which I explain why it represented a credible bet for achieving orchestration dominance.
The Orchestration Layer as AI's Lego Instructions
In "The 11% Paradox," I established how orchestration lock-in has become the dominant force in AI markets, with only 11% of enterprises switching providers despite model commoditization. That analysis revealed the power of behavioral lock-in created by orchestration dependencies.
GPT-5's launch and OpenAI's strategic maneuvers provide a compelling case study of how this competition unfolds.
Think of AI orchestration like Lego instructions. Individual AI models, tools, and data sources are sophisticated Lego pieces, powerful but useless in isolation. Orchestration provides the instruction manual: how pieces connect, in what sequence, toward what structure. Without orchestration, you have advanced blocks but cannot build. With it, those same blocks construct everything from simple tools to complex systems.
Coding serves as the perfect wedge because writing code is essentially creating Lego instructions at multiple abstraction levels. Functions assemble small components. Classes combine components into units. System architectures show how units create something greater. When AI models learn to write code, they learn these orchestration patterns—decomposing problems, managing dependencies, coordinating components.
This dynamic explains recent market shifts. According to Menlo Ventures data, Anthropic's surge from approximately 10-15% to 32% enterprise market share wasn't driven by marginally better benchmarks. Claude Opus 4.1 achieves 74.5% on SWE-bench Verified, statistically identical (and even a bit less) to GPT-5's 74.9%.
Instead, as Menlo notes, the Claude models released over the past year "introduced the first real glimpse of an agent-first LLM.” That difference was Anthropic's Model Context Protocol (MCP) becoming the industry's universal Lego connector, adopted even by competitors OpenAI and Google. When enterprises choose Claude, they're selecting an entire instruction system that becomes progressively harder to replace.


