
TL;DR: Does a model that cannot write change the economics of AI? I think it does. TypeSafe’s Jev charges $0.042 per million input tokens and nothing for output, separating cheap structured decisions from the generated reasoning frontier labs sell at up to $10 in and $50 out.
In barely two weeks, TypeSafe’s Jev went from launch to viral adoption, six open clones, and investor “fervor” that could value TypeSafe above $10 billion, according to a September 24 report by The Information.
That rapid industry recognition of Jev’s value reinforces my view that it is another tremor in the generative and agentic AI discontinuity: a shift in one layer that amplifies the others. These tremors are building the structural foundation for the Agentic Era in which intelligence is cheap, ubiquitous, and coordinated at machine scale.
My bet is that Jev is a tremor because it points to a broader architectural shift in which intelligence is being decomposed into specialized forms of computation, each invoked only where it is economically appropriate. Jev focuses on decisions: the bounded judgments software makes inside workflows. In doing so, they are becoming function calls priced toward zero. Deliberation remains the premium product for frontier labs.
That distinction changes the architecture of an AI workflow. A cheap decision model can handle routine forks. More capable models are called when the decision is uncertain, or when prose and deeper reasoning are required. The frontier does not disappear. It becomes more selective.
That matters after the backlash against “tokenmaxxing” and the push toward more efficient agentic systems. Jev is a signal that developers increasingly understand the frontier model need not adjudicate every step of a workflow. Where bounded decisions can be resolved reliably, software can buy less frontier inference while keeping control of the workflow, its thresholds, and its record.
If we carry that idea forward, my argument is that intelligence is beginning to separate into two products:
Decisions, the narrow, structured judgments that make up many of the routine judgments software asks models to make, have become a function call priced toward zero.
Deliberation, open-ended reasoning and writing, is what the frontier sells, at ten dollars in and fifty out.
For OpenAI and Anthropic, the competitive threat would not necessarily be a superior model that beats them. It would be software that buys less from them and retains control of the workflow and its record of decisions. Jev makes such a consequential architectural choice easier to implement.
Where those decisions can be resolved reliably, some pricing power moves away from the model layer and toward whoever controls the workflow and the record of what was decided.
TypeSafe itself may be copied away. The architecture it exposed may be harder to reverse.
The unresolved questions are how much work qualifies, whether the economics survive validation and escalation, and who ultimately captures the resulting value.
What is TypeSafe’s Jev decision model?
Jev reads text and answers predefined questions with a choice or score and a probability, rather than generated prose.
On September 15, San Francisco-based TypeSafe emerged from stealth with a $40 million seed round led by DCVC. Co-founder and CEO Diogo Almeida left OpenAI in 2024 and is one of the primary authors of the paper that taught language models to follow instructions, the paper that made ChatGPT possible.
The announcement of Jev caused an immediate shockwave. The launch video was watched about forty million times. The company explains that users give Jev the state of their program and a question with the possible answers declared in advance. It returns one answer and a probability. Input costs $0.042 per million tokens. Output is free, in the company’s words “too cheap to meter”.
TypeSafe calls it a System One model, after Daniel Kahneman’s fast, intuitive mode of thought. Developers submit up to about 32,000 tokens and one or more questions-a choice from a list, a score, a yes or no-and receive answers with probabilities in about a third of a second. Jev cannot choose “none of the above” unless that option is supplied. Almeida says Jev was trained exclusively on synthetic data using “reinforcement learning for calibrated decisions.”
Understanding the method behind Jev matters, because the product's value depends on it.
Where the training behind ChatGPT rewards answers people prefer, Jev’s stated objective is reliable confidence: among answers assigned 90 percent probability, nine in ten should be right. A related approach, described in “Rewarding Doubt”, rewards probability estimates while penalizing over- and underconfidence. That illustrates the idea, not Jev’s undisclosed implementation. TypeSafe has published neither the method nor the model’s weights or parameter count.
Outside analysis of the API suggests a language-model backbone repurposed for decisions, with a probability read directly from the network. Several open clones use that design. What remains is the ability to read text and answer declared questions without task-specific training. What disappears is the generation stage.
That is the economic hinge.
A language model first processes its input, then generates its answer one token at a time. This phase is called decode, where hidden reasoning also accrues. On that outside account, Jev processes the input and reads off a decision without decoding. The saving comes from removing generated text, not merely discounting it.
The trade-off is opacity. Simon Willison notes that a language model’s explanation may be unreliable, but Jev offers none: a move “even further towards black box machine learning”. For the audit-trail problem I described after Astra, that is a further loss of visibility. Almeida frames the change differently: intelligence “will be more like a database than a coworker”. You ask it questions, millions of times, cheaply.
Independent comparisons find Jev faster and cheaper, but the calibration claim is less secure. On a phishing benchmark, one broad question produced 62.6% accuracy, compared with Haiku’s 81.3%. Those results show how much accuracy depends on question structure. They do not, by themselves, establish whether Jev’s confidence scores are reliable.
Splitting the question into five and fitting a regression lifted Jev to 95%. Haiku, given the same decomposition, reached 93%. As one reviewer put it, “the accuracy and the calibration aren’t in the box.”
Should the evaluation measure calibration as well as accuracy: are high-confidence answers reliable enough to accept without escalation? The relevant cost comparison should include the cases that still require another model.
TypeSafe’s headline claims it is 193.6 times faster and 444.6 times cheaper. But its accuracy evaluation raises a separate question: the reference answers come from GPT-6 Astra and Fable 5.1, so agreement with those answers is not independent evidence of correctness.
Jev’s speed and cost are the attraction. Reliable confidence on the customer’s own questions remains the proposition to prove.
Why does Jev matter beyond TypeSafe?
Selling structured decisions separately from generated reasoning could change what software buys from AI, even if TypeSafe captures little of the value.
While we don’t know how much value it could potentially capture, consider what happened in just the first 10 days after Jev was announced:
Within days, six open clones emerged, one of them trained on a MacBook in under two hours.
On September 18, Vercel reported that nearly 13% of the paid teams on its AI Gateway had called Jev in its first twenty-four hours, twice the GPT-5.6 family’s share and six times Fable 5.1’s, adding that “the next test is whether that early adoption lasts.”
A few days later, TypeSafe paused new sign-ups under the demand.
Anthropic and OpenAI subsequently cut prices within ninety minutes of each other. That was a price war of their own that neither lab tied to Jev: the cuts share a week with Jev, not a cause, and what they show is the deflation Jev now prices at the floor.
On September 24 and 25, The Information and then the Financial Times reported investors offering to fund TypeSafe at ten billion dollars or more, against a last mark of about two hundred million; Bloomberg followed.
On September 25, OpenRouter, the marketplace Stripe is buying, listed a Jev Router: a decision model sitting in front of the frontier models, choosing for each request which one answers and how hard it should think. While this intense early interest is impressive, for me, I try to place such developments in the context of my larger Orchestration Economics thesis. In the Manifesto I published in April, I described the current generative and agentic AI discontinuity as a fault line marked by successive shifts across layers, each amplifying the others.
In that respect, I would ask: Is Jev another tremor...or a product launch?
I see Jev as another tremor in the strict sense I gave the term in July: “a phase shift in one layer that amplifies the others.” Jev’s shift concerns what software buys. The decision becomes a product, separated from the generated reasoning that previously accompanied it.
The bear case is TypeSafe’s defensibility. Jev’s form was reproduced in about nine hours on a consumer graphics card. Another project distilled half a million of Jev’s answers into a student that its creators say matches it on six public benchmarks. Vercel already tells developers to use a generic “evaluation API” rather than TypeSafe-specific naming. The examples still call Jev, but the name has gone from the slot. TypeSafe may be the first casualty of what it started.
In that case, the evidence points towards Jev being more of a tremor in terms of being a category that extends beyond the clones. Sakana’s Fugu already used a learned coordination head to delegate tasks, but sold the complete answer, including the text. Jev sells only the decision. Fastino now offers a 340-million-parameter competitor that runs on an ordinary server processor. The proposition is no longer confined to one company.
That changes accessibility: not merely a cheaper language model, but a narrower product, priced toward zero and copied in a weekend. It also points toward the architecture I argued in March, when I described agent-to-agent text as “slow, lossy, and expensive.” Some of the inner loop could run on decisions rather than prose.
One admission matters. I am naming a “decision tier” ahead of conclusive evidence. Speed, cost, and the ability to answer unfamiliar questions are demonstrated; reliable calibration is not. A fast classifier with no output charge can still change the bill. Whether it establishes a distinct product tier, rather than a cheaper tool, remains the larger claim to prove.
My bet is that it does: the moment software can buy the decision without buying the deliberation, intelligence has become two products.
Where does Jev fit in an AI workflow?
Jev handles bounded decisions inside a code-controlled workflow, with harder judgments escalated to another model or a person. The clearest way to see what changes is to follow one ticket through the two designs.

Picture a support desk. In the agent-loop design, a frontier model reads a ticket, reasons about its category and urgency, chooses a tool, checks the result, and decides what happens next. The model directs the workflow; it bills generated reasoning along the way.
TypeSafe’s rule reverses that relationship: “Code owns the workflow, and AI handles narrow, structured decisions.” The program receives the ticket, asks Jev which team should handle it, fetches the relevant records, and assesses urgency. Jev returns a label and a probability. Above a developer-set confidence threshold, the program continues; below it, another model or a person takes over. A generative model still writes the reply. The code decides what happens next; the model answers questions.
The architecture predates Jev. Anthropic’s December 2024 guide describes predefined workflows and routing, while Vercel describes production systems combining cheap classifiers with frontier reasoning. Jev’s contribution is a model packaged for the narrow question, without an output-token charge.
How much that saves depends on the question. In one developer’s forty-ticket test, a strict confidence threshold escalated four routing decisions and made the batch nearly ten times cheaper. Applied to urgency, the same threshold escalated thirty-eight and saved almost nothing. The frontier becomes an exception handler only where the cheaper model can reliably resolve most cases.

This also sharpens the question I raised in July about Stripe and OpenRouter: does model selection remain a defensible judgment, or become commodity plumbing? I identified two routes to commoditization: software absorbing routing upstream, or a stable frontier making the choice reducible to a table. Jev supplies a third: the routing table learned into a cheap decision model.
My argument therefore shifts toward what surrounds the choice. If selection becomes a commodity, the gateway’s value must rest on the traffic it sees, the constraints it enforces, and the settlement that turns a model call into an accountable transaction.
But the gateway need not see every decision. Where a decision model runs inside the customer’s systems, the customer’s program holds the thresholds, state, and record. The gateway may receive only the escalations. Those calls still lead to the expensive frontier models, but the distinction matters: controlling the paid escalation path is not the same as owning the workflow.
Is AI inference still getting cheaper?
Yes at the low-cost end, but not uniformly at the frontier. Jev adds another route to savings: avoiding generated reasoning for decisions that do not need it.
The change is not simply another reduction in token prices. It is the ability to buy a structured decision without also buying generated reasoning. This month, two curves are moving apart. In June, I wrote that the fight had left the per-token price, “which is going to zero, and moved to efficiency itself: the compute you burn per piece of useful work. On the outside account of its architecture, Jev takes the compute per decision down to one pass.
At the summit, the first half of 2026 was a price increase. OpenAI’s GPT-5.4 and GPT-5.5 each launched above their predecessors, and Anthropic’s new tokenizer in April produced about 30 percent more tokens for the same text at an unchanged list price. Below the summit, OpenAI’s turn came on July 30, when it cut its cheapest new model by 80 percent three weeks after launching it.
September’s cuts show the pressure below the summit. On September 22, Anthropic reduced Opus 5.5’s input and output prices by a fifth. OpenAI followed with GPT-6 Sol and Luna at half the preceding generation’s prices. Yet the flagship ceiling holds at $10 per million input tokens and $50 per million output tokens for Astra and Fable 5.1.
Price-performance is improving even when flagship prices are not. Epoch AI’s September study finds that the cost of achieving a given performance level has fallen about 47 percent per quarter since 2023, including hidden reasoning. As I wrote after Astra, “the price of a token has risen, but the price of a useful unit of work may still be falling”. Jev shifts the focus again, from the task to the decision.

The critical line on the bill is output. At the flagship rates, it costs five times as much as input, and both OpenAI and Anthropic charge for hidden reasoning as output. Their September cuts treated that line differently. Anthropic reduced Fable 5.1’s cached-input price by 75 percent while holding output unchanged. OpenAI cut output harder than input. One preserved the output premium; the other compressed it.

Consider an illustrative routing question with 2,000 input tokens and 500 hidden reasoning tokens. On Opus 5, it costs 2.25 cents: output represents one-fifth of the tokens but more than half the bill. Jev’s input-only charge is about 0.008 cents, roughly 270 times less. With the Opus prompt cached, the gap remains about 160-fold, and output accounts for 93 percent of its bill. Against GPT-6 Luna, the same workload produces a gap of about 5.4 times; these are price comparisons, not evidence of equal decision quality.
Compare input prices alone and Jev’s advantage against a cached frontier prompt is only five to twelve times; its input price is within 20 percent of GPT-5 nano’s. The much larger savings in this scenario come from avoiding generated reasoning, not simply buying cheaper input.
That exposes revenue but does not by itself establish the margin claim. The labs’ cost ratio between processing input and generating output is not public, so a fivefold price premium does not prove that output is disproportionately profitable. What Jev makes concrete is a different route to inference deflation: some decisions can become cheaper because they stop buying generated tokens at all. How much of the suppliers’ profit disappears with those tokens remains unproven.

What do decision models mean for OpenAI, Anthropic, and compute?
Decision models like Jev could remove paid reasoning from routine workflows. The effects on profits and compute demand depend on what is displaced and who supplies the replacement.
The exposure starts with work still bought at frontier prices. On September 22, Opus 5 led OpenRouter’s classification spending. One reported financial-services deployment spent more than $200,000 a month on inference, although over 70 percent of queries were within smaller models’ capabilities. When those requests move to a competing supplier or a locally run model, the lab loses revenue and visibility. Routing to its own cheaper model reduces spending without necessarily transferring the customer relationship.
Where the exposure sits is easier to say than how big it is.
Anthropic is the more usage-weighted business. Outside estimates put pay-per-token API at 70 to 75 percent of its revenue, against 15 to 20 percent for OpenAI by one count and closer to 30 percent by others. Neither company discloses the mix. But Anthropic’s API book is dominated by coding, with Claude Code alone at about $8 billion of a $47 billion run-rate in May, and agentic coding is the deliberation workload, long outputs and hidden reasoning, not the decision.
What share of either lab’s frontier spend is decision-shaped, nobody outside the labs knows. The one public deployment figure, more than 70 percent of queries routine in a financial-services rollout spending over $200,000 a month, is a share of queries, and routine queries are short, so the share of spend is far smaller.
We know the spend exists. On September 22, Opus 5 was the model OpenRouter’s users spent the most on for classification, and a cascade with Jev in front took 76 percent of those requests away from it. Run the multiplication on any plausible inputs and today’s revenue at risk is low single digits for either lab, a rounding error against growth of fourteen-fold a year.
That is not the claim. The claim is about pricing power on everything below the summit, and about the inner loop of the agentic workflows both labs are pricing their next decade on.
The decision tier removes that loop from the growth before it is booked. The recent price cuts align with pressure below the frontier, but they don't establish that Jev caused it. July’s cuts preceded Jev, and neither lab’s September announcement mentioned it.
The first counterargument is Jevons, the paradox TypeSafe named its model after: cheaper capability creates more demand. OpenRouter’s token usage rose nearly fourteen-fold after OpenAI’s July cuts, combining new usage with traffic taken from competitors. My concern isn't that activity shrinks, but that spending shifts toward cheaper suppliers. More tokens do not settle who captures the revenue.
Second, labs can copy the product. They have the models and distribution. As Arcturus Labs’ John Berryman writes, the architecture offers little obvious moat. Their silence has two readings. The simple one is that a young product with unpublished weights does not yet merit a response. The reading I hold is that a zero-output product is strategically awkward to copy because it gives away the very line the labs still price at a premium. If that reading is wrong, the third falsifier below is where it will show up.
Demand for the frontier isn't disappearing either. Astra took 7.7 percent of Vercel’s spending in twelve days at the top price. My argument concerns the routine work below it: whether labs retain those calls, their visibility and their pricing power. As I wrote in August, a system of execution is a token-minimizing business, while reserved capacity makes its supplier dependent on token volume. Decision models bring that tension into the workflow.
The economic question was neatly summarized at the Federal Reserve’s Jackson Hole symposium in August, when Chair Kevin Warsh asked: “Will token prices for older models fall to the level of their marginal cost?” September’s cuts suggest that process is already under way below the frontier.
Jev asks the same question one level deeper: what happens when the decision itself becomes a commodity?
Compute follows the split. Nvidia’s supply commitments rose from $119 billion to $279 billion in a quarter, “primarily memory and manufacturing facilities,” according to its Q2 FY2027 Form 10-Q. That buildout reflects a world in which large volumes of inference still pass through memory-intensive decode. On the outside, Jev’s architecture means a decision requires input processing and a readout without the generation phase.
That does not mean decision models reduce total compute demand; Jevons may well win again. The harder claim is about composition. If decision calls merely add to existing workloads, the infrastructure thesis barely changes. But if they replace a meaningful share of generated reasoning inside software loops, some demand shifts away from decode-heavy inference toward a different compute profile. AI usage can keep rising while the hardware mix changes underneath it.
Who captures the value when AI decisions get cheaper?
My hypothesis is that value shifts toward control of workflows and their records, and not automatically to the model maker or gateway. Some may instead become customer savings or remain with model suppliers.
My Orchestration Economics framework distinguishes three rings:
Ring 1, intelligence, is separating into a rapidly commoditizing decision tier and a deliberation tier that retains its premium.
Ring 2, the harness, manages that split, but the routing engine I called proprietary in June became open source in three days.
Ring 3, the orchestrator, poses questions, sets thresholds, and accumulates the record of decisions.
That record is not an explanation of the model’s internal reasoning. It documents what the system was asked, what it decided, and what happened next. That record follows the work: when a decision model runs inside the customer’s systems, the gateway need not own it. The distinction is between supplying a replaceable judgment and controlling the workflow that turns it into action.
Last week I argued that expertise inside software can be learned into weights, leaving the base model replaceable. The system running the workflow, however, retains its record. Or, as I put it then, the receipt becomes the source code. Jev’s rapid replication makes that replaceability concrete: what looked proprietary acquired alternatives within days.
In July, I described scarcity migrating toward the self-improving loop and margin moving outward from the model layer: “One fault line, one direction of travel.”
Jev adds a shift in what software buys-from generated tokens to structured decisions. The argument does not depend on TypeSafe surviving. It depends on whether that separation endures.
What would disprove the decision-model argument?
The argument weakens if adoption fades, decision quality disappoints, or the pricing proves unsustainable. If frontier labs supply the alternatives and keep the traffic, the split survives-but the claim about who captures its value changes. Four tests will determine which parts hold:
1. Does adoption outlast the launch? Vercel’s October and November Production Indexes are the first checkpoint. A collapse in Jev’s share of paid spending after its late-September free period would weaken the company’s case. If no alternative fills the role, the category’s case weakens too. Spending share is a signal, however, not a verdict on adoption: a cheap product need not capture much spending to find substantial use.
2. Does decision quality survive independent testing? By year-end, I want to see an independent evaluation using at least 1,000 held-out, human-labeled examples, comparing model families and testing beyond their training distributions. Decision quality should at least match a small language model. Without that evidence, the claim to a distinct product tier remains unproven; disappointing results would weaken it further.
3. Do the frontier labs keep both halves? If a lab ships a decision or calibrated-confidence endpoint, or bundles a free router, and retains the traffic, the separation has happened-but the labs own both sides. That challenges the argument about value capture, not the existence of the decision tier.
4. Does the cost advantage survive production? Over the next two quarters, decision models should retain a meaningful cost advantage at comparable service levels. An output charge or reduced rate limits would matter to the extent that they erode the total saving per reliable decision.
These tests need not deliver the same verdict. TypeSafe could fail while decision models thrive. The category could succeed while the frontier labs retain the value. Surviving the tests would strengthen the argument, not settle it.
The discontinuity does not care whether TypeSafe survives. If decisions remain a separate product priced toward zero, Jev will have repriced intelligence one layer below the token.
DISCLAIMER: The views and opinions expressed here are those of the author alone and are based on publicly available information. They do not constitute investment advice, a solicitation, or a recommendation to buy or sell any security or financial instrument. The author may hold positions in the securities of companies mentioned. Past performance is not indicative of future results. Readers should conduct their own independent due diligence and consult a qualified financial advisor before making any investment decision.

