Skip to content
Go back

Benchmark saturation is real. The price collapse is an illusion.

Two claims dominate conversation about the AI model market right now: benchmark saturation is killing differentiation, and prices are in freefall. Both need precision to mean anything. One is real but framed wrong. The other is a flat measurement error.

Finding 1: Saturation is compression, not stagnation

The frontier is clustering. In the last 8 weeks, the top 10 models’ intelligence-index standard deviation narrowed from 3.53 to 1.72—a 51% tightening. That’s real. More models are reaching capability levels once held by a few.

But here’s what the “saturation” narrative skips: the top-10 mean rose from 52.7 to 60.6 during that same window. A 15% absolute gain while the spread collapsed. The frontier isn’t plateauing—it’s compressing while climbing.

This is convergence, not stagnation. More labs are building models that actually work. That creates fierce competition, margin pressure, and a crowded leaderboard. But the bar is still rising.

The misframing matters because it feeds bad decisions. If you believe capability is saturated, you cut R&D. The data says the opposite: the frontier is rising; you just have more competition at every level.

Dual-axis chart: red line (top-10 mean) climbs from 52.7 to 60.6. Blue line (top-10 std dev) drops from 3.53 to 1.72. Green line (overall median) rises from 13.4 to 16.1 across all models.

The clustering is real (blue line down), but the frontier is still rising (red and green lines up).

Finding 2: Price stability, masked by composition

Median price in published API catalogs did rise from $0.835 to $1.10 between May and August. That looks like a spike. It’s not. It’s an optical illusion.

The catalog grew by 15% over that window. The new entrants are 2.35x pricier ($2.00 median) than the existing cohort ($0.85 median).

When I matched the 309 models that appeared in both snapshots, the median price was $0.85 at both dates. Zero change.

Distribution within the matched set:

What actually happened: the catalog absorbed a wave of premium new entrants. The “price collapse” framing came from computing a naive cross-section across an expanding list. Model-for-model, existing prices are flat to slightly cheaper.

Left: histogram of price % changes for the 309 matched models. Median bar at 0%. Right: overlaid histograms—existing cohort (n=309) peaked at $0.85, new entrants (n=77) shifted right to $2.00 median.

Most models didn’t move. New entrants pulled the overall median up.

Why this matters

If benchmark saturation means innovation is dead, you make the wrong bet on your infrastructure. The data says the opposite.

If prices are crashing, you panic-cut margins or rush to open-weight. The data shows prices are stable for existing models; new models are just more expensive. That’s stratification, not a race to the bottom. The competitive story is different: you’re not losing to cheaper commodities yet. You’re losing to tiers—premium reasoning models for complex problems, competent fast models for interactive work, cheap models for batch.

The open question: are the new entrants staying expensive, or does the median gap narrow over time? That’s worth tracking on the next snapshot.


Share this post on:


Next Post
Dictation and Meeting Notes, Fully Local: Rebuilding Wispr Flow and Granola Without the Cloud