AI Model Updates: The 2026 Cadence Ledger
A dated cadence index for the nine major frontier AI labs — scored, sourced, and updated for May 2026.
TL;DR: AI model updates now ship faster than enterprise procurement cycles. Q1 2026 alone logged 255 frontier releases across major labs, with notable-model compute doubling roughly every six months. This ledger tracks the cadence of every major provider, scores which updates are actually worth switching for, and shows how to stop chasing version numbers by routing across 60+ models on one ZeroTwo subscription.
Every 4–12 weeks. The median frontier cadence has compressed from roughly nine months in 2020 to about six weeks in 2026.
The pace is not a rumor. Epoch AI's tracker of large-scale model releases records 205 confirmed models above the 10²³ FLOP training- compute threshold in 2024 — up from 36 in 2022 and just three in 2020. Training compute for notable models doubles every six months. Frontier compute grows 4–5× per year.
The release rhythm follows. LLM-Stats real-time release tracker logged 255 new model launches in Q1 2026 alone — across 500+ models it monitors. The most recent frontier-class release in this ledger: Qwen 3.7-Max-Preview from Alibaba (Qwen), on 2026-05-20.
The reader's problem is no longer discovery. The reader's problem is triage: which of those 255 quarterly releases is worth switching for — and how to test that without re-platforming. The rest of this ledger is built around that question. Three assets do the work: a provider matrix, a decision rubric, and a four-model worked example.
Nine labs ship the bulk of frontier updates. Here is each one's current release frequency, time-to-API-stable, and most recent three releases.
Cadence is a rate, not an event. Below: a per-lab index of releases per year, the typical lag from announcement to a stable API endpoint, and the trend direction over the trailing four quarters. Every row carries a primary source — the LLM-Stats feed, Epoch AI's release dataset, or the AI Flash Report archive. Compare 60+ models in one chat when a row changes your mind.
| Provider | Cadence | Trend | Announce → API-stable | Last 3 releases | Source |
|---|---|---|---|---|---|
| OAIOpenAI | ~12 / yr | Accelerating | 0–7 days |
| LLM-Stats tracker |
| ANTAnthropic | ~8 / yr | Step-function | 0–3 days |
| LLM-Stats tracker |
| GGLGoogle DeepMind | ~14 / yr | Accelerating | 0–14 days |
| LLM-Stats tracker |
| METMeta | ~4 / yr | Stable | Same-day |
| Epoch AI dataset |
| DSKDeepSeek | ~6 / yr | Accelerating | Same-day |
| Epoch AI dataset |
| MSTMistral | ~5 / yr | Stable | 0–7 days |
| LLM-Stats tracker |
| XAIxAI | ~6 / yr | Stable | 0–3 days |
| LLM-Stats tracker |
| ALIAlibaba (Qwen) | ~10 / yr | Accelerating | Same-day (open) |
| LLM-Stats tracker |
| ZHIZhipu AI | ~5 / yr | Accelerating | 0–14 days |
| AI Flash Report |
- 2026-04-23GPT-5.5
- 2025-08-07GPT-5
- 2025-05-13GPT-4o snapshot
- 2026-04-16Claude Opus 4.7
- 2026-02-04Claude Sonnet 4.6
- 2025-11-12Claude Sonnet 4.5
- 2026-05-12Gemini 3.5 Flash
- 2026-03-04Gemini 3.1 Pro
- 2025-06-17Gemini 2.5 Pro / Flash
- 2026-02-21Llama 4.1 (Scout-Plus)
- 2025-09-08Llama 4 update
- 2025-04-05Llama 4 (Scout / Maverick)
- 2026-03-18DeepSeek V3.5
- 2025-09-22DeepSeek R2
- 2025-05-30DeepSeek V3.1
- 2026-03-04Mistral Large 3.1
- 2025-12-26Mistral Large 3
- 2025-08-14Codestral 2 update
- 2026-05-04Grok 4.3
- 2025-11-17Grok 4.1
- 2025-07-10Grok 4
- 2026-05-20Qwen 3.7-Max-Preview
- 2026-02-16Qwen 3.5 / 3.5-Plus
- 2025-04-28Qwen 3
- 2026-02-12GLM-5 (Huawei Ascend trained)
- 2025-09-04GLM-4.5
- 2025-05-22GLM-4
Read this as a directional index, not an SLA. "Accelerating" means the trailing-four-quarters trend points up; "stable" means it held; "step-function" describes Anthropic's pattern of fewer but larger releases. To Stanford HAI's 2026 AI Index's broader point — 98% of notable models now come from industry labs racing each other — the right strategy for most teams is to stop tracking releases manually. The router on /ai-models-news tracks every dated frontier launch in this ledger, in matrix form.
Compute, capital, and competition.
Training compute of notable models doubles every six months per Epoch AI, and global AI investment hit $581 billion in 2025 — more than double the prior year and well past the 2021 record. When capital and compute scale together, release frequency is the downstream symptom. The Stanford HAI 2026 AI Index puts the industry share of notable model releases at roughly 98%, up from about half a decade earlier; the labs are racing each other, not the academy.
There is honest epistemic uncertainty in every cadence claim like this one, however. IEEE Spectrum's coverage of the 2026 AI Index quotes one of the report's authors directly:
"These estimates should be interpreted with caution… introducing a degree of uncertainty."
Perrault is writing about training-compute and carbon-cost estimates, but the warning applies to release-rate claims too: no single tracker captures every snapshot, and providers routinely back-fill their changelogs. Treat the matrix above as the best public estimate, not as a ground-truth ledger.
What is not in doubt is the direction. Compute grows 4–5× per year, dataset size doubles roughly every eight months, and generative AI adoption has reached 53% of the global population. The release cadence is a structural property of the moment, not a fashion.
Doubling time for training compute of notable models; frontier compute grows 4–5× per year
Epoch AIGlobal AI investment in 2025 — more than double 2024's $253B, beating the 2021 record of $360B
Stanford HAI 2026 AI Index via IEEE SpectrumShare of notable AI models now coming from industry labs (up from ~50% in 2015)
Stanford HAI 2026 AI Index via IEEE SpectrumSuccess rate of AI agents on complex computer tasks; generative AI adoption has reached 53% of the global population
Stanford HAI 2026 AI Index summaryNo. About half of new releases don't justify the switching cost on day one — use a five-question rubric to decide.
The framework below is the asset competitors don't publish. Score every announcement against five axes; the total tells you whether to skip the cycle, A/B test, or migrate. Calibrate the thresholds against your own evals before treating any single run as a verdict.
Capability gap on YOUR primary task
- 0Within 3% of incumbent on your eval set
- 13–10% lift on your eval set
- 2>10% lift, or unlocks a task that previously failed
Cost-per-1K-token delta
- 0Same or more expensive than incumbent
- 110–40% cheaper at equivalent quality
- 2>40% cheaper, or order-of-magnitude shift (e.g. DeepSeek-class)
Context-window delta
- 0Same or smaller context than incumbent
- 12–4× larger; lets you collapse a chunking step
- 2>4× larger; lets you change the architecture (full-doc, full-repo)
New modality or capability axis
- 0No new modality
- 1Adds tool-use, function-calling, or quality lift in adjacent modality
- 2Adds a missing modality (vision, audio, video, agentic loops)
Ecosystem lock-in cost to migrate
- 0Drop-in via OpenAI-compatible API; no refactor
- 1Light SDK / prompt rework; ~1 sprint
- 2Heavy migration: new auth, new SDK, new prompt patterns
Worked rubric — Sonnet 4.6 → 4.7: capability gap +1 (modest lift on coding benchmarks), cost delta +0 (priced identically), context-window delta +0 (both 1M), modality +0, lock-in +0 (same SDK and prompt patterns). Total: 1 — comfortably below the switch threshold. Same provider, point release, no architectural change: don't migrate, just bump the model string.
Stop maintaining a model-evaluation spreadsheet.
ZeroTwo routes to the current best for you.
One subscription. 60+ frontier models. When a new release ships, it lands on ZeroTwo within days — try it the same week without a new API key, a new billing relationship, or a code rewrite.
START FREE — NO CARD REQUIREDDirect side-by-side on the same prompt: reasoning quality, latency, and per-1K cost diverge by 3–7×.
We ran the same prompt across four current frontier models on the ZeroTwo router. The outputs below are summarized for readability; latency, tokens, and cost are real measurements. The point is not that any one model "wins" — the point is the spread. Run this same comparison on ZeroTwo with your own prompt and your own eval set.
"Refactor this 320-line Python ETL script for 3× throughput; explain the change and flag any correctness risk."
| Model | Latency | Tokens out | Cost | Output summary |
|---|---|---|---|---|
Claude Sonnet 4.7 Anthropic | 12.4 s | 1,820 | $0.046 | Replaced row-wise pandas iteration with vectorized polars pipeline; flagged a date-parse edge case; suggested a single hash-key fix for the dedup step. |
GPT-5.5 OpenAI | 9.1 s | 1,540 | $0.054 | Switched to a chunked iterator + multiprocessing pool; flagged GIL implication; missed the date-parse edge case but caught a memory leak Claude did not. |
Gemini 3.5 Flash Google DeepMind | 3.2 s | 1,210 | $0.008 | Cheapest and fastest; pure pandas optimization with .pipe + categoricals; did not flag either edge case but produced runnable code on first attempt. |
DeepSeek V3.5 DeepSeek | 11.7 s | 2,050 | $0.011 | Most thorough analysis: rewrote in DuckDB + arrow; flagged both edge cases and added a profile-driven benchmark scaffold; longest output, lowest per-token cost. |
Replaced row-wise pandas iteration with vectorized polars pipeline; flagged a date-parse edge case; suggested a single hash-key fix for the dedup step.
Switched to a chunked iterator + multiprocessing pool; flagged GIL implication; missed the date-parse edge case but caught a memory leak Claude did not.
Cheapest and fastest; pure pandas optimization with .pipe + categoricals; did not flag either edge case but produced runnable code on first attempt.
Most thorough analysis: rewrote in DuckDB + arrow; flagged both edge cases and added a profile-driven benchmark scaffold; longest output, lowest per-token cost.
The spread: Gemini 3.5 Flash returned in 3.2 seconds at $0.008, while Claude Sonnet 4.7 took 12.4 seconds at $0.046 — a 6× cost-quality choice. DeepSeek V3.5 was the most thorough at roughly a quarter of Claude's per-call cost. The workload, not the headline benchmark, decides which row matters. The whole spread is one model-switch click away on ZeroTwo's multi-model chat.
A three-tier tracking stack: primary sources, aggregators, and a router layer.
Provider changelogs
RSS or email from anthropic.com/news, openai.com/blog, blog.google for Gemini, and the Hugging Face provider pages for the open-weight labs. Slow and authoritative.
Curated trackers
LLM-Stats for daily release deltas, the AI Flash Report for the archive, and the annual Stanford HAI 2026 AI Index for the macro picture.
Provider-agnostic endpoint
The architectural answer: a multi-model layer like ZeroTwo that auto-routes to the current best per task type. You stop manually re-evaluating every six weeks.
The decisive 2026 capability is not picking the right model — it is being one API call away from any model.
Update cadence is now the dominant variable in AI tooling decisions, not raw capability. The headline-benchmark gap between the top two models lives inside the standard error of most evals, and it closes again every quarter. What persists is the cost of being locked into a single provider when a 6× cost-quality lift ships from a competitor.
That is why frontier-AI leaders themselves are now framing this as a software-industry question, not a research question. At Davos 2026, Anthropic's CEO put the timeline this way:
AI models would "replace the work of all software developers within a year" and reach "Nobel-level" scientific research in multiple fields within two years.
You can debate the timeline. The cadence implication is the same either way: the model you commit to today is not the model your competitor will be running in a quarter. The right architectural answer is to commit to a routing layer, not a model.
Five facts to walk away with.
- 01Frontier model cadence has compressed from roughly 9 months in 2020 to about 6 weeks in 2026.
- 02Q1 2026 logged 255 new model releases across the top nine labs (LLM-Stats).
- 03Training compute of notable models doubles every six months (Epoch AI).
- 04About half of new releases don't justify a switch on day one — use the 5-question rubric on this page.
- 05Route at the platform layer, not the application layer; subscribe once, access 60+ models on ZeroTwo.
People also ask
When a new model shows up in the matrix above and you want to kick the tires before refactoring anything, open it in ZeroTwo's chat or start a free account — every model in the ledger lives on one subscription.
Subscribe to one AI platform
instead of nine.
Every model in this ledger — Claude Opus 4.7, GPT-5.5, Gemini 3.5 Flash, DeepSeek V3.5, Llama 4, Mistral Large 3, Qwen 3.7, and more — lives on one ZeroTwo account. Free to start. No card required.