Log in
zerotwo / release-wire
LIVE · UPDATED 2026-05-21
◆ ISSUE 14Frontier Release Tracker · 2026

AI Models News:
Frontier Release Tracker & 2026 Matrix

Frontier AI models now ship every 3–6 weeks. This is the dated matrix of every major launch from DeepSeek V3 in December 2024 through Qwen 3.7-Max-Preview in May 2026 — provider, benchmark claim, license, and the source link for every row. ZeroTwo unifies all twelve in one chat.

12 dated releases8 providers60+ models on ZeroTwoBiweekly refresh
Last 90-day pulse
Most recent launch
Qwen 3.7-Max-Preview
May 20, 2026 · Alibaba Cloud
Top MMLU score
Gemini 3.1 Pro · 94.3%
Source: LLM-Stats Leaderboard
Gap, #1 vs #2
0.7 pp
Source: Stanford HAI 2025
DEC 26 · 2024·DSKDeepSeek V3·671B · 37B active·128K ctx·Open
APR 05 · 2025·METLlama 4 (Scout / Maverick)·17B active · 128 experts·10M (Scout) ctx·Open
APR 28 · 2025·ALIQwen 3·235B MoE · 22B active·128K ctx·Apache 2.0
JUN 17 · 2025·GGLGemini 2.5 Pro & Flash·n/d (closed)·2M ctx·API
JUL 10 · 2025·XAIGrok 4·n/d (closed)·256K ctx·API
AUG 07 · 2025·OAIGPT-5·n/d (closed)·400K ctx·API
NOV 17 · 2025·XAIGrok 4.1·n/d (closed)·256K ctx·API
DEC 26 · 2025·MSTMistral Large 3·MoE (undisclosed mix)·256K ctx·API + research
FEB 16 · 2026·ALIQwen 3.5 & Qwen 3.5-Plus·MoE (undisclosed mix)·128K ctx·Apache 2.0 (3.5)
APR 16 · 2026·ANTClaude Opus 4.7·n/d (closed)·1M ctx·API
APR 23 · 2026·OAIGPT-5.5·n/d (closed)·400K ctx·API
MAY 20 · 2026·ALIQwen 3.7-Max-Preview·n/d (preview)·128K ctx·API preview
DEC 26 · 2024·DSKDeepSeek V3·671B · 37B active·128K ctx·Open
APR 05 · 2025·METLlama 4 (Scout / Maverick)·17B active · 128 experts·10M (Scout) ctx·Open
APR 28 · 2025·ALIQwen 3·235B MoE · 22B active·128K ctx·Apache 2.0
JUN 17 · 2025·GGLGemini 2.5 Pro & Flash·n/d (closed)·2M ctx·API
JUL 10 · 2025·XAIGrok 4·n/d (closed)·256K ctx·API
AUG 07 · 2025·OAIGPT-5·n/d (closed)·400K ctx·API
NOV 17 · 2025·XAIGrok 4.1·n/d (closed)·256K ctx·API
DEC 26 · 2025·MSTMistral Large 3·MoE (undisclosed mix)·256K ctx·API + research
FEB 16 · 2026·ALIQwen 3.5 & Qwen 3.5-Plus·MoE (undisclosed mix)·128K ctx·Apache 2.0 (3.5)
APR 16 · 2026·ANTClaude Opus 4.7·n/d (closed)·1M ctx·API
APR 23 · 2026·OAIGPT-5.5·n/d (closed)·400K ctx·API
MAY 20 · 2026·ALIQwen 3.7-Max-Preview·n/d (preview)·128K ctx·API preview
01Latest frontier releases

May 2026: three frontier launches in four weeks

The last 35 days delivered three frontier releases: Claude Opus 4.7 on April 16, GPT-5.5 on April 23, and Qwen 3.7-Max-Preview on May 20. That cadence — roughly one frontier launch every 10–14 days across the eight main labs — is the new normal. The Stanford HAI 2026 AI Index reports training compute now doubles every five months, datasets every eight months, and roughly 90% of notable models come from industry — up from ~60% in 2023. Read this section first; the dated matrix below is the long form.

02Month-over-month matrix

Every frontier release, dated and sourced

Twelve dated releases from December 2024 through May 2026 — the full frontier wire across OpenAI, Anthropic, Google DeepMind, Meta, xAI, Alibaba, DeepSeek, and Mistral. Each row links to the primary source (provider announcement, release blog, or tracker). Reasoning and coding bars are normalized composite scores against published benchmark families (MMLU, GPQA, SWE-bench, HumanEval, MATH 500) — useful for scanning, never a substitute for a workload-specific eval.

Filter:
DateProviderModelLicenseParamsCtxReasoningCodingSource
MAY 20 · 2026ALI
Alibaba
Qwen 3.7-Max-Preview
Announced at Alibaba Cloud Summit; API rollout on Alibaba Cloud; preview-only — no open weights released yet.
API previewn/d (preview)128K
R
88
C
84
Codersera
APR 23 · 2026OAI
OpenAI
GPT-5.5
API availability from April 24; incremental gains on reasoning, code, math; same agent and tool-use scaffolding as GPT-5.
APIn/d (closed)400K
R
92
C
91
OpenAI
APR 16 · 2026ANT
Anthropic
Claude Opus 4.7
13% lift on coding evals, 3× more production tasks resolved than prior Opus; 1M-token context; $5 / $25 per 1M input/output tokens.
APIn/d (closed)1M
R
93
C
92
Anthropic
FEB 16 · 2026ALI
Alibaba
Qwen 3.5 & Qwen 3.5-Plus
Incremental update over Qwen 3; multilingual gains, hybrid reasoning carried forward; Plus tier on Alibaba Cloud API.
Apache 2.0 (3.5)MoE (undisclosed mix)128K
R
86
C
82
Alibaba Cloud
DEC 26 · 2025MST
Mistral
Mistral Large 3
Most capable Mistral to date; MoE architecture; European data-residency story remains a key positioning lever in EU deployments.
API + researchMoE (undisclosed mix)256K
R
84
C
82
BentoML guide
NOV 17 · 2025XAI
xAI
Grok 4.1
Incremental reasoning and coding improvements over Grok 4; positioned as the daily-driver xAI tier ahead of Grok 4.3.
APIn/d (closed)256K
R
87
C
85
Manifold tracker
AUG 07 · 2025OAI
OpenAI
GPT-5
Initial release across ChatGPT, API, and GitHub Models Playground; significant lift on reasoning over GPT-4 generation.
APIn/d (closed)400K
R
91
C
90
ChatGPT timeline
JUL 10 · 2025XAI
xAI
Grok 4
Flagship reasoning push; gains on factuality and code; wider rollout reached general availability through 2026 with the 4.3 update.
APIn/d (closed)256K
R
86
C
83
xAI announcement
JUN 17 · 2025GGL
Google
Gemini 2.5 Pro & Flash
Public GA; thinking mode on Pro; ~2× inference speedup vs 2.0; native multimodal across text, vision, audio.
APIn/d (closed)2M
R
89
C
86
Google blog
APR 28 · 2025ALI
Alibaba
Qwen 3
Six dense plus two MoE variants under Apache 2.0; 30B-A3B and 235B-A22B configurations; hybrid reasoning toggle.
Apache 2.0235B MoE · 22B active128K
R
84
C
79
Alibaba Cloud
APR 05 · 2025MET
Meta
Llama 4 (Scout / Maverick)
First MoE Llama; Maverick beats GPT-4o and Gemini 2.0 Flash on benchmarks; Scout claims a 10M-token context window.
Open17B active · 128 experts10M (Scout)
R
78
C
81
TechCrunch
DEC 26 · 2024DSK
DeepSeek
DeepSeek V3
Matches frontier performance at $5.6M training cost; 671B-param MoE with 37B active per token; competitive on SWE-Bench Verified and MATH 500.
Open671B · 37B active128K
R
82
C
87
Interconnects
MAY 20 · 2026ALI
Qwen 3.7-Max-Preview
Alibaba · API preview · 128K ctx
Announced at Alibaba Cloud Summit; API rollout on Alibaba Cloud; preview-only — no open weights released yet.
R
88
C
84
Codersera
APR 23 · 2026OAI
GPT-5.5
OpenAI · API · 400K ctx
API availability from April 24; incremental gains on reasoning, code, math; same agent and tool-use scaffolding as GPT-5.
R
92
C
91
OpenAI
APR 16 · 2026ANT
Claude Opus 4.7
Anthropic · API · 1M ctx
13% lift on coding evals, 3× more production tasks resolved than prior Opus; 1M-token context; $5 / $25 per 1M input/output tokens.
R
93
C
92
Anthropic
FEB 16 · 2026ALI
Qwen 3.5 & Qwen 3.5-Plus
Alibaba · Apache 2.0 (3.5) · 128K ctx
Incremental update over Qwen 3; multilingual gains, hybrid reasoning carried forward; Plus tier on Alibaba Cloud API.
R
86
C
82
Alibaba Cloud
DEC 26 · 2025MST
Mistral Large 3
Mistral · API + research · 256K ctx
Most capable Mistral to date; MoE architecture; European data-residency story remains a key positioning lever in EU deployments.
R
84
C
82
BentoML guide
NOV 17 · 2025XAI
Grok 4.1
xAI · API · 256K ctx
Incremental reasoning and coding improvements over Grok 4; positioned as the daily-driver xAI tier ahead of Grok 4.3.
R
87
C
85
Manifold tracker
AUG 07 · 2025OAI
GPT-5
OpenAI · API · 400K ctx
Initial release across ChatGPT, API, and GitHub Models Playground; significant lift on reasoning over GPT-4 generation.
R
91
C
90
ChatGPT timeline
JUL 10 · 2025XAI
Grok 4
xAI · API · 256K ctx
Flagship reasoning push; gains on factuality and code; wider rollout reached general availability through 2026 with the 4.3 update.
R
86
C
83
xAI announcement
JUN 17 · 2025GGL
Gemini 2.5 Pro & Flash
Google · API · 2M ctx
Public GA; thinking mode on Pro; ~2× inference speedup vs 2.0; native multimodal across text, vision, audio.
R
89
C
86
Google blog
APR 28 · 2025ALI
Qwen 3
Alibaba · Apache 2.0 · 128K ctx
Six dense plus two MoE variants under Apache 2.0; 30B-A3B and 235B-A22B configurations; hybrid reasoning toggle.
R
84
C
79
Alibaba Cloud
APR 05 · 2025MET
Llama 4 (Scout / Maverick)
Meta · Open · 10M (Scout) ctx
First MoE Llama; Maverick beats GPT-4o and Gemini 2.0 Flash on benchmarks; Scout claims a 10M-token context window.
R
78
C
81
TechCrunch
DEC 26 · 2024DSK
DeepSeek V3
DeepSeek · Open · 128K ctx
Matches frontier performance at $5.6M training cost; 671B-param MoE with 37B active per token; competitive on SWE-Bench Verified and MATH 500.
R
82
C
87
Interconnects
Stop juggling 8 subscriptions
Every model in the matrix lives on one ZeroTwo account.
Start free
03Provider spotlights

Eight labs. Eight release cadences.

Each provider has carved out a distinctive release pattern. OpenAI ships step-changes; Anthropic prefers narrow lifts on Opus; Google maintains a Pro/Flash split with rapid Flash refreshes; Meta and DeepSeek anchor the open-weight frontier; Alibaba runs quarterly increments capped with API-only previews; xAI iterates in tenths; Mistral stays the European data-residency story. Spotlight summaries below, each with a primary source.

OAI

OpenAI

~9 months between flagship steps

GPT-5 arrived August 2025; GPT-5.5 dropped April 2026 with incremental reasoning and code gains while keeping the same 400K context window and agent/tool-use scaffolding. Tooling remains the strongest of any closed lab.

ChatGPT timeline · enter.pro
ANT

Anthropic

Step-function lifts, narrow cadence

Claude Opus 4.7 (April 2026) cites a 13% lift on coding evals and resolves roughly 3× more production tasks than its predecessor, with a 1M-token context window and explicit pricing of $5 / $25 per million input/output tokens.

Anthropic announcement
GGL

Google DeepMind

Pro + Flash split, rapid Flash refresh

Gemini 2.5 Pro and Flash reached general availability in June 2025 with thinking mode, ~2× faster inference than Gemini 2.0, and the largest mainstream context window at 2M tokens. May 2026 brought further Ultra-tier and Spark-Omni rollouts on the consumer side.

Google AI blog
MET

Meta

Single open-weight flagship per year

Llama 4 (Scout, Maverick) shipped April 2025 as the first MoE Llama family. Maverick claimed wins over GPT-4o and Gemini 2.0 Flash on multiple benchmarks; Scout headlined a 10M-token context window, the largest open-weight context published to date.

TechCrunch
ALI

Alibaba (Qwen)

Quarterly increments + preview tiers

Qwen 3 (April 2025) shipped six dense and two MoE variants under Apache 2.0. Qwen 3.5 followed in February 2026, and Qwen 3.7-Max-Preview was announced at the Alibaba Cloud Summit on May 20, 2026 — API-only on Alibaba Cloud, no open weights yet.

Alibaba Cloud
DSK

DeepSeek

Cost-defining open releases

DeepSeek V3 (December 2024) reset the conversation around training economics — frontier-grade performance at a claimed $5.6M training cost using a 671B-parameter MoE with 37B active per token. The model continues to drive open-weight cost benchmarks into 2026.

Interconnects
XAI

xAI

Stepwise 4 → 4.1 → 4.3

Grok 4 shipped July 2025; Grok 4.1 followed in November 2025 with incremental reasoning and coding gains. The 4.3 rollout in May 2026 widened availability across xAI's app and API surfaces.

xAI
MST

Mistral

Annual flagship + open mid-tier

Mistral Large 3 landed in December 2025 as the most capable model in the family to date, retaining the European data-residency narrative that drives much of its EU deployment story. Mid-tier weights continue to ship under permissive licenses.

BentoML guide
04Benchmark convergence

The numbers behind the wire

Frontier benchmark scores are converging — every flagship now clears 88% on MMLU and 93%+ on HumanEval, so legacy multiple-choice tests no longer discriminate. The real contest has shifted to expert-reasoning evals (GPQA Diamond, Humanity's Last Exam) and agentic engineering benchmarks (SWE-bench, MLe-bench). Seven of the most-cited 2025–2026 numbers, each with a primary source.

94.3%
Gemini 3.1 Pro MMLU score — every frontier model now clears 88%
LLM-Stats Leaderboard 2026
0.7%
Gap between #1 and #2 models — down from 4.9% one year prior (Stanford HAI)
Stanford HAI, 2025 AI Index Report
+67.3 pp
SWE-bench improvement on real software-engineering tasks (2023→2024)
Stanford HAI, 2025 AI Index Report
+30% YoY
Reasoning lift on Humanity's Last Exam (HLE) — 2,500 frontier questions across math, science, languages
Edge AI Vision · May 2026
Every 5 mo
Training compute doubling time — datasets double every 8 months; ~90% of notable models now from industry
Stanford HAI, 2026 AI Index Report
94.6%
GPQA Diamond leader (PhD-level science Q&A) — non-expert PhDs score ~34%
LXT.ai, LLM Benchmarks 2026
65%
Success rate on MLe-bench (hour-long ML engineering tasks) — up from ~0% in 2023
Edge AI Vision · May 2026
05How ZeroTwo's lineup compares

Every model in the matrix, one subscription

ZeroTwo unifies more than 60 frontier models in a single chat. The full matrix above — GPT-5.5, Claude Opus 4.7, Gemini 2.5 Pro, Grok 4.1, Llama 4 (Scout / Maverick), Qwen 3 and 3.5, DeepSeek V3, and Mistral Large 3 — is accessible on one account. Switch models turn-by-turn inside the same thread; the conversation history travels with you. For context on how that compares vs single-vendor stacks, see our best AI platforms 2026 scorecard and the head-to-head AI models for text generation comparison. For the full tool catalog including image, video, research, and code, browse our AI tools directory.

Free
$0forever
  • Rotating access to frontier models
  • Daily message limits
  • Multi-model picker
Pro · MOST POPULAR
$29.99/ month
  • 60+ frontier models
  • Image, video, deep research
  • All tools included
Pro 2x
$59.98/ month
  • 2× request limits vs Pro
  • Higher-volume agent workflows
  • Multi-model side-by-side runs
Ultra
$120/ month
  • Highest request limits
  • Team and research workloads
  • Priority access to new models
06What to expect next

Mid-2026 outlook: reasoning, MoE economics, calibration

Three patterns will define the next 90 days. First, reasoning models that scale inference-time compute — Claude Mythos Preview already leads GPQA Diamond at 94.6%, and expect more o-series-style architectures from every major lab. Second, MoE cost economics will keep compressing the open-weight frontier behind the DeepSeek V3 cost curve — training compute is doubling every five months, but per-token inference cost is falling. Third, benchmark saturation forces the field toward harder evals: Humanity's Last Exam, agentic SWE-bench, and MLe-bench replace MMLU as the meaningful signal. The risk Edge AI Vision flags is calibration: even top models still report confident-sounding wrong answers.

"Even highly capable reasoning models can fail to reliably know when they do not know, and their errors can arrive with the same confident tone and high internal probability as correct answers."
Edge AI and Vision Alliance · May 2026 · source
07Key takeaways

The five things to remember

  • 01Frontier launches now arrive every 3–6 weeks across the eight main labs. Twelve frontier releases in the 17 months from December 2024 to May 2026.
  • 02Benchmark gaps have collapsed. The Stanford HAI 2025 AI Index reports a 0.7 pp gap between #1 and #2 — choose by task, not by leaderboard rank.
  • 03Open-weight frontier is real. DeepSeek V3, Llama 4, and Qwen 3 trade benchmarks with closed flagships at a fraction of training cost.
  • 04Context windows are diverging. Llama 4 Scout claims 10M tokens; Gemini 2.5 Pro ships 2M; Claude Opus 4.7 ships 1M; GPT-5.5 stays at 400K.
  • 05The cheapest way to keep up is to stop committing. ZeroTwo unifies 60+ frontier models on one subscription so switching models becomes a click, not a credit-card decision.
08FAQ · People also ask

Frequently asked questions

Where can I try the latest frontier models?

Every model in the matrix — GPT-5.5, Claude Opus 4.7, Gemini 2.5 Pro, Grok 4.1, Llama 4, Qwen 3 & 3.5, DeepSeek V3, Mistral Large 3 — runs on ZeroTwo with one account. Free tier ships rotating frontier-model access; Pro at $29.99/mo unlocks the full lineup; Pro 2x at $59.98/mo doubles request limits; Ultra at $120/mo handles team and research workloads.

Stop reading model news.
Start using all of them.

Every frontier model in the matrix — twelve dated releases and counting — lives on one ZeroTwo account. Free to start. No card required.

60+ models·Free / $29.99 / $59.98 / $120·Published 2026-05-21 · Updated 2026-05-21