AI Models News:
Frontier Release Tracker & 2026 Matrix
Frontier AI models now ship every 3–6 weeks. This is the dated matrix of every major launch from DeepSeek V3 in December 2024 through Qwen 3.7-Max-Preview in May 2026 — provider, benchmark claim, license, and the source link for every row. ZeroTwo unifies all twelve in one chat.
May 2026: three frontier launches in four weeks
The last 35 days delivered three frontier releases: Claude Opus 4.7 on April 16, GPT-5.5 on April 23, and Qwen 3.7-Max-Preview on May 20. That cadence — roughly one frontier launch every 10–14 days across the eight main labs — is the new normal. The Stanford HAI 2026 AI Index reports training compute now doubles every five months, datasets every eight months, and roughly 90% of notable models come from industry — up from ~60% in 2023. Read this section first; the dated matrix below is the long form.
Every frontier release, dated and sourced
Twelve dated releases from December 2024 through May 2026 — the full frontier wire across OpenAI, Anthropic, Google DeepMind, Meta, xAI, Alibaba, DeepSeek, and Mistral. Each row links to the primary source (provider announcement, release blog, or tracker). Reasoning and coding bars are normalized composite scores against published benchmark families (MMLU, GPQA, SWE-bench, HumanEval, MATH 500) — useful for scanning, never a substitute for a workload-specific eval.
| Date | Provider | Model | License | Params | Ctx | Reasoning | Coding | Source |
|---|---|---|---|---|---|---|---|---|
| MAY 20 · 2026 | ALI Alibaba | Qwen 3.7-Max-Preview Announced at Alibaba Cloud Summit; API rollout on Alibaba Cloud; preview-only — no open weights released yet. | API preview | n/d (preview) | 128K | R 88 | C 84 | Codersera |
| APR 23 · 2026 | OAI OpenAI | GPT-5.5 API availability from April 24; incremental gains on reasoning, code, math; same agent and tool-use scaffolding as GPT-5. | API | n/d (closed) | 400K | R 92 | C 91 | OpenAI |
| APR 16 · 2026 | ANT Anthropic | Claude Opus 4.7 13% lift on coding evals, 3× more production tasks resolved than prior Opus; 1M-token context; $5 / $25 per 1M input/output tokens. | API | n/d (closed) | 1M | R 93 | C 92 | Anthropic |
| FEB 16 · 2026 | ALI Alibaba | Qwen 3.5 & Qwen 3.5-Plus Incremental update over Qwen 3; multilingual gains, hybrid reasoning carried forward; Plus tier on Alibaba Cloud API. | Apache 2.0 (3.5) | MoE (undisclosed mix) | 128K | R 86 | C 82 | Alibaba Cloud |
| DEC 26 · 2025 | MST Mistral | Mistral Large 3 Most capable Mistral to date; MoE architecture; European data-residency story remains a key positioning lever in EU deployments. | API + research | MoE (undisclosed mix) | 256K | R 84 | C 82 | BentoML guide |
| NOV 17 · 2025 | XAI xAI | Grok 4.1 Incremental reasoning and coding improvements over Grok 4; positioned as the daily-driver xAI tier ahead of Grok 4.3. | API | n/d (closed) | 256K | R 87 | C 85 | Manifold tracker |
| AUG 07 · 2025 | OAI OpenAI | GPT-5 Initial release across ChatGPT, API, and GitHub Models Playground; significant lift on reasoning over GPT-4 generation. | API | n/d (closed) | 400K | R 91 | C 90 | ChatGPT timeline |
| JUL 10 · 2025 | XAI xAI | Grok 4 Flagship reasoning push; gains on factuality and code; wider rollout reached general availability through 2026 with the 4.3 update. | API | n/d (closed) | 256K | R 86 | C 83 | xAI announcement |
| JUN 17 · 2025 | GGL Google | Gemini 2.5 Pro & Flash Public GA; thinking mode on Pro; ~2× inference speedup vs 2.0; native multimodal across text, vision, audio. | API | n/d (closed) | 2M | R 89 | C 86 | Google blog |
| APR 28 · 2025 | ALI Alibaba | Qwen 3 Six dense plus two MoE variants under Apache 2.0; 30B-A3B and 235B-A22B configurations; hybrid reasoning toggle. | Apache 2.0 | 235B MoE · 22B active | 128K | R 84 | C 79 | Alibaba Cloud |
| APR 05 · 2025 | MET Meta | Llama 4 (Scout / Maverick) First MoE Llama; Maverick beats GPT-4o and Gemini 2.0 Flash on benchmarks; Scout claims a 10M-token context window. | Open | 17B active · 128 experts | 10M (Scout) | R 78 | C 81 | TechCrunch |
| DEC 26 · 2024 | DSK DeepSeek | DeepSeek V3 Matches frontier performance at $5.6M training cost; 671B-param MoE with 37B active per token; competitive on SWE-Bench Verified and MATH 500. | Open | 671B · 37B active | 128K | R 82 | C 87 | Interconnects |
Eight labs. Eight release cadences.
Each provider has carved out a distinctive release pattern. OpenAI ships step-changes; Anthropic prefers narrow lifts on Opus; Google maintains a Pro/Flash split with rapid Flash refreshes; Meta and DeepSeek anchor the open-weight frontier; Alibaba runs quarterly increments capped with API-only previews; xAI iterates in tenths; Mistral stays the European data-residency story. Spotlight summaries below, each with a primary source.
OpenAI
GPT-5 arrived August 2025; GPT-5.5 dropped April 2026 with incremental reasoning and code gains while keeping the same 400K context window and agent/tool-use scaffolding. Tooling remains the strongest of any closed lab.
ChatGPT timeline · enter.proAnthropic
Claude Opus 4.7 (April 2026) cites a 13% lift on coding evals and resolves roughly 3× more production tasks than its predecessor, with a 1M-token context window and explicit pricing of $5 / $25 per million input/output tokens.
Anthropic announcementGoogle DeepMind
Gemini 2.5 Pro and Flash reached general availability in June 2025 with thinking mode, ~2× faster inference than Gemini 2.0, and the largest mainstream context window at 2M tokens. May 2026 brought further Ultra-tier and Spark-Omni rollouts on the consumer side.
Google AI blogMeta
Llama 4 (Scout, Maverick) shipped April 2025 as the first MoE Llama family. Maverick claimed wins over GPT-4o and Gemini 2.0 Flash on multiple benchmarks; Scout headlined a 10M-token context window, the largest open-weight context published to date.
TechCrunchAlibaba (Qwen)
Qwen 3 (April 2025) shipped six dense and two MoE variants under Apache 2.0. Qwen 3.5 followed in February 2026, and Qwen 3.7-Max-Preview was announced at the Alibaba Cloud Summit on May 20, 2026 — API-only on Alibaba Cloud, no open weights yet.
Alibaba CloudDeepSeek
DeepSeek V3 (December 2024) reset the conversation around training economics — frontier-grade performance at a claimed $5.6M training cost using a 671B-parameter MoE with 37B active per token. The model continues to drive open-weight cost benchmarks into 2026.
InterconnectsxAI
Grok 4 shipped July 2025; Grok 4.1 followed in November 2025 with incremental reasoning and coding gains. The 4.3 rollout in May 2026 widened availability across xAI's app and API surfaces.
xAIMistral
Mistral Large 3 landed in December 2025 as the most capable model in the family to date, retaining the European data-residency narrative that drives much of its EU deployment story. Mid-tier weights continue to ship under permissive licenses.
BentoML guideThe numbers behind the wire
Frontier benchmark scores are converging — every flagship now clears 88% on MMLU and 93%+ on HumanEval, so legacy multiple-choice tests no longer discriminate. The real contest has shifted to expert-reasoning evals (GPQA Diamond, Humanity's Last Exam) and agentic engineering benchmarks (SWE-bench, MLe-bench). Seven of the most-cited 2025–2026 numbers, each with a primary source.
Every model in the matrix, one subscription
ZeroTwo unifies more than 60 frontier models in a single chat. The full matrix above — GPT-5.5, Claude Opus 4.7, Gemini 2.5 Pro, Grok 4.1, Llama 4 (Scout / Maverick), Qwen 3 and 3.5, DeepSeek V3, and Mistral Large 3 — is accessible on one account. Switch models turn-by-turn inside the same thread; the conversation history travels with you. For context on how that compares vs single-vendor stacks, see our best AI platforms 2026 scorecard and the head-to-head AI models for text generation comparison. For the full tool catalog including image, video, research, and code, browse our AI tools directory.
- Rotating access to frontier models
- Daily message limits
- Multi-model picker
- 60+ frontier models
- Image, video, deep research
- All tools included
- 2× request limits vs Pro
- Higher-volume agent workflows
- Multi-model side-by-side runs
- Highest request limits
- Team and research workloads
- Priority access to new models
Mid-2026 outlook: reasoning, MoE economics, calibration
Three patterns will define the next 90 days. First, reasoning models that scale inference-time compute — Claude Mythos Preview already leads GPQA Diamond at 94.6%, and expect more o-series-style architectures from every major lab. Second, MoE cost economics will keep compressing the open-weight frontier behind the DeepSeek V3 cost curve — training compute is doubling every five months, but per-token inference cost is falling. Third, benchmark saturation forces the field toward harder evals: Humanity's Last Exam, agentic SWE-bench, and MLe-bench replace MMLU as the meaningful signal. The risk Edge AI Vision flags is calibration: even top models still report confident-sounding wrong answers.
"Even highly capable reasoning models can fail to reliably know when they do not know, and their errors can arrive with the same confident tone and high internal probability as correct answers."
The five things to remember
- 01Frontier launches now arrive every 3–6 weeks across the eight main labs. Twelve frontier releases in the 17 months from December 2024 to May 2026.
- 02Benchmark gaps have collapsed. The Stanford HAI 2025 AI Index reports a 0.7 pp gap between #1 and #2 — choose by task, not by leaderboard rank.
- 03Open-weight frontier is real. DeepSeek V3, Llama 4, and Qwen 3 trade benchmarks with closed flagships at a fraction of training cost.
- 04Context windows are diverging. Llama 4 Scout claims 10M tokens; Gemini 2.5 Pro ships 2M; Claude Opus 4.7 ships 1M; GPT-5.5 stays at 400K.
- 05The cheapest way to keep up is to stop committing. ZeroTwo unifies 60+ frontier models on one subscription so switching models becomes a click, not a credit-card decision.
Frequently asked questions
Every model in the matrix — GPT-5.5, Claude Opus 4.7, Gemini 2.5 Pro, Grok 4.1, Llama 4, Qwen 3 & 3.5, DeepSeek V3, Mistral Large 3 — runs on ZeroTwo with one account. Free tier ships rotating frontier-model access; Pro at $29.99/mo unlocks the full lineup; Pro 2x at $59.98/mo doubles request limits; Ultra at $120/mo handles team and research workloads.
Stop reading model news.
Start using all of them.
Every frontier model in the matrix — twelve dated releases and counting — lives on one ZeroTwo account. Free to start. No card required.