Release archive
Every AI model we've tracked, newest first. 63 of 185 have a plain-language summary of what changed versus the previous version — more get added as models ship.
2026
No. 2 on the Arena text-to-image leaderboard, +79 Elo over MAI-Image-2.5 across all eight categories.
OpenAI's new flagship: far fewer made-up answers and a third of the tokens on coding-agent work, but the same overall score as GPT-5.6 Sol — at 2.5 times the price.
Compared with GPT-5.6 Sol · the model it replaces
Invents facts on half as many trick questions
AA hallucination rate — how often it makes up an answer instead of admitting it doesn't know
51%
▲ down from 92% · 41 points fewer
Uses about a third of the tokens on coding-agent tasks
AA Coding Agent Index, max effort — text consumed per task
~⅓ tokens
▲ vs GPT-5.6 Sol · 70% more token-efficient
Costs 2.5 times more to run than Sol
Price per 1M tokens — what developers pay, input / output
$10 / $50
up from $4 / $20 · 2.5× the price
Google's new Flash: a big jump on command-line tasks and a few points smarter, at the same price — though it burns about 30% more tokens per task, and the price doubles in 2027.
Compared with Gemini 3.7 Flash · the model it replaces
Finishes 9 in 10 multi-step tasks at the command line
Terminal-Bench 2.1 — running real commands to complete a task
90.8%
▲ up from 81.6% · +9 points
A few points smarter on the all-round score
Artificial Analysis Intelligence Index — a broad average across many tests
59
▲ up from 56 · +3 points
Same price as 3.7 Flash — until the end of 2026
Price per 1M tokens — what developers pay, input / output
$0.75 / $3.75
unchanged · doubles to $1.50 / $7.50 on Jan 1, 2027
Meta's flagship edges closer to the leaders: four points up on the all-round score and clear gains on office and terminal tasks, at the same price and the same 1M-token context.
Compared with Muse Spark 1.2 · the model it replaces
Four points higher on the all-round intelligence score
Artificial Analysis Intelligence Index — a broad average across many tests
61
▲ up from 57 · +4 points
Does better on realistic office work
GDPval-AA v2 — real knowledge-work tasks, scored as an Elo rating
1709
▲ up from 1615 · +94 Elo
Same price as Muse Spark 1.2
Price per 1M tokens — what developers pay, input / output
$1.25 / $4.25
unchanged from Muse Spark 1.2
Anthropic's top model, refreshed: big jumps on long coding and scientific-research tasks, and cached input now costs 75% less — same headline price as Fable 5.
Compared with Claude Fable 5 · the model it replaces
Finishes over half of long multi-step coding tasks
Terminal-Bench 4.0 — long agentic coding sessions in a terminal
55.8%
▲ up from 42.0% · +14 points
Doubles its score on agentic scientific research
Terminal-Bench-Science 0.1 — running real research workflows end to end
52.6%
▲ up from 24.7% · more than doubled
Same price — and cached input now costs 75% less
Price per 1M tokens — what developers pay, input / output
$10 / $50
same as Fable 5 · cache reads now $0.25
Open-weight mixture-of-experts model with 125B total parameters but only 6B active per token, a 262K native context window and a built-in vision encoder, previewing the Qwen4 architecture.
Alibaba's open 27B model, now multimodal: big gains on computer use, coding and instruction following over Qwen3.6-27B — still free to download under Apache 2.0.
Compared with Qwen3.6-27B · the model it replaces
Operates a computer on its own far more reliably
OSWorld-Verified — completing real tasks inside a desktop OS
84.3%
▲ up from 63.9% · +20 points
Solves 6 in 10 hard real-world coding tasks
SWE-bench Pro — fixing real bugs in real software projects
61.7%
▲ up from 53.5% · +8 points
Free to download and run
License — open weights, commercial use allowed
Apache 2.0
same license as Qwen3.6-27B
DeepSeek's V4-Pro leaves preview with huge agentic gains — terminal, long-running coding and security tasks all jump — and stays free under the MIT license.
Compared with DeepSeek-V4-Pro (preview) · the preview it replaces
Finishes almost 9 in 10 multi-step terminal tasks
Terminal Bench 2.1 — running real commands to complete a task
87.9%
▲ up from 72.1% · +16 points
Five times better at long-running coding work
DeepSWE — multi-hour software-engineering tasks
62.7%
▲ up from 12.8% · about 5×
Free to download and run
License — open weights, commercial use allowed
MIT
same license as the preview
Open-weights release of Qwen 3.8-Max, without the cloud model's image input and non-thinking mode; a commercial licence is required above $50M in annual revenue.
Meta's first open-weight model since Llama 4: 30B parameters under Apache 2.0, able to run offline on a single 24GB consumer GPU.
xAI's new image model, built for precise editing, sharp text rendering and consistent multi-image workflows.
Introduced a contributor tier at $0.10 per million input tokens and $0.20 per million output tokens in exchange for permission to train on the user's prompts and completions.
2.4 trillion parameters with about 95 billion active per token and a 1M-token context window, served through Alibaba Cloud Model Studio.
General-availability release of the 284B-parameter V4 sibling with 13B active per token, under the MIT licence with a 1M-token context window.
Released alongside 3.5 Flash-Lite; Google reported up to 17% fewer tokens used than the previous Flash model.
The lowest-cost model of Google's Flash line, released the same day as 3.6 Flash.
Released without benchmarks, a model card or open weights — a break from the open Qwen-Image line.
Added configurable reasoning effort (low, medium, high) with a 500K-token context window, priced at $2 per million input tokens and $6 per million output tokens.
Meta's first in-house image model, free in the Meta AI app, Instagram Stories and WhatsApp.
Apache 2.0 model for Lean 4 proof engineering with 119B parameters and 6B active, solving 587 of 672 PutnamBench problems.
Replaced Sonnet 4.6 as the default model on Anthropic's Free and Pro plans, priced at $2 per million input tokens and $10 per million output tokens.
Google's first open-weight text diffusion model: 26B parameters with 3.8B active, Apache 2.0, generating over 1,000 tokens per second on a single H100.
Anthropic's new top-end model: a big jump on hard coding tasks over Opus 4.8 — at twice the price.
Compared with Claude Opus 4.8 · the model it replaces
Solves 8 in 10 hard real-world coding tasks
SWE-bench Pro — fixing real bugs in real software projects
80.3%
▲ up from 69.2% · +11 points
Solves about 3 in 10 research-grade coding problems
FrontierCode Diamond — the hardest tier of coding evals
29.3%
▲ up from 13.4% · more than doubled
Costs developers twice as much to run
Price per 1M tokens — what developers pay, input / output
$10 / $50
price doubled — was $5 / $25
Microsoft's first in-house coding model — small, cheap, and ahead of Claude Haiku 4.5 on the coding tests Microsoft published.
Compared with Claude Haiku 4.5 · a competitor — comparison chosen by Microsoft
Solves about half of hard real-world coding tasks
SWE-bench Pro — fixing real bugs in real software projects
51.2%
▲ vs 35.2% for Haiku 4.5 · +16 points
Follows written instructions more precisely than Haiku 4.5
IF Bench — gap only; Microsoft didn't publish raw scores
+28.9 points
▲ ahead of Haiku 4.5
Costs less than a dollar per million input tokens
Price per 1M tokens — what developers pay, input / output
$0.75 / $4.50
✦ launch price — nothing earlier to compare
Alibaba's proprietary multimodal agent model with a 1M-token context, generally available API-only on Alibaba Cloud Model Studio at $0.40/$1.60 per 1M tokens.
Anthropic's updated top-end model: steady coding gains over Opus 4.7 at the same price, plus a new Fast mode that runs the model at a third of the cost.
Compared with Claude Opus 4.7 · the model it replaces
Finishes more multi-step tasks at the command line
Terminal-Bench 2.1 — running real commands to complete a task
74.6%
▲ up from 66.1% · +8 points
Solves 7 in 10 hard real-world coding tasks
SWE-bench Pro — fixing real bugs in real software projects
69.2%
▲ up from 64.3% · +5 points
Costs the same to run as before
Price per 1M tokens — what developers pay, input / output
$5 / $25
same price as Opus 4.7
Google's lightweight Flash model: it edges out the bigger Gemini 3.1 Pro on a coding test, while running far more token-efficient than the Flash it replaces — at the same price.
Compared with Gemini 3 Flash · the model it replaces
Scores higher on a coding test than the bigger Gemini 3.1 Pro
Terminal-Bench 2.1 — solving coding tasks from the command line
76.2%
▲ vs 70.3 for 3.1 Pro · +6 points
Uses far fewer tokens to do the same work
Token use — how much text the model burns per task
72% fewer
▲ vs Gemini 3 Flash · same task
Costs $1.50 per million input tokens
Price per 1M tokens — what developers pay, input / output
$1.50 / $9.00
same price as Gemini 3 Flash
Google's new any-to-video model: text, images, audio or video in, physics-aware video out, editable in plain English — 10-second clips, free in YouTube Shorts and the Gemini app.
Alibaba's new top Qwen model: clear gains on terminal coding work and hard reasoning tests over 3.6-Plus, plus a much larger memory — at the same price.
Compared with Qwen 3.6-Plus · the model it replaces
Solves 7 in 10 hands-on coding tasks in a terminal
Terminal-Bench 2.0 — doing real work in a command-line terminal
69.7%
▲ up from 61.6% · +8 points
Answers about 4 in 10 of the hardest expert exam questions
Humanity's Last Exam — graduate-level questions across many fields
38.1%
▲ up from 28.9% · +9 points
Costs the same to run as before
Price per 1M tokens — what developers pay, input / output
$2.50 / $7.50
same price as Qwen 3.6-Plus
OpenAI's new lightweight model: far fewer made-up answers and much stronger math than GPT-5.3 Instant Mini — and it's now the default free model in ChatGPT.
Compared with GPT-5.3 Instant Mini · the model it replaces
Makes up about half as many wrong answers on tricky questions
High-stakes prompts — questions where a wrong answer really matters
52% fewer
▲ vs GPT-5.3 Instant Mini · same tests
Solves 8 in 10 competition-level math problems
AIME 2025 — a hard high-school math competition
81.2
▲ up from 65.4 · +16 points
Now the model free ChatGPT users get by default
Availability — which users get this model without paying
default free ChatGPT
✦ new in GPT-5.5 Instant
xAI's update to Grok 4.20: it can now hold a whole book of text at once, watch video, and costs 40% less for the text you send in — with a few points higher on a broad skills test.
Compared with Grok 4.20 · the model it replaces
Can hold a whole book in mind at once
Context window — how much text it can read in one go
1M tokens
✦ new in Grok 4.3
Can now watch video, not just read text and see images
Video input — understanding moving footage directly
native
✦ new in Grok 4.3
Costs 40% less for the text you send it
Input price — what developers pay for the text they send in
40% cheaper
▲ vs Grok 4.20 · input only
DeepSeek's new open flagship: about 8 times more text in memory than V3.2, while using a fraction of the computing power — and it solves about 8 in 10 real-world coding tasks.
Compared with DeepSeek-V3.2 · the version it replaces
Holds about 8 times more text in memory at once
Context window — how much text the model can read at one time
1M tokens
▲ up from 128K · about 8× more
Uses far less computing power to run at full memory
Compute at 1M context — the processing work needed to run it
27% of V3.2
▲ vs V3.2 · same 1M context
Costs developers about 14 cents per million input tokens
Price per 1M tokens — what developers pay for the Flash version's input
$0.14 / M
✦ launch price — nothing earlier to compare
OpenAI's GPT-5.5 roughly doubles its reasoning over very long documents and writes shorter answers than GPT-5.4 — but it costs twice as much to run.
Compared with GPT-5.4 · the model it replaces
Roughly doubles its reasoning over book-length documents
1M-token input — handling very long documents in one go
74.0%
▲ up from 36.6% · roughly doubled
Writes shorter answers to long prompts
Output length — how much text the model produces
19-34% shorter
▲ vs GPT-5.4 · same prompts
Costs developers twice as much to run
Price per 1M tokens — what developers pay, input / output
$5 / $30
price doubled — was $2.50 / $15
The first open-weight variant of the Qwen3.6 family — a coding-focused 27B dense model under Apache 2.0 with a 262K native context, extensible to 1M.
Reasons before it draws: up to 2K output, up to eight coherent images per prompt, and reliable text rendering including non-Latin scripts.
Anthropic's updated flagship: a clear jump on real-world coding and much sharper at reading charts than 4.6 — at the same price.
Compared with Claude Opus 4.6 · the model it replaces
Solves nearly 9 in 10 real-world coding tasks
SWE-bench Verified — fixing real bugs in real software projects
87.6%
▲ up from 80.8% · +7 points
Makes the right code edits 7 times in 10 inside an editor
CursorBench — editing code the way developers do in an IDE
70%
▲ up from 58% · +12 points
Costs developers the same as before
Price per 1M tokens — what developers pay, input / output
$5 / $25
same price as Opus 4.6
Unified architecture for image generation and editing; the video models followed the week after, with select models back under Apache 2.0.
An open-weight (Apache 2.0) code agent for the Lean 4 proof assistant that generates code and formally proves implementations against specifications.
Google's updated flagship: its score on brand-new reasoning puzzles more than doubled over Gemini 3 Pro — at the same price.
Compared with Gemini 3 Pro · the model it replaces
Solves about 8 in 10 brand-new reasoning puzzles
ARC-AGI-2 — novel puzzles a model hasn't seen before
77.1%
▲ up from 31.1% · more than doubled
Handles 7 in 10 tasks that use outside tools
MCP Atlas — calling tools and services to get work done
69.2%
▲ up from 59.5% · +10 points
Costs the same to run as before
Price per 1M tokens — what developers pay, input / output
$2 / $12
same price as Gemini 3 Pro
xAI's new flagship Grok: the first version that retrains itself every week from real-world feedback, and it makes up fewer answers than any model tested so far.
Compared with Grok 4.1 · the version it replaces
Retrains itself every week from real-world feedback
Update cadence — how often the model is refreshed
weekly
✦ new in Grok 4.20 · first Grok to do this
Makes up fewer answers than any model tested so far
AA Omniscience — how often the model avoids inventing facts
78%
▲ best of any model tested · no earlier score to compare
Lands among the top-scoring models in public head-to-head votes
Arena Elo — early ranking from people comparing answers blind
~1505-1535
▲ early measurement · nothing earlier to compare
Unifies generation and editing in a single 7B model — nearly 3× smaller than Qwen-Image, yet ahead on every major benchmark reported.
Anthropic's Opus 4.6: a big jump on brand-new reasoning puzzles and the first Opus that can read 1M tokens of text at once — but everyday coding is unchanged.
Compared with Claude Opus 4.5 · the model it replaces
Solves far more puzzle-style problems it has never seen before
ARC-AGI-2 — reasoning through brand-new problems, not memorized ones
68.8%
▲ up from 37.6% · +31 points
First Opus that can read 1M tokens of text at once
Context window — how much text it can hold in one go
1M tokens
✦ new in Opus 4.6 — earlier Opus models held less
Everyday coding is unchanged from the last version
SWE-bench — fixing real bugs in real software projects
80.8%
no change from Opus 4.5 — was 80.9%
An open-weight (Apache 2.0) coding-agent model with 80B total and 3B active parameters and a 256K context, built for coding agents and local development.
2025
Shipped as a closed API only — no weights, a break from the open 2.1 and 2.2.
Adds 4K and native portrait output on top of Veo 3's synchronized audio.
Microsoft's first in-house image model, debuting in the top 10 on LMArena and inside Bing Image Creator.
Google's image generation moves into Gemini itself: consistent characters across edits and multi-image blending, at Flash speed.
Alibaba's open image model (Apache 2.0), strong at rendering text inside images.
Open weights again under Apache 2.0 — the last fully open Wan video release.
Google's video model gains synchronized audio — dialogue, ambient sound and effects generated with the footage.