Release archive

Every AI model we've tracked, newest first. 63 of 185 have a plain-language summary of what changed versus the previous version — more get added as models ship.

2026

Sep
4

No. 2 on the Arena text-to-image leaderboard, +79 Elo over MAI-Image-2.5 across all eight categories.

Sources: Microsoft AI
Sep
3

OpenAI's new flagship: far fewer made-up answers and a third of the tokens on coding-agent work, but the same overall score as GPT-5.6 Sol — at 2.5 times the price.

Compared with GPT-5.6 Sol · the model it replaces

Invents facts on half as many trick questions

AA hallucination rate — how often it makes up an answer instead of admitting it doesn't know

51%

down from 92% · 41 points fewer

Uses about a third of the tokens on coding-agent tasks

AA Coding Agent Index, max effort — text consumed per task

~⅓ tokens

vs GPT-5.6 Sol · 70% more token-efficient

Costs 2.5 times more to run than Sol

Price per 1M tokens — what developers pay, input / output

$10 / $50

up from $4 / $20 · 2.5× the price

Sep
2

Google's new Flash: a big jump on command-line tasks and a few points smarter, at the same price — though it burns about 30% more tokens per task, and the price doubles in 2027.

Compared with Gemini 3.7 Flash · the model it replaces

Finishes 9 in 10 multi-step tasks at the command line

Terminal-Bench 2.1 — running real commands to complete a task

90.8%

up from 81.6% · +9 points

A few points smarter on the all-round score

Artificial Analysis Intelligence Index — a broad average across many tests

59

up from 56 · +3 points

Same price as 3.7 Flash — until the end of 2026

Price per 1M tokens — what developers pay, input / output

$0.75 / $3.75

unchanged · doubles to $1.50 / $7.50 on Jan 1, 2027

Sources: Google · Vellum · 9to5Google
Sep
2

Meta's flagship edges closer to the leaders: four points up on the all-round score and clear gains on office and terminal tasks, at the same price and the same 1M-token context.

Compared with Muse Spark 1.2 · the model it replaces

Four points higher on the all-round intelligence score

Artificial Analysis Intelligence Index — a broad average across many tests

61

up from 57 · +4 points

Does better on realistic office work

GDPval-AA v2 — real knowledge-work tasks, scored as an Elo rating

1709

up from 1615 · +94 Elo

Same price as Muse Spark 1.2

Price per 1M tokens — what developers pay, input / output

$1.25 / $4.25

unchanged from Muse Spark 1.2

Sep
1

Anthropic's top model, refreshed: big jumps on long coding and scientific-research tasks, and cached input now costs 75% less — same headline price as Fable 5.

Compared with Claude Fable 5 · the model it replaces

Finishes over half of long multi-step coding tasks

Terminal-Bench 4.0 — long agentic coding sessions in a terminal

55.8%

up from 42.0% · +14 points

Doubles its score on agentic scientific research

Terminal-Bench-Science 0.1 — running real research workflows end to end

52.6%

up from 24.7% · more than doubled

Same price — and cached input now costs 75% less

Price per 1M tokens — what developers pay, input / output

$10 / $50

same as Fable 5 · cache reads now $0.25

Aug
26
Qw

Open-weight mixture-of-experts model with 125B total parameters but only 6B active per token, a 262K native context window and a built-in vision encoder, previewing the Qwen4 architecture.

Aug
14
Qw

Alibaba's open 27B model, now multimodal: big gains on computer use, coding and instruction following over Qwen3.6-27B — still free to download under Apache 2.0.

Compared with Qwen3.6-27B · the model it replaces

Operates a computer on its own far more reliably

OSWorld-Verified — completing real tasks inside a desktop OS

84.3%

up from 63.9% · +20 points

Solves 6 in 10 hard real-world coding tasks

SWE-bench Pro — fixing real bugs in real software projects

61.7%

up from 53.5% · +8 points

Free to download and run

License — open weights, commercial use allowed

Apache 2.0

same license as Qwen3.6-27B

Aug
13
DS

DeepSeek's V4-Pro leaves preview with huge agentic gains — terminal, long-running coding and security tasks all jump — and stays free under the MIT license.

Compared with DeepSeek-V4-Pro (preview) · the preview it replaces

Finishes almost 9 in 10 multi-step terminal tasks

Terminal Bench 2.1 — running real commands to complete a task

87.9%

up from 72.1% · +16 points

Five times better at long-running coding work

DeepSWE — multi-hour software-engineering tasks

62.7%

up from 12.8% · about 5×

Free to download and run

License — open weights, commercial use allowed

MIT

same license as the preview

Sources: Hugging Face
Aug
12
xAI

A post-training step on Grok 4.5 that adds an xhigh reasoning-effort level; cached input pricing rose from $0.30 to $0.50 per million tokens.

Sources: xAI · LLM Stats
Aug
12
Qw

Open-weights release of Qwen 3.8-Max, without the cloud model's image input and non-thinking mode; a commercial licence is required above $50M in annual revenue.

Aug
10

Meta's first open-weight model since Llama 4: 30B parameters under Apache 2.0, able to run offline on a single 24GB consumer GPU.

Aug
7
xAI

xAI's new image model, built for precise editing, sharp text rendering and consistent multi-image workflows.

Aug
5

Introduced a contributor tier at $0.10 per million input tokens and $0.20 per million output tokens in exchange for permission to train on the user's prompts and completions.

Sources: OpenRouter
Aug
3
Qw

2.4 trillion parameters with about 95 billion active per token and a 1M-token context window, served through Alibaba Cloud Model Studio.

Jul
31
DS

General-availability release of the 284B-parameter V4 sibling with 13B active per token, under the MIT licence with a 1M-token context window.

Sources: Hugging Face
Jul
24

Priced at $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8, and became the default model for Claude Max subscribers.

Sources: Anthropic · Axios
Jul
21

Released alongside 3.5 Flash-Lite; Google reported up to 17% fewer tokens used than the previous Flash model.

Sources: TechCrunch
Jul
21

The lowest-cost model of Google's Flash line, released the same day as 3.6 Flash.

Sources: TechCrunch
Jul
21
Qw

Released without benchmarks, a model card or open weights — a break from the open Qwen-Image line.

Sources: Unite.AI
Jul
9

The most capable of the three GPT-5.6 models (Sol, Terra and Luna), priced at $5 per million input tokens and $30 per million output tokens.

Sources: OpenAI · CNBC
Jul
9

The mid tier of the GPT-5.6 trio (Sol / Terra / Luna), positioned by OpenAI as the balanced everyday model: $2 per million input tokens and $12 per million output tokens after the 30 July price cut.

Sources: OpenAI · Vellum
Jul
9

The lightest tier of the GPT-5.6 trio, positioned by OpenAI as the fast, affordable model: $0.20 per million input tokens and $1.20 per million output tokens after the 30 July price cut.

Sources: OpenAI · Vellum
Jul
9

Opened the Meta Model API to outside developers in public preview, priced at $1.25 per million input tokens and $4.25 per million output tokens.

Sources: Meta AI · Axios
Jul
8
xAI

Added configurable reasoning effort (low, medium, high) with a 500K-token context window, priced at $2 per million input tokens and $6 per million output tokens.

Sources: xAI · TechCrunch
Jul
7

Meta's first in-house image model, free in the Meta AI app, Instagram Stories and WhatsApp.

Sources: Meta · TechCrunch
Jul
2

Apache 2.0 model for Lean 4 proof engineering with 119B parameters and 6B active, solving 587 of 672 PutnamBench problems.

Sources: Mistral · Hugging Face
Jun
30

Replaced Sonnet 4.6 as the default model on Anthropic's Free and Pro plans, priced at $2 per million input tokens and $10 per million output tokens.

Sources: Anthropic
Jun
10

Google's first open-weight text diffusion model: 26B parameters with 3.8B active, Apache 2.0, generating over 1,000 tokens per second on a single H100.

Jun
9

Anthropic's new top-end model: a big jump on hard coding tasks over Opus 4.8 — at twice the price.

Compared with Claude Opus 4.8 · the model it replaces

Solves 8 in 10 hard real-world coding tasks

SWE-bench Pro — fixing real bugs in real software projects

80.3%

up from 69.2% · +11 points

Solves about 3 in 10 research-grade coding problems

FrontierCode Diamond — the hardest tier of coding evals

29.3%

up from 13.4% · more than doubled

Costs developers twice as much to run

Price per 1M tokens — what developers pay, input / output

$10 / $50

price doubled — was $5 / $25

Jun
2

Microsoft's first in-house coding model — small, cheap, and ahead of Claude Haiku 4.5 on the coding tests Microsoft published.

Compared with Claude Haiku 4.5 · a competitor — comparison chosen by Microsoft

Solves about half of hard real-world coding tasks

SWE-bench Pro — fixing real bugs in real software projects

51.2%

vs 35.2% for Haiku 4.5 · +16 points

Follows written instructions more precisely than Haiku 4.5

IF Bench — gap only; Microsoft didn't publish raw scores

+28.9 points

ahead of Haiku 4.5

Costs less than a dollar per million input tokens

Price per 1M tokens — what developers pay, input / output

$0.75 / $4.50

launch price — nothing earlier to compare

Jun
2

Launched at No. 2 for image editing on Arena, unveiled at Build 2026.

Sources: Microsoft AI
Jun
1
Qw

Alibaba's proprietary multimodal agent model with a 1M-token context, generally available API-only on Alibaba Cloud Model Studio at $0.40/$1.60 per 1M tokens.

May
28

Anthropic's updated top-end model: steady coding gains over Opus 4.7 at the same price, plus a new Fast mode that runs the model at a third of the cost.

Compared with Claude Opus 4.7 · the model it replaces

Finishes more multi-step tasks at the command line

Terminal-Bench 2.1 — running real commands to complete a task

74.6%

up from 66.1% · +8 points

Solves 7 in 10 hard real-world coding tasks

SWE-bench Pro — fixing real bugs in real software projects

69.2%

up from 64.3% · +5 points

Costs the same to run as before

Price per 1M tokens — what developers pay, input / output

$5 / $25

same price as Opus 4.7

Sources: Vellum · VentureBeat
May
19

Google's lightweight Flash model: it edges out the bigger Gemini 3.1 Pro on a coding test, while running far more token-efficient than the Flash it replaces — at the same price.

Compared with Gemini 3 Flash · the model it replaces

Scores higher on a coding test than the bigger Gemini 3.1 Pro

Terminal-Bench 2.1 — solving coding tasks from the command line

76.2%

vs 70.3 for 3.1 Pro · +6 points

Uses far fewer tokens to do the same work

Token use — how much text the model burns per task

72% fewer

vs Gemini 3 Flash · same task

Costs $1.50 per million input tokens

Price per 1M tokens — what developers pay, input / output

$1.50 / $9.00

same price as Gemini 3 Flash

May
19

Google's new any-to-video model: text, images, audio or video in, physics-aware video out, editable in plain English — 10-second clips, free in YouTube Shorts and the Gemini app.

May
19
Qw

Alibaba's new top Qwen model: clear gains on terminal coding work and hard reasoning tests over 3.6-Plus, plus a much larger memory — at the same price.

Compared with Qwen 3.6-Plus · the model it replaces

Solves 7 in 10 hands-on coding tasks in a terminal

Terminal-Bench 2.0 — doing real work in a command-line terminal

69.7%

up from 61.6% · +8 points

Answers about 4 in 10 of the hardest expert exam questions

Humanity's Last Exam — graduate-level questions across many fields

38.1%

up from 28.9% · +9 points

Costs the same to run as before

Price per 1M tokens — what developers pay, input / output

$2.50 / $7.50

same price as Qwen 3.6-Plus

Sources: TechNode · DataCamp
May
5

OpenAI's new lightweight model: far fewer made-up answers and much stronger math than GPT-5.3 Instant Mini — and it's now the default free model in ChatGPT.

Compared with GPT-5.3 Instant Mini · the model it replaces

Makes up about half as many wrong answers on tricky questions

High-stakes prompts — questions where a wrong answer really matters

52% fewer

vs GPT-5.3 Instant Mini · same tests

Solves 8 in 10 competition-level math problems

AIME 2025 — a hard high-school math competition

81.2

up from 65.4 · +16 points

Now the model free ChatGPT users get by default

Availability — which users get this model without paying

default free ChatGPT

new in GPT-5.5 Instant

Apr
30
xAI

xAI's update to Grok 4.20: it can now hold a whole book of text at once, watch video, and costs 40% less for the text you send in — with a few points higher on a broad skills test.

Compared with Grok 4.20 · the model it replaces

Can hold a whole book in mind at once

Context window — how much text it can read in one go

1M tokens

new in Grok 4.3

Can now watch video, not just read text and see images

Video input — understanding moving footage directly

native

new in Grok 4.3

Costs 40% less for the text you send it

Input price — what developers pay for the text they send in

40% cheaper

vs Grok 4.20 · input only

Apr
24
DS

DeepSeek's new open flagship: about 8 times more text in memory than V3.2, while using a fraction of the computing power — and it solves about 8 in 10 real-world coding tasks.

Compared with DeepSeek-V3.2 · the version it replaces

Holds about 8 times more text in memory at once

Context window — how much text the model can read at one time

1M tokens

up from 128K · about 8× more

Uses far less computing power to run at full memory

Compute at 1M context — the processing work needed to run it

27% of V3.2

vs V3.2 · same 1M context

Costs developers about 14 cents per million input tokens

Price per 1M tokens — what developers pay for the Flash version's input

$0.14 / M

launch price — nothing earlier to compare

Sources: Hugging Face · NxCode
Apr
23

OpenAI's GPT-5.5 roughly doubles its reasoning over very long documents and writes shorter answers than GPT-5.4 — but it costs twice as much to run.

Compared with GPT-5.4 · the model it replaces

Roughly doubles its reasoning over book-length documents

1M-token input — handling very long documents in one go

74.0%

up from 36.6% · roughly doubled

Writes shorter answers to long prompts

Output length — how much text the model produces

19-34% shorter

vs GPT-5.4 · same prompts

Costs developers twice as much to run

Price per 1M tokens — what developers pay, input / output

$5 / $30

price doubled — was $2.50 / $15

Sources: OpenAI · llm-stats
Apr
22
Qw

The first open-weight variant of the Qwen3.6 family — a coding-focused 27B dense model under Apache 2.0 with a 262K native context, extensible to 1M.

Apr
21

Reasons before it draws: up to 2K output, up to eight coherent images per prompt, and reliable text rendering including non-Latin scripts.

Sources: Wikipedia · MindStudio
Apr
16

Anthropic's updated flagship: a clear jump on real-world coding and much sharper at reading charts than 4.6 — at the same price.

Compared with Claude Opus 4.6 · the model it replaces

Solves nearly 9 in 10 real-world coding tasks

SWE-bench Verified — fixing real bugs in real software projects

87.6%

up from 80.8% · +7 points

Makes the right code edits 7 times in 10 inside an editor

CursorBench — editing code the way developers do in an IDE

70%

up from 58% · +12 points

Costs developers the same as before

Price per 1M tokens — what developers pay, input / output

$5 / $25

same price as Opus 4.6

Sources: anthropic.com · Vellum
Apr
2
Apr
1
Qw

Unified architecture for image generation and editing; the video models followed the week after, with select models back under Apache 2.0.

Mar
16

An open-weight (Apache 2.0) code agent for the Lean 4 proof assistant that generates code and formally proves implementations against specifications.

Feb
19

Google's updated flagship: its score on brand-new reasoning puzzles more than doubled over Gemini 3 Pro — at the same price.

Compared with Gemini 3 Pro · the model it replaces

Solves about 8 in 10 brand-new reasoning puzzles

ARC-AGI-2 — novel puzzles a model hasn't seen before

77.1%

up from 31.1% · more than doubled

Handles 7 in 10 tasks that use outside tools

MCP Atlas — calling tools and services to get work done

69.2%

up from 59.5% · +10 points

Costs the same to run as before

Price per 1M tokens — what developers pay, input / output

$2 / $12

same price as Gemini 3 Pro

Sources: blog.google · DataCamp
Feb
17
xAI

xAI's new flagship Grok: the first version that retrains itself every week from real-world feedback, and it makes up fewer answers than any model tested so far.

Compared with Grok 4.1 · the version it replaces

Retrains itself every week from real-world feedback

Update cadence — how often the model is refreshed

weekly

new in Grok 4.20 · first Grok to do this

Makes up fewer answers than any model tested so far

AA Omniscience — how often the model avoids inventing facts

78%

best of any model tested · no earlier score to compare

Lands among the top-scoring models in public head-to-head votes

Arena Elo — early ranking from people comparing answers blind

~1505-1535

early measurement · nothing earlier to compare

Feb
17
Qw
Feb
10
Qw

Unifies generation and editing in a single 7B model — nearly 3× smaller than Qwen-Image, yet ahead on every major benchmark reported.

Feb
5

Anthropic's Opus 4.6: a big jump on brand-new reasoning puzzles and the first Opus that can read 1M tokens of text at once — but everyday coding is unchanged.

Compared with Claude Opus 4.5 · the model it replaces

Solves far more puzzle-style problems it has never seen before

ARC-AGI-2 — reasoning through brand-new problems, not memorized ones

68.8%

up from 37.6% · +31 points

First Opus that can read 1M tokens of text at once

Context window — how much text it can hold in one go

1M tokens

new in Opus 4.6 — earlier Opus models held less

Everyday coding is unchanged from the last version

SWE-bench — fixing real bugs in real software projects

80.8%

no change from Opus 4.5 — was 80.9%

Sources: anthropic.com · Vellum
Feb
4
Qw

An open-weight (Apache 2.0) coding-agent model with 80B total and 3B active parameters and a 256K context, built for coding agents and local development.

Feb
3
xAI

xAI's image-and-video generator: 10-second 720p clips with native audio, including lip-synced dialogue.

Sources: Basenor · xAI docs

2025

Dec
16

Image inputs and outputs 20% cheaper in the API than GPT Image 1.

Sources: Wikipedia
Dec
16
Qw

Shipped as a closed API only — no weights, a break from the open 2.1 and 2.2.

Nov
20

Built on Gemini 3 Pro: far better text rendering inside images, consistent identity across up to five subjects, and 2K/4K output.

Sources: Google · TechRadar
Oct
15

Adds 4K and native portrait output on top of Veo 3's synchronized audio.

Oct
13

Microsoft's first in-house image model, debuting in the top 10 on LMArena and inside Bing Image Creator.

Sources: Learn AI
Oct
6

A smaller GPT Image that costs 80% less in the API.

Sources: Wikipedia
Sep
30

Adds synchronized sound and 'cameos' (insert yourself into clips), launched as a social app. OpenAI shut the Sora app down on April 26, 2026 and ends the API on September 24, 2026.

Sources: OpenAI · Wikipedia
Aug
26

Google's image generation moves into Gemini itself: consistent characters across edits and multi-image blending, at Flash speed.

Aug
4
Qw

Alibaba's open image model (Apache 2.0), strong at rendering text inside images.

Sources: GitHub · Hugging Face
Jul
28
Qw

Open weights again under Apache 2.0 — the last fully open Wan video release.

May
22
May
20

Google's video model gains synchronized audio — dialogue, ambient sound and effects generated with the footage.

Apr
28
Qw
Apr
5
Llama 4 Scout/MaverickMeta · LlamaOpen weightsDiscontinued
Mar
12
Feb
25
Qw

Alibaba's open-weight video model, released under Apache 2.0.

Sources: HowAIWorks
Feb
17
xAI

2024

Dec
12
Dec
9
SoraOpenAI · SoraDiscontinued
Dec
6
Llama 3.3Meta · LlamaOpen weightsDiscontinued
Sep
25
Llama 3.2Meta · LlamaOpen weightsDiscontinued
Sep
19
Qw
Aug
20
Aug
13
xAI
Jul
23
Llama 3.1Meta · LlamaOpen weightsDiscontinued
Jul
1
DS
Jun
27
Jun
6
Qw
Apr
23
Apr
18
Llama 3Meta · LlamaOpen weightsDiscontinued
Feb
21
Jan
15
Qw

2023

2022

2020