Where the line stands
How long it has been since the last release against the line's usual gap and its recent ones — the rhythm, not a date.
Every Qwen Omni release
Newest first — what each version changed versus the one before, in Alibaba's own numbers.
Qwen's first omni-modal model built for agents: it takes text, images, audio and video and returns text, with a 1M-token context window (up from 256K) and output tokens priced about 90% below Qwen3.5-Omni-Plus.
Compared with Qwen3.5-Omni-Plus · the model it replaces
Completes 7 in 10 multimodal agent tasks
WildClawBench-MM — agent tasks that combine audio, video and tool use
71.0
▲ up from 34.5 for Qwen3.5-Omni-Plus · +36.5 points
Costs about 90% less per output token than the model it replaces
Price per 1M tokens — what developers pay, input / output
$0.15 / $0.47
▲ down from $0.40 / $4.80 for Qwen3.5-Omni-Plus · output 90% cheaper
Alibaba's first closed-weight Qwen model: a native omni-modal API model that takes text, images, audio and video, with a 256K-token context window.
Get the weekly drop
One email a week: what AI models shipped, the actual differences vs the previous version, and what's coming next. No filler.