Every release.
Every claim, tested.
Providers publish the benchmark. We publish what happened when someone actually tried it.
AllReleasedAnnouncedRumoredAll providersAlibabaAnthropicByteDanceDeepSeekGoogle DeepMindMetaMiniMaxMoonshot AIOpenAISakana AIXiaomiZ.AIxAI
FEB 2026 — FEB 2027
FEBRUARY 2026
FEB 05
Anthropic · FEB 05, 2026
Reasoning-focused flagship aimed at long-running agent workflows.
$15.00/Mtok
No reviewed reports yet.
Read more →MARCH 2026
APRIL 2026
MAY 2026
MAY 28
Anthropic · MAY 28, 2026
Further reasoning gains over Opus 4.6 on long-horizon tasks.
No reviewed reports yet.
Read more →JUNE 2026
JUN 30
Anthropic · JUN 30, 2026
Mid-tier sibling to Opus, tuned for latency-sensitive production work.
$3.00/Mtok
No reviewed reports yet.
Read more →JULY 2026
JUL 24
Anthropic · JUL 24, 2026
Anthropic's flagship model, positioned for long-horizon agentic work and coding.
SWE-bench Verified 96.0%SWE-bench Pro 79.2%$15.00/Mtok
- CodingHandled a multi-file refactor across an unfamiliar Rust codebase without breaking the build, but needed a second pass to fix a borrow-checker edge case.
- AgenticRan a 40-step browser automation task end to end, but got stuck retrying a flaky selector instead of asking for help.
AUGUST 2026
TODAY
SEPTEMBER 2026
OCTOBER 2026
NOVEMBER 2026
DECEMBER 2026
JANUARY 2027
FEBRUARY 2027