Every release.
Every claim, tested.
Providers publish the benchmark. We publish what happened when someone actually tried it.
Anthropic · FEB 05, 2026
Reasoning-focused flagship aimed at long-running agent workflows.
$15.00/Mtok
No reviewed reports yet.
Read more →Z.AI · FEB 15, 2026
Open-weights flagship from Z.AI, strong on tool use for its size.
$0.60/Mtok
No reviewed reports yet.
Read more →MiniMax · FEB 15, 2026
Long-context model with an emphasis on cost per token.
No reviewed reports yet.
Read more →Alibaba · FEB 16, 2026
Sparse mixture-of-experts release in the Qwen3.5 line.
No reviewed reports yet.
Read more →Google DeepMind · FEB 19, 2026
Native multimodal model with a very large context window.
$7.00/Mtok
No reviewed reports yet.
Read more →OpenAI · FEB 24, 2026
Code-specialised variant tuned for repository-scale edits.
No reviewed reports yet.
Read more →OpenAI · MAR 06, 2026
General-purpose update with improved tool-use reliability.
$10.00/Mtok
No reviewed reports yet.
Read more →OpenAI · MAR 17, 2026
Smallest GPT-5.4 tier, aimed at high-volume classification.
No reviewed reports yet.
Read more →Xiaomi · MAR 18, 2026
Xiaomi's multimodal model, positioned for on-device work.
No reviewed reports yet.
Read more →Z.AI · APR 07, 2026
Point update to GLM-5 with a longer context window.
No reviewed reports yet.
Read more →ByteDance · APR 14, 2026
Video generation model from ByteDance's Seed group.
No reviewed reports yet.
Read more →Moonshot AI · APR 20, 2026
Agentic-focused release in the Kimi K2 line.
No reviewed reports yet.
Read more →OpenAI · APR 23, 2026
Flagship update with a substantially larger context window.
No reviewed reports yet.
Read more →DeepSeek · APR 24, 2026
Open-weights reasoning model at an aggressive price point.
$0.90/Mtok
- CodingMatched a much pricier model on a LeetCode-style set, but degraded noticeably once the prompt exceeded ~60k tokens.
Anthropic · MAY 28, 2026
Further reasoning gains over Opus 4.6 on long-horizon tasks.
No reviewed reports yet.
Read more →Sakana AI · JUN 23, 2026
Sakana AI's frontier model, developed in Tokyo.
No reviewed reports yet.
Read more →Anthropic · JUN 30, 2026
Mid-tier sibling to Opus, tuned for latency-sensitive production work.
$3.00/Mtok
No reviewed reports yet.
Read more →Moonshot AI · JUL 16, 2026
New generation of Moonshot's open-weights line.
$0.50/Mtok
- WritingHeld a consistent house style across a 12-section document without re-prompting, which the previous version could not do.
Anthropic · JUL 24, 2026
Anthropic's flagship model, positioned for long-horizon agentic work and coding.
SWE-bench Verified 96.0%SWE-bench Pro 79.2%$15.00/Mtok
- CodingHandled a multi-file refactor across an unfamiliar Rust codebase without breaking the build, but needed a second pass to fix a borrow-checker edge case.
- AgenticRan a 40-step browser automation task end to end, but got stuck retrying a flaky selector instead of asking for help.
Google DeepMind · AUG 13, 2026
Cost-optimised Gemini tier with the full context window.
$0.40/Mtok
- VisionCorrectly read a handwritten whiteboard photo end to end, including a crossed-out line most tools misread as included text.
Z.AI · AUG 14, 2026
Third point release in the GLM-5 series.
- AgenticRan an overnight scraping agent on a single GPU without falling over, though it silently skipped two failed pages.