trackai.

Every release.
Every claim, tested.

Providers publish the benchmark. We publish what happened when someone actually tried it.

FEB 2026 — FEB 2027
FEBRUARY 2026
FEB 05

Anthropic · FEB 05, 2026

Reasoning-focused flagship aimed at long-running agent workflows.

$15.00/Mtok

No reviewed reports yet.

Read more →
FEB 152 releases

Z.AI · FEB 15, 2026

Open-weights flagship from Z.AI, strong on tool use for its size.

$0.60/Mtok

No reviewed reports yet.

Read more →

MiniMax · FEB 15, 2026

Long-context model with an emphasis on cost per token.

No reviewed reports yet.

Read more →
FEB 16

Alibaba · FEB 16, 2026

Sparse mixture-of-experts release in the Qwen3.5 line.

No reviewed reports yet.

Read more →
FEB 19

Google DeepMind · FEB 19, 2026

Native multimodal model with a very large context window.

$7.00/Mtok

No reviewed reports yet.

Read more →
FEB 24

OpenAI · FEB 24, 2026

Code-specialised variant tuned for repository-scale edits.

No reviewed reports yet.

Read more →
MARCH 2026
MAR 06

OpenAI · MAR 06, 2026

General-purpose update with improved tool-use reliability.

$10.00/Mtok

No reviewed reports yet.

Read more →
MAR 172 releases

OpenAI · MAR 17, 2026

Smallest GPT-5.4 tier, aimed at high-volume classification.

No reviewed reports yet.

Read more →

OpenAI · MAR 17, 2026

Smaller, cheaper sibling to GPT-5.4.

No reviewed reports yet.

Read more →
MAR 182 releases

Xiaomi · MAR 18, 2026

Xiaomi's multimodal model, positioned for on-device work.

No reviewed reports yet.

Read more →

MiniMax · MAR 18, 2026

Incremental update to the M2 line.

No reviewed reports yet.

Read more →
APRIL 2026
APR 07

Z.AI · APR 07, 2026

Point update to GLM-5 with a longer context window.

No reviewed reports yet.

Read more →
APR 14

ByteDance · APR 14, 2026

Video generation model from ByteDance's Seed group.

No reviewed reports yet.

Read more →
APR 20

Moonshot AI · APR 20, 2026

Agentic-focused release in the Kimi K2 line.

No reviewed reports yet.

Read more →
APR 232 releases

OpenAI · APR 23, 2026

Extended-reasoning tier of GPT-5.5.

No reviewed reports yet.

Read more →

OpenAI · APR 23, 2026

Flagship update with a substantially larger context window.

No reviewed reports yet.

Read more →
APR 24

DeepSeek · APR 24, 2026

Open-weights reasoning model at an aggressive price point.

$0.90/Mtok

  • CodingMatched a much pricier model on a LeetCode-style set, but degraded noticeably once the prompt exceeded ~60k tokens.
Read more →
MAY 2026
MAY 28

Anthropic · MAY 28, 2026

Further reasoning gains over Opus 4.6 on long-horizon tasks.

No reviewed reports yet.

Read more →
JUNE 2026
JUN 01

MiniMax · JUN 01, 2026

New generation of the MiniMax line.

No reviewed reports yet.

Read more →
JUN 13

Z.AI · JUN 13, 2026

Second point release in the GLM-5 series.

No reviewed reports yet.

Read more →
JUN 23

Sakana AI · JUN 23, 2026

Sakana AI's frontier model, developed in Tokyo.

No reviewed reports yet.

Read more →
JUN 30

Anthropic · JUN 30, 2026

Mid-tier sibling to Opus, tuned for latency-sensitive production work.

$3.00/Mtok

No reviewed reports yet.

Read more →
JULY 2026
JUL 08

xAI · JUL 08, 2026

Reasoning update to the Grok 4 line.

No reviewed reports yet.

Read more →
JUL 093 releases

OpenAI · JUL 09, 2026

Conversational tier of the GPT-5.6 family.

No reviewed reports yet.

Read more →

OpenAI · JUL 09, 2026

Reasoning tier of the GPT-5.6 family.

No reviewed reports yet.

Read more →

OpenAI · JUL 09, 2026

Agentic tier of the GPT-5.6 family.

No reviewed reports yet.

Read more →
JUL 16

Moonshot AI · JUL 16, 2026

New generation of Moonshot's open-weights line.

$0.50/Mtok

  • WritingHeld a consistent house style across a 12-section document without re-prompting, which the previous version could not do.
Read more →
JUL 24

Anthropic · JUL 24, 2026

Anthropic's flagship model, positioned for long-horizon agentic work and coding.

SWE-bench Verified 96.0%SWE-bench Pro 79.2%$15.00/Mtok

  • CodingHandled a multi-file refactor across an unfamiliar Rust codebase without breaking the build, but needed a second pass to fix a borrow-checker edge case.
  • AgenticRan a 40-step browser automation task end to end, but got stuck retrying a flaky selector instead of asking for help.
Read more →
AUGUST 2026
AUG 02

Alibaba · AUG 02, 2026

Largest tier of the Qwen3.8 series.

No reviewed reports yet.

Read more →
AUG 062 releases

Meta · AUG 06, 2026

Meta's creative generation model.

No reviewed reports yet.

Read more →

xAI · AUG 06, 2026

Latest Grok 4 point release.

No reviewed reports yet.

Read more →
AUG 13

Google DeepMind · AUG 13, 2026

Cost-optimised Gemini tier with the full context window.

$0.40/Mtok

  • VisionCorrectly read a handwritten whiteboard photo end to end, including a crossed-out line most tools misread as included text.
Read more →
AUG 14

Z.AI · AUG 14, 2026

Third point release in the GLM-5 series.

  • AgenticRan an overnight scraping agent on a single GPU without falling over, though it silently skipped two failed pages.
Read more →
AUG 17

Z.AI · AUG 17, 2026

Throughput-optimised variant of GLM-5.2.

No reviewed reports yet.

Read more →
TODAY
SEPTEMBER 2026
OCTOBER 2026
NOVEMBER 2026
DECEMBER 2026
JANUARY 2027
FEBRUARY 2027