trackai.

Every release.
Every claim, tested.

Providers publish the benchmark. We publish what happened when someone actually tried it.

FEB 2026 — FEB 2027
FEBRUARY 2026
FEB 05

Anthropic · FEB 05, 2026

Reasoning-focused flagship aimed at long-running agent workflows.

$15.00/Mtok

No reviewed reports yet.

Read more →
MARCH 2026
APRIL 2026
MAY 2026
MAY 28

Anthropic · MAY 28, 2026

Further reasoning gains over Opus 4.6 on long-horizon tasks.

No reviewed reports yet.

Read more →
JUNE 2026
JUN 30

Anthropic · JUN 30, 2026

Mid-tier sibling to Opus, tuned for latency-sensitive production work.

$3.00/Mtok

No reviewed reports yet.

Read more →
JULY 2026
JUL 24

Anthropic · JUL 24, 2026

Anthropic's flagship model, positioned for long-horizon agentic work and coding.

SWE-bench Verified 96.0%SWE-bench Pro 79.2%$15.00/Mtok

  • CodingHandled a multi-file refactor across an unfamiliar Rust codebase without breaking the build, but needed a second pass to fix a borrow-checker edge case.
  • AgenticRan a 40-step browser automation task end to end, but got stuck retrying a flaky selector instead of asking for help.
Read more →
AUGUST 2026
TODAY
SEPTEMBER 2026
OCTOBER 2026
NOVEMBER 2026
DECEMBER 2026
JANUARY 2027
FEBRUARY 2027