Claims are easy. Reality is the hard part.
Every AI lab publishes benchmark numbers when it ships a model. Most of those numbers are accurate. Few of them tell you what the model is actually like to use. trackai tracks both, side by side, and treats them differently — because they come from different places and deserve different levels of trust.
Two layers of data
CLAIM — WHAT THE PROVIDER SAYS
For each model we record what the lab itself announced: a short summary of what shipped, whichever benchmark figures that lab chose to publish, and a link straight to its announcement. Nothing here is independently verified — it is the provider’s own account of its own product, presented as exactly that.
Two things worth knowing when you read these numbers. Labs quote the benchmarks that flatter them, so the figures on one model page are rarely directly comparable to another’s. And where a summary was drafted automatically from the announcement, the page says so and links the source, so you can check it in one click.
REALITY — REVIEWED REPORTS
Reports start as posts on Hacker News, or submissions from readers, describing what happened when someone actually used a model on a real task. Each one is summarized into a short takeaway and a task tag — never a copy of the original text — with a link back to the source. Nothing here is confirmed information the way a release date is; it’s one person’s account, and it’s presented that way.
Review, before anything is public
Because reality-check reports are lower-confidence than confirmed releases, nothing from that layer goes live automatically. Every report — sourced from Hacker News or submitted directly — sits in a review queue until it’s approved. Rejected reports never appear on the site. There’s no algorithmic ranking or voting behind what gets published; it’s a single editorial pass.
Reading the status marker
The dot next to a model’s name on the timeline shows how confirmed it is:
- Rumored — talked about, nothing official yet.
- Announced — the provider has confirmed it’s coming.
- Released — shipped, with benchmark data attached.
Submit a report
If you’ve tried a model on a real task and it’s not reflected here yet, submit a report. One line on what happened, a link to back it up, and the task category it falls under.