Benchmarks published by labs measure what labs optimise for. We measure what buyers actually do, so our scores sometimes disagree with headline leaderboards.
The task suite
We maintain a rotating set of roughly two hundred tasks drawn from real professional workflows: refactoring a live codebase, reconciling two contradictory documents, drafting from a style guide, extracting structured data from messy PDFs, and multi-step research with source verification.
Scoring
- Reasoning — correctness on tasks with a verifiable answer.
- Coding — does the output run, and how much editing before it ships.
- Writing — edit distance from draft to publishable.
- Speed — median time to first useful token under normal load.
- Value — quality per dollar at realistic usage volumes.
Independence
We buy our own subscriptions at retail price. Scores are never adjusted for commercial relationships, and any page that carries an affiliate link says so at the top of the page.
We re-run the full suite quarterly and spot-check monthly. When a score changes materially, the provider page records the date of the change.
The aiz.to editorial team
We buy every subscription at retail price, run the same task suite across providers, and publish the scores unedited. Read the full methodology on our about page.
How we test