How We Test AI Tools at aiz.to

Our evaluation methodology: the task suite, the scoring rubric, how often we re-test, and how we handle vendor relationships.

The aiz.to editorial team5 min read

Benchmarks published by labs measure what labs optimise for. We measure what buyers actually do, so our scores sometimes disagree with headline leaderboards.

The task suite

We maintain a rotating set of roughly two hundred tasks drawn from real professional workflows: refactoring a live codebase, reconciling two contradictory documents, drafting from a style guide, extracting structured data from messy PDFs, and multi-step research with source verification.

Scoring

  • Reasoning — correctness on tasks with a verifiable answer.
  • Coding — does the output run, and how much editing before it ships.
  • Writing — edit distance from draft to publishable.
  • Speed — median time to first useful token under normal load.
  • Value — quality per dollar at realistic usage volumes.

Independence

We buy our own subscriptions at retail price. Scores are never adjusted for commercial relationships, and any page that carries an affiliate link says so at the top of the page.

We re-run the full suite quarterly and spot-check monthly. When a score changes materially, the provider page records the date of the change.

MethodologyTestingIndependence
Found this useful? Pass it on.
About the author

The aiz.to editorial team

We buy every subscription at retail price, run the same task suite across providers, and publish the scores unedited. Read the full methodology on our about page.

How we test