How we test and score AI tools
The nine criteria behind every score, how long each tool is tested for, who pays for the subscriptions, and why no vendor can buy a ranking.
The five rules we do not break
We pay for everything
Every account is bought at full retail price on a normal plan. We do not accept free licences, extended trials, or press accounts, because a vendor-provisioned account is not the account you will get.
Minimum two weeks of real use
A tool has to survive at least two weeks of live client work or internal operations before it gets a score. Most get four to six. The testing period is stated at the top of every review.
Same tasks, same rubric
Every tool in a category is given the identical set of briefs, files, and edge cases, then scored against the nine criteria below. That is what makes scores comparable across reviews.
Nothing is for sale
No vendor has ever paid for a score, a badge, a position on a list, or a mention. We do not run sponsored posts, we do not send drafts to vendors for approval, and we do not remove critical findings on request.
Re-tested twice a year
Software moves. Every published review is re-tested at least every six months, and sooner if pricing changes or a major release lands. The last-updated date at the top of each article is the date of the most recent hands-on check.
The nine scoring criteria
Each criterion is scored from 1 to 5. The weighted average becomes the overall score you see on every review card, rounded to one decimal place.
| Criterion | Weight | What it measures |
|---|---|---|
| Output quality | 20% | How usable the result is without a rewrite, judged on the same set of real briefs given to every tool in the category. |
| Ease of setup | 10% | Time from signup to first genuinely useful output, measured with a stopwatch on a clean account. |
| Day-to-day speed | 10% | Latency on the tasks you run twenty times a day, not the ones you run once a quarter. |
| Reliability | 15% | Failed jobs, silent errors, and how the tool behaves when we deliberately feed it bad input. |
| Integrations | 10% | Whether it connects to the stack a US small business already runs, and whether those connections survive a token refresh. |
| Value for money | 15% | Cost against the alternative, including the cost of the plan you actually need rather than the headline tier. |
| Support | 8% | We open at least one real support ticket per tool from a paid account and record the response time and the answer. |
| Data handling | 7% | Training opt-outs, retention windows, export paths, and whether the terms match what the marketing page claims. |
| Fit for a small team | 5% | Seat minimums, admin overhead, and whether a one-person business can run it without an IT contractor. |
What the scores mean
- 4.5 to 5.0
- We would put our own money on it and recommend it without caveats.
- 4.0 to 4.4
- Strong. Buy it if the caveats in the review do not apply to you.
- 3.5 to 3.9
- Good at one thing, weak at others. Read the who-it-is-for section.
- 3.0 to 3.4
- Works, but something cheaper or simpler probably does the job.
- Below 3.0
- We would not spend our own money on it in its current state.
How this is funded
Tested AI makes money in two ways: display advertising, and affiliate commission on some outbound links. Both are disclosed. Neither influences a score.
Advertising is sold and filled by a third-party network. Advertisers cannot choose which articles they appear next to at the level of an individual review, and they have no editorial input whatsoever. Affiliate links are marked, and the reviewer writing a piece does not have access to affiliate revenue reporting for the tools they are scoring.
If a tool we rank first has no affiliate program at all, it still ranks first. That has happened, and it will happen again.