Methodology · v2 · 10 September 2026
How we research
We do not run controlled tests or generate our own clips. We read the testing that already exists — official documentation, independent reviews, technical evaluations, public demonstrations and published pricing — weigh each source, and publish the finding with a confidence level and the sources behind it.
What this site is not
Not a testing laboratory. We never claim or imply that we generated, ran or frame-by-frame inspected a tool's output. Where a finding rests on somebody else's test, we name them and say how much weight that test can carry.
Sources are weighted, not counted
Two reviews are not twice the evidence of one if both looked at the wrong model version. Each accepted source carries a composite weight built from three factors:
| Factor | Strongest | Weakest | Why |
|---|---|---|---|
| Version match | Names the exact model | Version unknown | A review of last year's release is not evidence about this year's |
| Directness | Tests the question directly | Indirect or contextual | An impression in passing is weaker than a measurement |
| Recency | Within six months | Over eighteen months | These products change monthly |
We prefer primary and authoritative sources, and we separate what a vendor claims from what an independent source observed. Official documentation is excellent evidence of a feature existing and poor evidence of it working well.
Five confidence levels
Findings carry a label rather than a number, because a number implies a precision the evidence does not have.
| Level | What it means |
|---|---|
| Strongly supported | Multiple independent, recent sources on the exact model version agree |
| Moderately supported | A useful conclusion with identified limitations |
| Limited evidence | Provisional. Often one source, or a close-but-not-exact version match |
| Conflicting evidence | Credible sources disagree, and we could not resolve it. We show both |
| Not established | We researched it and the public evidence does not support a conclusion |
A missing finding is never converted into a low score. "Not established" means we looked and could not tell — it is not a judgement that the tool performs badly.
Evidence coverage, and withheld verdicts
Alongside each tool we publish evidence coverage: how much of our criteria set the available sources actually answer. It is the honest answer to "how much of this do you really know?"
Where coverage falls below our threshold, we withhold the overall verdict rather than publishing a number the evidence cannot carry. At the time of writing that applies to two of the three tools we have researched in depth, and their pages say so.
Tool-level finding: 60% coverage · Category recommendation: 75% coverage and Moderate confidence or better
We also do not publish a single winner across product categories. An avatar platform and a cinematic generative model do different jobs; ranking them against each other on one scale would be a category error, so the site is structured to prevent it.
Conflicts are recorded, not smoothed over
When credible sources disagree we log it — what each said, and what it would take to settle. Those open questions are published rather than quietly resolved in favour of whichever reads better. The current research carries twenty-six such gaps.
What we do not accept
- Vendor-supplied accounts, embargoed previews or review copy
- Paid placements, sponsored positions or pay-for-inclusion
- Aggregated star ratings from sources that do not publish their method
- Marketing claims presented as performance evidence
Where a source we cite disclosed a complimentary account, we note it next to the finding.
Direct inspection
Where a public demonstration can be attributed to a named model version, watching it can strengthen a finding, and we do that when it is possible. It is not required for every article, and when a clip cannot be reliably attributed we say so instead of relying on it.
Updating and disputes
Every page carries the date it was last researched. We re-research a tool when it ships a new named model, when pricing changes, or when a reader shows us a source we missed.
If you think a finding is wrong, write to admin@testedclips.com and say why. We re-research it and publish the outcome either way, including when the re-check says the original was right. Corrections are noted on the page rather than quietly edited in.
Independence
Some links are affiliate links and we may earn a commission at no extra cost to you. Findings are not influenced by that. Tools with no affiliate programme are researched to the same standard and appear in the same tables, flagged so you know which is which. Full detail on the disclosure.