Methodology
How we test
Same prompt, same seed, every model. Then three independent verdicts.
1. Fixed prompts per use case
Each category (for example product photos) has a set of prompts written to test specific, checkable things — listed under every prompt. Every model gets the identical prompt text, aspect ratio and seed, with default settings, through the fal.ai API.
2. Four criteria, scored 1–5
- Prompt adherence
Is every requested subject, attribute, composition and lighting detail there?
- Aesthetics
Would a professional publish this image for its purpose?
- Text accuracy
Only when the prompt asks for text: exact spelling and legibility.
- Artifacts
Absence of warped objects, melted details and broken geometry. Higher is cleaner.
3. Three verdicts
Editor — a human reviewer scores every image against the rubric and writes a short comment.
AI judge — a Claude vision model scores the same image with the same rubric and explains why. We publish which judge model was used.
Community — coming soon: head-to-head votes from visitors.
Rankings use the editor score and fall back to the AI judge where the editor has not scored yet. Where the verdicts disagree, we show both.
4. Cost and time to image
Cost per image is estimated from the provider’s published price and the output size at test time.
Time matters as much as quality when you generate at volume, so we measure two numbers for every image: time to image — the full wait from sending the request to receiving the finished image, including queueing — and generation time, the model’s own compute time as reported by the provider. Rankings show the median of each.
5. Freshness
Models change. Every page shows when it was last tested, and we re-test when a model gets a new version.