Fixed standards
The test rows, prompts, metrics and decision rules were fixed before any competitor's test result existed. Test rows were drawn by a fixed seed and were identical for every model: 500 per headline dataset, 250 for FLIVE and TAD66K, and all 89 labelled test rows of GFIQA-20k.
Metrics
The primary metric is Spearman rank correlation with the human ratings. Pairwise accuracy is the share of photo pairs, whose human ratings differ by at least half a standard deviation, that a model places in the correct order; clear-cut pairs differ by at least one standard deviation. Up to 2,000 pairs are used per output, and a tied prediction counts as half correct. Intervals are 95% bootstrap intervals from 2,000 draws, resampling linked images together, or whole pages for likability. A headline claim requires at least 500 rows and a paired interval on the gap that excludes zero. Reported p-values are two-sided bootstrap p-values from the same paired resampling, extended to 100,000 draws; where no draw reaches zero, the p-value is reported as below 0.00002.
Competitors
GPT-6 Sol ran through the OpenAI API, Claude Opus 5.5 through the Anthropic API and Claude Haiku 4.5 through Amazon Bedrock, all on September 26, 2026. GPT-6 Sol and Claude Opus 5.5 ran at low reasoning effort; their higher default effort did not improve on it by more than the 0.02 SRCC threshold. On each of the four headline datasets, every model chose its prompt from four variants using 100 development photos. Images were sent at each provider's native resolution, one call per image. Across 9,450 competitor test calls there were no refusals, no unparsable replies and no out-of-range scores.
Standalone scores
Where this page quotes Cwupid's score on its own, it uses the full locked test set: KonIQ-10k 0.942, CGFIQA-40k 0.987, PARA 0.923, AVA 0.803, GFIQA-20k 0.972, FLIVE 0.598 and TAD66K 0.493. Head-to-head figures use the shared test rows.
Scope and disclosures
- The Cwupid 1.0 entrant is a Cwupid 1.0 checkpoint that was never trained on the test post, though it may have seen other posts by the same person. It scored 64.7%, ahead of Claude Haiku 4.5 and statistically tied with Claude Opus 5.5 and GPT-6 Sol.
- An approximation in our harness sent 439 of Claude Haiku 4.5's 500 PARA images at a mean 95.5% of the pixels Claude's resize rule allows.
- Specialist image-quality models and open-weight vision models were not part of this evaluation. Every claim on this page concerns the three named models.
- Provider batch inputs were deleted after collection, OpenAI calls used
store: false, and no provider account opted in to training.