cwupid Open Cwupid

Cwupid outperforms GPT-6 Sol and Claude Opus 5.5 at judging photographs, while being more than 10,000 times smaller.

We evaluated Cwupid against three frontier vision models on seven public photo-rating datasets and predicted likes from real Instagram posts. Cwupid achieved first on all of the 14 metrics by a landslide.

Cwupid won 42 of 42 direct comparisons against GPT-6 Sol, Claude Opus 5.5 and Claude Haiku 4.5. Cwupid 1.2 ranks first on all 13 photo-rating outputs, across face quality, technical quality and aesthetics. On face photos it makes 17 times fewer ordering mistakes than Claude Opus 5.5, the strongest competitor there. Shown two Instagram posts from the same account, Cwupid 1.1 picks the one that earned more likes 70.7% of the time, compared with 63.2% for the next-best model, all with 86 million parameters.*

500 photos and their corresponding human and Cwupid ranks. KonIQ-10k test set. Each point is one photograph.

Measuring against the frontier:

Each benchmark asks a model to rate photographs judged by people. The score is Spearman rank correlation (SRCC), which is how closely the model's ordering of the photos matches the human ordering, where 1.0 is identical. Every model saw the same test photos, and every axis below runs from 0 to 1.

KonIQ-10k · Technical image quality

Cwupid leads GPT-6 Sol by +0.103 SRCC

Cwupid picks the wrong photo in a pair 2.7% of the time. GPT-6 Sol does so 10.7% of the time, so Cwupid makes four times fewer mistakes. On the full KonIQ test set, Cwupid scores 0.942.

CGFIQA-40k · Face image quality

0.987 SRCC, and 2,000 of 2,000 clear-cut pairs

Cwupid ordered every clear-cut pair correctly, meaning every pair of faces whose quality ratings clearly differ. Across all pairs it made 17× fewer ordering errors than Claude Opus 5.5 (0.2% against 3.4%).

PARA · Overall aesthetics

Cwupid leads GPT-6 Sol by +0.136

When people clearly preferred one photo over another, Cwupid agreed 99.9% of the time. It also came first on all seven of PARA's attributes: aesthetics, color, composition, content, depth of field, light and quality.

PARA · Content

Cwupid leads Claude Opus 5.5 by +0.178 on content

Content asks how interesting a photo's subject is, which is about meaning more than pixels. It is also where Cwupid's lead over the best competitor was widest.

AVA · Photography-contest aesthetics

Cwupid leads GPT-6 Sol by +0.152

Each AVA score is the average vote of around 200 photography-contest members. Cwupid scores 0.803 on the full AVA test set; the 500-photo sample used here happens to run a little higher.

Instagram likability · 561 held-out posts

70.7% within-page accuracy

Shown two posts from the same account, Cwupid 1.1 picked the one that earned more likes 70.7% of the time, from the photo alone: no caption, follower count or posting time. Guessing would score 50%, marked by the line.

First on all 13 photo-rating outputs.

Cwupid beat three frontier models on all seven datasets and thirteen outputs. On every row Cwupid 1.2 sits to the right of every other model. General-purpose language models also sample their answers, so the same photo can get a different score each time you ask; Cwupid gives the same photo the same score every time.

Each row is one output, and each marker is one model's rank correlation with human scores on the same test photos. The number on the right is how far Cwupid leads the strongest competitor on that row.

It orders photographs the way people do.

Each panel plots all 500 test photos: the human ranking across, the model's ranking up. Agreement collapses the cloud onto the diagonal. The horizontal bands show where a model gave many photos the same score. The frontier models keep landing on the same few numbers, most likely because they write their scores out as text instead of computing them.

Language models have favorite numbers.

Cwupid regresses a continuous score for every image. The general-purpose models write a number as text, and they return to the same few values: Claude Haiku 4.5 gave the identical KonIQ score of 72.5 to 264 of 500 photographs, and Claude Opus 5.5 gave 72.4 to 122. On KonIQ, GPT-6 Sol spread its answers more widely than either Claude model, but still used fewer than half as many distinct values as Cwupid: 159 against 401.

Each column is one score a model gave, and its height is how many of the 500 photos received it. All four rows share the same scale, and a score given to just one photo still shows as a small tick.

Cwupid is built for this.

Frontier models are trained to do a little of everything. Cwupid is trained for this one task, and our benchmarks show the advantages of specialization.

Cwupid 1.2

85.8M

parameters in the complete scoring model: an adapted DINOv3 ViT-B/16 vision transformer with Cwupid's own spatial pooling and readout. Each score costs fractions of a cent in compute, and is tens of thousands of times smaller than open source LLMs.*

One pass, the whole image

Cwupid reads the full photograph once and returns every score together. There is no face detector, no crop and no prompt to tune.

Continuous by construction

Scores are regressed, not written out as text, so two photographs that differ slightly receive different scores. Across 500 KonIQ photos Cwupid returned 401 distinct values.

Every task, every model.

Rank correlation with human scores, with every model scored on the same test photos. The last column is Cwupid's lead over the strongest competitor on each row, with its 95% confidence interval. Rows marked table-only have fewer than 500 test photos, so we report them here but don't use them for headline claims.

Photo rating: Spearman rank correlation with human ratings
Pairwise accuracy on headline outputs: all pairs / clear-cut pairs

How we measured.

Fixed standards

The test rows, prompts, metrics and decision rules were fixed before any competitor's test result existed. Test rows were drawn by a fixed seed and were identical for every model: 500 per headline dataset, 250 for FLIVE and TAD66K, and all 89 labelled test rows of GFIQA-20k.

Metrics

The primary metric is Spearman rank correlation with the human ratings. Pairwise accuracy is the share of photo pairs, whose human ratings differ by at least half a standard deviation, that a model places in the correct order; clear-cut pairs differ by at least one standard deviation. Up to 2,000 pairs are used per output, and a tied prediction counts as half correct. Intervals are 95% bootstrap intervals from 2,000 draws, resampling linked images together, or whole pages for likability. A headline claim requires at least 500 rows and a paired interval on the gap that excludes zero. Reported p-values are two-sided bootstrap p-values from the same paired resampling, extended to 100,000 draws; where no draw reaches zero, the p-value is reported as below 0.00002.

Competitors

GPT-6 Sol ran through the OpenAI API, Claude Opus 5.5 through the Anthropic API and Claude Haiku 4.5 through Amazon Bedrock, all on September 26, 2026. GPT-6 Sol and Claude Opus 5.5 ran at low reasoning effort; their higher default effort did not improve on it by more than the 0.02 SRCC threshold. On each of the four headline datasets, every model chose its prompt from four variants using 100 development photos. Images were sent at each provider's native resolution, one call per image. Across 9,450 competitor test calls there were no refusals, no unparsable replies and no out-of-range scores.

Standalone scores

Where this page quotes Cwupid's score on its own, it uses the full locked test set: KonIQ-10k 0.942, CGFIQA-40k 0.987, PARA 0.923, AVA 0.803, GFIQA-20k 0.972, FLIVE 0.598 and TAD66K 0.493. Head-to-head figures use the shared test rows.

Scope and disclosures

  • The Cwupid 1.0 entrant is a Cwupid 1.0 checkpoint that was never trained on the test post, though it may have seen other posts by the same person. It scored 64.7%, ahead of Claude Haiku 4.5 and statistically tied with Claude Opus 5.5 and GPT-6 Sol.
  • An approximation in our harness sent 439 of Claude Haiku 4.5's 500 PARA images at a mean 95.5% of the pixels Claude's resize rule allows.
  • Specialist image-quality models and open-weight vision models were not part of this evaluation. Every claim on this page concerns the three named models.
  • Provider batch inputs were deleted after collection, OpenAI calls used store: false, and no provider account opted in to training.
Appendix A · Prompts, word for word
Appendix B · Prompt selection on development data

Put the benchmark leader to work on your photos.

Cwupid 1.1 and Cwupid 1.2, the models benchmarked on this page, are the models behind Cwupid's scores and photo reports.

Try Cwupid

* Parameter counts for GPT-6 Sol and Claude Opus 5.5 have not been published. Size comparisons on this page assume each has more parameters than Moonshot AI's Kimi K3, the largest open-weight model, at 2.8 trillion. Kimi K3 alone is about 32,600 times the size of Cwupid 1.2, which has 85.8 million parameters. Cwupid 1.1, which produced the likability result, runs 180 million parameters in total.