cwupid Open Cwupid

Cwupid outperforms GPT-6 Sol and Claude Opus 5.5 at judging photographs, while being more than 10,000 times smaller*.

We evaluated Cwupid against three frontier vision models on seven public photo-rating datasets and predicted likes from real Instagram posts. Across photo rating and likability, Cwupid led 42 of 42 direct comparisons.

Cwupid 1.2 returns 13 photo-rating scores across face quality, technical quality and aesthetics from one 86 million-parameter model; Cwupid 1.1 predicts which of two posts earned more likes. Together they led all 42 comparisons with GPT-6 Sol, Claude Opus 5.5 and Claude Haiku 4.5. On face photos, Cwupid 1.2 makes 17 times fewer ordering mistakes than Claude Opus 5.5, the strongest competitor there. Cwupid 1.1 picks the more-liked post 70.7% of the time, compared with 63.2% for the next-best model.

500 photos and their corresponding human and Cwupid ranks. KonIQ-10k test set. Each point is one photograph.

Measuring against the frontier:

Each benchmark asks a model to rate photographs judged by people. The score is Spearman rank correlation (SRCC), which is how closely the model's ordering of the photos matches the human ordering, where 1.0 is identical. Every model saw the same test photos, and every axis below runs from 0 to 1.

KonIQ-10k · Technical image quality

Cwupid leads GPT-6 Sol by +0.103 SRCC

Cwupid picks the wrong photo in a pair 2.7% of the time. GPT-6 Sol does so 10.7% of the time, so Cwupid makes four times fewer mistakes. On the full KonIQ test set, Cwupid scores 0.942.

CGFIQA-40k · Face image quality

0.987 SRCC, and 2,000 of 2,000 clear-cut pairs

Cwupid ordered every clear-cut pair correctly, meaning every pair of faces whose quality ratings clearly differ. Across all pairs it made 17× fewer ordering errors than Claude Opus 5.5 (0.2% against 3.4%).

PARA · Overall aesthetics

Cwupid leads GPT-6 Sol by +0.136

When people clearly preferred one photo over another, Cwupid agreed 99.9% of the time. It also came first on all seven of PARA's attributes: aesthetics, color, composition, content, depth of field, light and quality.

PARA · Content

Cwupid leads Claude Opus 5.5 by +0.178 on content

Content asks how interesting a photo's subject is, which is about meaning more than pixels. It is also where Cwupid's lead over the best competitor was widest.

AVA · Photography-contest aesthetics

Cwupid leads GPT-6 Sol by +0.152

Each AVA score is the average vote of around 200 photography-contest members. Cwupid scores 0.803 on the full AVA test set.

Instagram likability · 561 held-out posts

70.7% within-page accuracy

Shown two posts from the same account, Cwupid 1.1 picked the one that earned more likes 70.7% of the time, from the photo alone: no caption, follower count or posting time. Guessing would score 50%, marked by the line.

The photo-rating comparisons.

Cwupid 1.2 led every displayed photo-rating row against the three frontier models. Across photo rating and Cwupid 1.1's likability test, the two models led 42 of 42 direct comparisons. General-purpose language models also sample their answers, so the same photo can get a different score each time you ask; Cwupid gives the same photo the same score every time.

Each row is one output, and each marker is one model's rank correlation with human scores on the same test photos. The number on the right is how far Cwupid leads the strongest competitor on that row. Read as text: Cwupid 1.2 scored 0.987 versus 0.939 on face quality, 0.937 versus 0.834 on KonIQ technical quality, 0.929 versus 0.793 on PARA overall aesthetics, and 0.846 versus 0.694 on AVA aesthetics. The full table gives all 13 outputs.

It orders photographs the way people do.

Each panel plots all 500 test photos: the human ranking across, the model's ranking up. Agreement collapses the cloud onto the diagonal. The horizontal bands show where a model gave many photos the same score. The frontier models keep landing on the same few numbers, most likely because they write their scores out as text instead of computing them.

Language models have favorite numbers.

Cwupid regresses a continuous score for every image. The general-purpose models write a number as text, and they return to the same few values: Claude Haiku 4.5 gave the identical KonIQ score of 72.5 to 264 of 500 photographs, and Claude Opus 5.5 gave 72.4 to 122. On KonIQ, GPT-6 Sol spread its answers more widely than either Claude model, but still used fewer than half as many distinct values as Cwupid: 159 against 401.

Each column is one score a model gave, and its height is how many of the 500 photos received it. All four rows share the same scale, and a score given to just one photo still shows as a small tick.

How Cwupid compares with larger published photo scorers.

The larger specialists below were built for narrower rating jobs, including models trained on KonIQ-10k quality or AVA aesthetics. Several have more than ten times Cwupid's parameters and still report lower scores on the task they target. Cwupid 1.2 reports higher figures on those tasks while using one 85.8M-parameter model for 13 ratings across technical quality, face quality, aesthetics and seven PARA attributes.

Cwupid 1.2 (85.8M parameters): Technical quality, face quality, aesthetics and seven PARA attributes, with 13 scores from one photograph.

Q-ReAlign Mini (0.8B): Technical quality, aesthetics and video quality.

OneAlign (about 8B): Technical quality, aesthetics and video quality in one model.

Q-Align (8.20B, about 96 times Cwupid's size): Its KonIQ-trained checkpoint rates technical image quality; the paper trains separate scorers for other rating tasks.

Q-Insight (8.29B, about 97 times Cwupid's size): Trained with KonIQ quality ratings for image quality and degradation perception, without face-quality or PARA scores.

Compare2Score (mPLUG-Owl2 with a 7B language model): Trained across image-quality datasets to score technical quality; it does not score faces, aesthetics or PARA attributes.

IAA-LQ (1B-parameter vision backbone alone, more than 11 times Cwupid's size): Trained on AVA for aesthetics only.

MUSIQ, TOPIQ, MANIQA and LIQE (about 27M to 150M): Technical quality, or aesthetics in a separately trained copy; usually one dataset per trained model.

The source links and full size-and-scope table are in the comparison below.

More than ten times larger, focused on fewer ratings, and still below Cwupid's published scores

KonIQ-10k, the specialists' technical-quality task: Cwupid 1.2 reports 0.942 SRCC across its full locked test set. The Q-Align checkpoint trained for KonIQ reports 0.940, despite using an 8.20B-parameter model, about 96 times Cwupid's size. Q-Insight, trained with KonIQ quality ratings, reports 0.916 with 8.29B parameters, about 97 times Cwupid's size. Both specialize in image-quality judgment; neither reports Cwupid's face, aesthetic and PARA outputs from the same model.

AVA, the aesthetics specialist's task: Cwupid 1.2 reports 0.803 SRCC across its full locked test set. IAA-LQ was trained for AVA aesthetics and reports 0.791. Its frozen vision backbone alone has about 1B parameters, more than 11 times Cwupid's entire model. It returns an aesthetic judgment, while Cwupid returns aesthetics alongside technical quality, face quality and the PARA attributes.

Further larger quality scorers: Compare2Score reports 0.931 KonIQ SRCC using an mPLUG-Owl2 architecture with a 7B-parameter language model, trained across image-quality datasets. Q-ReAlign Mini reports 0.935 KonIQ and 0.797 AVA with 0.8B parameters. Q-ReAlign does combine technical quality, aesthetics and video quality; it does not report face quality or the seven PARA attributes. These examples expand the list of larger published models with lower reported figures without treating every specialist as a single-dataset model.

Why the combination matters: The KonIQ and AVA specialists can devote their models to a narrower target, and their published scores are still below Cwupid's corresponding figures. Cwupid uses the same 85.8M weights for all 13 outputs across seven datasets and produces them in one pass. The result is a compact unified scorer whose published technical-quality and aesthetic scores stand up even beside much larger, more narrowly trained models.

Cwupid 1.2 is the state of the art in unified photo rating · How it compares with published models

Cwupid 1.2 is a unified, multi-task photo-rating model: a complete vision transformer of 85.8M parameters, fine-tuned end to end on 352,103 images, not a small head on a frozen backbone. It reads a photograph once and returns thirteen scores learned from human ratings, spanning the three tasks the research field divides photo judgment into: technical image quality (KonIQ-10k, FLIVE), face image quality (CGFIQA-40k, GFIQA-20k), aesthetics (AVA, TAD66K) and all seven of PARA's aesthetic attributes. Its companion model, Cwupid 1.1, predicts Instagram likability.

Cwupid 1.2 is the state of the art in unified photo rating across technical quality, face quality and aesthetics, and it sits on the frontier of the field in both capability and size. We know of no other model, at any size, that scores all three with one set of weights. Unified assessment is the field's own ambition, pursued by OneAlign, Q-ReAlign and LIQE, and Cwupid 1.2 is the one that also scores faces. Across Cwupid 1.2's photo-rating tests and Cwupid 1.1's likability test, they led the frontier general-purpose models GPT-6 Sol, Claude Opus 5.5 and Claude Haiku 4.5 in 42 of 42 direct comparisons. And at 85.8M parameters Cwupid 1.2 is the smallest unified photo-rating model we know of: the selected unified models nearest in scope cover technical quality and aesthetics but not faces, and run on 0.8B or about 8B parameters, from about 9 to 96 times its size. Both terms are used here as the field uses them: state of the art for the best result in a task or class of models, as Q-Align claims it for image quality and aesthetic assessment and MobileNetV3 for mobile models, and frontier for the leading edge on one capability or on efficiency, as OpenAI uses it when it says GPT-6 Astra "sets a new frontier on computer and browser use" and GPT-6 Sol is "advancing the frontier on cost efficiency".

Published image-assessment models by size and scope
ModelParametersWhat one trained model scoresBuilt on
Cwupid 1.285.8MTechnical quality, face quality, aesthetics and 7 PARA attributes: 13 scores in one passDINOv3 ViT-B/16, included in the count
Q-ReAlign Mini0.8BTechnical quality, aesthetics and video qualityQwen3.5-VL language model
OneAlignabout 8BTechnical quality, aesthetics and video qualitymPLUG-Owl2 language model
Q-Align8.20BTechnical quality in its KonIQ-trained checkpoint; other rating tasks use separate trained checkpointsmPLUG-Owl2 language model
Q-Insight8.29BImage-quality ratings and degradation perceptionQwen2.5-VL-7B-Instruct
Compare2Score7B language model plus vision componentsTechnical image quality across multiple datasetsmPLUG-Owl2 architecture
IAA-LQabout 1BAesthetics onlyFrozen EVA-CLIP ViT-G/14
MUSIQ, TOPIQ, MANIQA, LIQEabout 27M to 150MTechnical quality, or aesthetics in a separately trained copy; usually one dataset per trained modelImageNet or CLIP backbones

Most of these models are trained and tested one dataset at a time, each on its own home ground, as were earlier specialists such as NIMA and HyperIQA, and face-quality models such as DSL-FIQA score faces only. Covering seven datasets with a single set of weights is the harder test, and it is the one every real photograph sets: its technical quality, its faces and its aesthetics all need judging at once. Every model here starts from a pretrained backbone. Cwupid 1.2's 85.8M is its complete inference path, backbone included, with no face detector or crop. Other models' sizes are as published by their authors.

Is Cwupid a frontier model? Yes, at photo judgment · Questions answered

Is Cwupid state of the art?

Yes. State of the art is claimed task by task and class by class across the field: Q-Align reports "state-of-the-art performance on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) tasks", and Google's MobileNetV3 reports "new state of the art results for mobile classification, detection and segmentation". In that established sense, Cwupid 1.2 is the state of the art in unified photo rating: one 85.8M-parameter model that scores technical quality, face quality, aesthetics and PARA's seven attributes in a single pass, and no published model we know of, at any size, covers all of them with one set of weights. Q-Align's authors went on to "unify the three tasks into one model, termed the OneAlign"; Cwupid 1.2 covers more of a photograph than OneAlign or any other unified scorer, adding faces and PARA's seven attributes at a fraction of their size. Across Cwupid 1.2's photo-rating tests and Cwupid 1.1's likability test, they led GPT-6 Sol, Claude Opus 5.5 and Claude Haiku 4.5 in 42 of 42 direct comparisons. On its full locked test set Cwupid 1.2 scores KonIQ-10k 0.942, CGFIQA-40k 0.987 and PARA 0.923 SRCC.

Is Cwupid a frontier model?

Yes. OpenAI uses the word in both of these senses: GPT-6 Astra "sets a new frontier on computer and browser use", a frontier on one capability, and GPT-6 Sol and Luna are "advancing the frontier on cost efficiency", a frontier on efficiency (GPT-6 Astra announcement, GPT-6 Sol announcement). By both measures, Cwupid is at the frontier on photo judgment. Across Cwupid 1.2's photo-rating tests and Cwupid 1.1's likability test, they outperformed GPT-6 Sol and Claude Opus 5.5, today's frontier general-purpose models, and Claude Haiku 4.5 in 42 of 42 direct comparisons. On efficiency, Cwupid 1.2 sits on the frontier of accuracy against size: we know of no smaller unified photo-rating model, and none of any size that also scores face quality. General-purpose frontier models are measured across many tasks; on this one, they trail a model of 86 million parameters.

Why does photo judgment matter?

Because the demand is enormous and the judgment is not one task. Humanity took an estimated 2.1 trillion photos in 2025 (PetaPixel), and people now ask chatbots to judge theirs (TechRadar). The research field puts it directly: "the explosion of visual content available online underscores the requirement for an accurate machine assessor" (Q-Align). It also divides photo judgment into separate tasks, each with its own datasets and leaderboards: technical image quality assessment, image aesthetic assessment, and face image quality assessment, which its researchers call "crucial in improving image restoration algorithms and selecting high-quality face images" (DSL-FIQA). Cwupid 1.2 covers all three in one model.

How does an 86-million-parameter model beat GPT-6 Sol and Claude Opus 5.5?

By being built for the judgment end to end. Cwupid 1.2 is a complete vision transformer, a DINOv3 ViT-B/16 fine-tuned end to end for photo rating with its own learned spatial pooler and readout, trained on 352,103 images. It regresses all 13 scores as continuous values in one pass, so photographs that differ slightly get different scores: across 500 KonIQ photos it returned 401 distinct values. General-purpose models write their answer as text and keep landing on favorite numbers: Claude Haiku 4.5 gave the identical score of 72.5 to 264 of those 500 photos. Cwupid's path has no face detector, no crop and no prompt, only the photograph.

Why compare Cwupid with general-purpose models?

Because frontier general-purpose models are what people actually ask to judge their photos. Rating a photo is not an off-label task for them: judging image quality and aesthetics is a standard test of general-purpose vision models (Q-Bench), and the leading unified scorers, Q-Align, OneAlign and Q-ReAlign, are multimodal language models fine-tuned for exactly this. Writing a score as text is how these models judge photos for everyone who asks them, so it is the capability being measured, not a handicap placed on it. If sheer scale decided photo judgment, the frontier models would have won. They did not.

Would GPT-6 Astra have changed the result?

OpenAI itself presents GPT-6 Sol as frontier intelligence: it trained Sol "with similar methods as GPT-6 Astra" and introduces it as one of "more ways to bring frontier intelligence into the work you do every day" (GPT-6 Sol announcement). The areas where OpenAI says Astra is state of the art are computer use, browsing, software engineering, cybersecurity, science and professional work (GPT-6 Astra announcement), and neither announcement reports an image-quality or aesthetics result for either model. On photo judgment, GPT-6 Sol trailed Cwupid throughout the reported results.

Did the frontier models get a fair fight?

Each model chose among four prompts, including rubric and anchored versions, using 100 development photos. Across all 9,450 competitor test calls there were no refusals and no unreadable replies.

How well does Cwupid predict likes?

Better than every frontier model tested. Shown two posts from the same Instagram account, Cwupid 1.1 picked the one that earned more likes 70.7% of the time, against 63.2% for Claude Opus 5.5, 61.5% for GPT-6 Sol and 56.9% for Claude Haiku 4.5. Every lead holds after correcting for all 42 comparisons on this page.

How large are the models Cwupid is compared with?

Specialist quality models such as MUSIQ, TOPIQ, MANIQA and LIQE have about 27M to 150M parameters, and each trained copy covers one kind of rating. IAA-LQ scores aesthetics only, on a backbone of about 1B parameters. OneAlign (about 8B, on the mPLUG-Owl2 language model) and Q-ReAlign Mini (0.8B), the selected unified scorers nearest in scope, were trained on two of Cwupid's seven datasets, KonIQ-10k and AVA, and have no face-quality or PARA outputs. GPT-6 Sol and Claude Opus 5.5 have not published their sizes; Kimi K3, the largest open-weight model, has 2.8 trillion parameters, about 32,600 times Cwupid 1.2.

Is Cwupid one model or several?

One. Cwupid 1.2 is a single set of weights: an adapted DINOv3 ViT-B/16 backbone, one learned spatial pooler and one linear readout. It reads the whole photograph once, with no face detector or crop, and all 13 scores come from the same forward pass. Its 85.8M parameters are the complete inference path, backbone included.

Is the size comparison like for like?

Yes. Every model compared here starts from a pretrained backbone: Q-Align from mPLUG-Owl2, IAA-LQ from a frozen EVA-CLIP ViT-G/14, and the smaller specialists from ImageNet or CLIP backbones. Cwupid 1.2 starts from DINOv3 ViT-B/16, fine-tuned end to end, and its 85.8M count includes that backbone.

Can these results be checked?

Yes. The test rows, prompts, metrics and decision rules were fixed before any competitor's test result existed. Every prompt is published word for word in Appendix A, every dataset is public, and all 42 wins hold after correcting for multiple comparisons (largest adjusted p = 0.0061), so anyone can rerun the published prompts against the same public datasets.

Every task, every model.

Rank correlation with human scores, with every model scored on the same test photos. The last column is Cwupid's lead over the strongest competitor on each row, with its 95% confidence interval. Rows marked table-only have fewer than 500 test photos; their sample sizes appear in the table.

Photo rating: Spearman rank correlation with human ratings
Dataset · outputTest rowsCwupid 1.2GPT-6 SolClaude Opus 5.5Claude Haiku 4.5Lead over best competitor
CGFIQA-40k · Face image quality5000.9870.9320.9390.751+0.048 vs Claude Opus 5.595% CI [+0.039, +0.059]
KonIQ-10k · Technical image quality5000.9370.8340.7790.658+0.103 vs GPT-6 Sol95% CI [+0.075, +0.133]
PARA · Overall aesthetics5000.9290.7930.7630.762+0.136 vs GPT-6 Sol95% CI [+0.101, +0.172]
PARA · Color5000.9080.8140.7820.748+0.094 vs GPT-6 Sol95% CI [+0.064, +0.125]
PARA · Composition5000.9080.7530.7580.758+0.149 vs Claude Opus 5.595% CI [+0.113, +0.191]
PARA · Content5000.8940.7010.7160.685+0.178 vs Claude Opus 5.595% CI [+0.137, +0.221]
PARA · Depth of field5000.8920.6630.7620.766+0.127 vs Claude Haiku 4.595% CI [+0.088, +0.167]
PARA · Light5000.8940.7670.7190.675+0.127 vs GPT-6 Sol95% CI [+0.092, +0.163]
PARA · Image quality5000.9250.8070.7980.792+0.118 vs GPT-6 Sol95% CI [+0.088, +0.152]
AVA · Aesthetics5000.8460.6940.6870.488+0.152 vs GPT-6 Sol95% CI [+0.108, +0.197]
FLIVE · Image qualitytable only2500.5520.3960.3110.262+0.156 vs GPT-6 Sol95% CI [+0.054, +0.260]
TAD66K · Aestheticstable only2500.5410.3870.3710.321+0.154 vs GPT-6 Sol95% CI [+0.057, +0.246]
GFIQA-20k · Face image qualitytable only890.9720.8080.7570.599+0.164 vs GPT-6 Sol95% CI [+0.087, +0.260]
Pairwise accuracy on headline outputs: all pairs / clear-cut pairs
Dataset · outputCwupid 1.2GPT-6 SolClaude Opus 5.5Claude Haiku 4.5
CGFIQA-40k · Face image quality99.8% / 100.0%96.2% / 99.5%96.6% / 99.5%83.8% / 90.0%
KonIQ-10k · Technical image quality97.3% / 99.3%89.3% / 96.1%86.9% / 94.4%77.6% / 85.5%
PARA · Overall aesthetics97.3% / 99.9%90.0% / 95.1%87.7% / 93.5%88.1% / 92.1%
PARA · Color95.7% / 99.1%89.2% / 94.3%87.2% / 92.5%85.9% / 91.0%
PARA · Composition95.8% / 99.5%85.4% / 92.3%85.8% / 92.9%86.1% / 91.6%
PARA · Content94.4% / 99.0%83.3% / 89.1%84.8% / 90.5%83.8% / 88.9%
PARA · Depth of field95.2% / 99.1%82.7% / 90.6%86.4% / 93.3%87.4% / 93.0%
PARA · Light94.8% / 99.1%87.1% / 92.5%84.2% / 90.2%81.1% / 88.9%
PARA · Image quality97.5% / 99.9%91.1% / 97.2%89.9% / 96.5%90.0% / 95.9%
AVA · Aesthetics90.0% / 94.8%80.8% / 87.8%82.3% / 86.7%71.3% / 75.4%

How we measured.

Fixed standards

The test rows, prompts, metrics and decision rules were fixed before any competitor's test result existed. Test rows were drawn by a fixed seed and were identical for every model: 500 per headline dataset, 250 for FLIVE and TAD66K, and all 89 labelled test rows of GFIQA-20k.

Metrics

The primary metric is Spearman rank correlation with the human ratings. Pairwise accuracy is the share of photo pairs, whose human ratings differ by at least half a standard deviation, that a model places in the correct order; clear-cut pairs differ by at least one standard deviation. Up to 2,000 pairs are used per output, and a tied prediction counts as half correct. Intervals are 95% bootstrap intervals from 2,000 draws, resampling linked images together, or whole pages for likability. Reported p-values are two-sided bootstrap p-values from the same paired resampling, extended to 100,000 draws; where no draw reaches zero, the p-value is reported as below 0.00002.

Competitors

GPT-6 Sol ran through the OpenAI API, Claude Opus 5.5 through the Anthropic API and Claude Haiku 4.5 through Amazon Bedrock, all on September 26, 2026. On each of the four main photo-rating datasets, every model chose its prompt from four variants using 100 development photos. Images were sent at each provider's native resolution, one call per image. Across 9,450 competitor test calls there were no refusals, no unparsable replies and no out-of-range scores.

Standalone scores

Where this page quotes Cwupid's score on its own, it uses the full locked test set: KonIQ-10k 0.942, CGFIQA-40k 0.987, PARA 0.923, AVA 0.803, GFIQA-20k 0.972, FLIVE 0.598 and TAD66K 0.493. Head-to-head figures use the shared test rows.

Additional method details

  • Provider batch inputs were deleted after collection, OpenAI calls used store: false, and no provider account opted in to training.
Appendix A · Photo-rating prompts, word for word

For the photo-rating tasks, every request carried this system prompt, followed by the photograph and the task prompt. Each reply was constrained to a JSON schema with the named numeric fields.

You are taking part in a published research benchmark of photo-rating models. Judge only the photograph as an image. Reply with the requested JSON and nothing else.

CGFIQA-40k

V1 · GPT-6 Sol

This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).

Predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "score".

V2 · not selected

This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).

First, in "notes", briefly assess sharpness of the face, noise, compression artifacts, exposure and lighting on the face, resolution, and anything that obscures the face. Then predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "notes" followed by "score".

V3 · Claude Opus 5.5, Claude Haiku 4.5

This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).

For reference: below 0.3 means the face is badly degraded (heavy blur, noise, compression or very poor lighting); around 0.5 is noticeably imperfect but usable; above 0.7 is a clean, sharp, well-lit face.

Predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "score".

V4 · not selected

This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).

Predict the mean opinion score as a number from 0 to 1 with three decimal places. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "score".

KonIQ-10k

V1 · Claude Opus 5.5, Claude Haiku 4.5

This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).

Predict the mean opinion score as a number from 0 to 100 with one decimal place. Return JSON with "score".

V2 · not selected

This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).

First, in "notes", briefly assess sharpness or blur, noise, compression artifacts, exposure, color cast, and how severe the overall distortion is. Then predict the mean opinion score as a number from 0 to 100 with one decimal place. Return JSON with "notes" followed by "score".

V3 · GPT-6 Sol

This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).

For reference: below 30 means heavily distorted (strong blur, noise or artifacts); around 60 is a typical everyday photo with minor flaws; above 80 is clean and sharp with no visible distortion.

Predict the mean opinion score as a number from 0 to 100 with one decimal place. Return JSON with "score".

V4 · not selected

This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).

Predict the mean opinion score as a number from 0 to 100 with one decimal place. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "score".

PARA

V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5

Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.

Predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Return JSON with "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".

V2 · not selected

Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.

First, in "notes", briefly assess composition, lighting, color, focus and depth of field, how interesting the subject is, and technical execution. Then predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Return JSON with "notes" followed by "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".

V3 · not selected

Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.

For reference: around 2 means clearly weak on that dimension; around 3 is an ordinary, competent photograph; above 4 is excellent on that dimension.

Predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Return JSON with "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".

V4 · not selected

Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.

Predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".

AVA

V1 · GPT-6 Sol

This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.

Predict the mean rating as a number from 1 to 10 with one decimal place. Return JSON with "score".

V2 · Claude Opus 5.5, Claude Haiku 4.5

This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.

First, in "notes", briefly assess composition, lighting, color, focus and depth of field, how interesting the subject is, and technical execution. Then predict the mean rating as a number from 1 to 10 with one decimal place. Return JSON with "notes" followed by "score".

V3 · not selected

This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.

For reference: around 3 means most voters found it poor (badly exposed, out of focus or uninteresting); around 5.5 is typical for the contest; above 7 is exceptional work that stands out among skilled amateur photographers.

Predict the mean rating as a number from 1 to 10 with one decimal place. Return JSON with "score".

V4 · not selected

This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.

Predict the mean rating as a number from 1 to 10 with one decimal place. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "score".

FLIVE

V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5

This is an everyday photograph from the LIVE-FB (PaQ-2-PiQ) image-quality dataset. Crowd workers rated its overall technical picture quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a continuous scale from 0 (worst) to 100 (best). The score is the mean rating.

Predict the mean rating as a number from 0 to 100 with one decimal place. Return JSON with "score".

TAD66K

V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5

This photograph is from the TAD66K theme-oriented aesthetics dataset. Raters judged its aesthetic quality against other photographs of the same theme (for example portraits, landscapes, food or animals) on a 1 to 10 scale. The score is the average rating.

Predict the average rating as a number from 1 to 10 with one decimal place. Return JSON with "score".

GFIQA-20k

V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5

This is a face image from the GFIQA-20k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).

Predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "score".
Appendix B · Tested models' prompt selection on development data

Each model ran four prompt variants on the same 100 development photos per task. The highest development SRCC was selected, and V1 was kept unless another variant beat it by more than 0.02. PARA uses the mean of its seven outputs. V1 asks for the dataset score; V2 adds a rubric; V3 adds low, middle and high anchors; V4 asks for fine distinctions.

GPT-6 Sol, CGFIQA-40k: development SRCC V1 / V2 / V3 / V4 = 0.911 / 0.911 / 0.912 / 0.915; selected V1.

GPT-6 Sol, KonIQ-10k: development SRCC V1 / V2 / V3 / V4 = 0.806 / 0.834 / 0.843 / 0.801; selected V3.

GPT-6 Sol, PARA: development SRCC V1 / V2 / V3 / V4 = 0.674 / 0.677 / 0.674 / 0.691; selected V1.

GPT-6 Sol, AVA: development SRCC V1 / V2 / V3 / V4 = 0.719 / 0.678 / 0.702 / 0.682; selected V1.

Claude Opus 5.5, CGFIQA-40k: development SRCC V1 / V2 / V3 / V4 = 0.891 / 0.921 / 0.922 / 0.902; selected V3.

Claude Opus 5.5, KonIQ-10k: development SRCC V1 / V2 / V3 / V4 = 0.818 / 0.813 / 0.824 / 0.799; selected V1.

Claude Opus 5.5, PARA: development SRCC V1 / V2 / V3 / V4 = 0.692 / 0.198 / 0.710 / 0.695; selected V1.

Claude Opus 5.5, AVA: development SRCC V1 / V2 / V3 / V4 = 0.721 / 0.742 / 0.705 / 0.715; selected V2.

Claude Haiku 4.5, CGFIQA-40k: development SRCC V1 / V2 / V3 / V4 = 0.730 / 0.757 / 0.795 / 0.764; selected V3.

Claude Haiku 4.5, KonIQ-10k: development SRCC V1 / V2 / V3 / V4 = 0.559 / 0.552 / 0.490 / 0.549; selected V1.

Claude Haiku 4.5, PARA: development SRCC V1 / V2 / V3 / V4 = 0.674 / 0.681 / 0.657 / 0.687; selected V1.

Claude Haiku 4.5, AVA: development SRCC V1 / V2 / V3 / V4 = 0.572 / 0.638 / 0.567 / 0.575; selected V2.

Put the benchmark leader to work on your photos.

Cwupid 1.1 and Cwupid 1.2, the models benchmarked on this page, are the models behind Cwupid's scores and photo reports.

Try Cwupid

Complete results in text

Text equivalents of the photo-rating, pairwise and likability results above.

All photo rating scores in text

Each line gives Spearman rank correlation for models that scored the same held-out photos. These are the values plotted in the output chart.

CGFIQA-40k · Face image quality: 500 test photos; SRCC: Cwupid 1.2 0.987, GPT-6 Sol 0.932, Claude Opus 5.5 0.939, Claude Haiku 4.5 0.751. Lead over Claude Opus 5.5: +0.048 (95% CI [+0.039, +0.059]).

KonIQ-10k · Technical image quality: 500 test photos; SRCC: Cwupid 1.2 0.937, GPT-6 Sol 0.834, Claude Opus 5.5 0.779, Claude Haiku 4.5 0.658. Lead over GPT-6 Sol: +0.103 (95% CI [+0.075, +0.133]).

PARA · Overall aesthetics: 500 test photos; SRCC: Cwupid 1.2 0.929, GPT-6 Sol 0.793, Claude Opus 5.5 0.763, Claude Haiku 4.5 0.762. Lead over GPT-6 Sol: +0.136 (95% CI [+0.101, +0.172]).

PARA · Color: 500 test photos; SRCC: Cwupid 1.2 0.908, GPT-6 Sol 0.814, Claude Opus 5.5 0.782, Claude Haiku 4.5 0.748. Lead over GPT-6 Sol: +0.094 (95% CI [+0.064, +0.125]).

PARA · Composition: 500 test photos; SRCC: Cwupid 1.2 0.908, GPT-6 Sol 0.753, Claude Opus 5.5 0.758, Claude Haiku 4.5 0.758. Lead over Claude Opus 5.5: +0.149 (95% CI [+0.113, +0.191]).

PARA · Content: 500 test photos; SRCC: Cwupid 1.2 0.894, GPT-6 Sol 0.701, Claude Opus 5.5 0.716, Claude Haiku 4.5 0.685. Lead over Claude Opus 5.5: +0.178 (95% CI [+0.137, +0.221]).

PARA · Depth of field: 500 test photos; SRCC: Cwupid 1.2 0.892, GPT-6 Sol 0.663, Claude Opus 5.5 0.762, Claude Haiku 4.5 0.766. Lead over Claude Haiku 4.5: +0.127 (95% CI [+0.088, +0.167]).

PARA · Light: 500 test photos; SRCC: Cwupid 1.2 0.894, GPT-6 Sol 0.767, Claude Opus 5.5 0.719, Claude Haiku 4.5 0.675. Lead over GPT-6 Sol: +0.127 (95% CI [+0.092, +0.163]).

PARA · Image quality: 500 test photos; SRCC: Cwupid 1.2 0.925, GPT-6 Sol 0.807, Claude Opus 5.5 0.798, Claude Haiku 4.5 0.792. Lead over GPT-6 Sol: +0.118 (95% CI [+0.088, +0.152]).

AVA · Aesthetics: 500 test photos; SRCC: Cwupid 1.2 0.846, GPT-6 Sol 0.694, Claude Opus 5.5 0.687, Claude Haiku 4.5 0.488. Lead over GPT-6 Sol: +0.152 (95% CI [+0.108, +0.197]).

FLIVE · Image quality (table only): 250 test photos; SRCC: Cwupid 1.2 0.552, GPT-6 Sol 0.396, Claude Opus 5.5 0.311, Claude Haiku 4.5 0.262. Lead over GPT-6 Sol: +0.156 (95% CI [+0.054, +0.260]).

TAD66K · Aesthetics (table only): 250 test photos; SRCC: Cwupid 1.2 0.541, GPT-6 Sol 0.387, Claude Opus 5.5 0.371, Claude Haiku 4.5 0.321. Lead over GPT-6 Sol: +0.154 (95% CI [+0.057, +0.246]).

GFIQA-20k · Face image quality (table only): 89 test photos; SRCC: Cwupid 1.2 0.972, GPT-6 Sol 0.808, Claude Opus 5.5 0.757, Claude Haiku 4.5 0.599. Lead over GPT-6 Sol: +0.164 (95% CI [+0.087, +0.260]).

Pairwise accuracy in text

For each model, the first percentage covers all eligible pairs; the second covers clear-cut pairs.

CGFIQA-40k · Face image quality: Cwupid 1.2 99.8% / 100.0%, GPT-6 Sol 96.2% / 99.5%, Claude Opus 5.5 96.6% / 99.5%, Claude Haiku 4.5 83.8% / 90.0%.

KonIQ-10k · Technical image quality: Cwupid 1.2 97.3% / 99.3%, GPT-6 Sol 89.3% / 96.1%, Claude Opus 5.5 86.9% / 94.4%, Claude Haiku 4.5 77.6% / 85.5%.

PARA · Overall aesthetics: Cwupid 1.2 97.3% / 99.9%, GPT-6 Sol 90.0% / 95.1%, Claude Opus 5.5 87.7% / 93.5%, Claude Haiku 4.5 88.1% / 92.1%.

PARA · Color: Cwupid 1.2 95.7% / 99.1%, GPT-6 Sol 89.2% / 94.3%, Claude Opus 5.5 87.2% / 92.5%, Claude Haiku 4.5 85.9% / 91.0%.

PARA · Composition: Cwupid 1.2 95.8% / 99.5%, GPT-6 Sol 85.4% / 92.3%, Claude Opus 5.5 85.8% / 92.9%, Claude Haiku 4.5 86.1% / 91.6%.

PARA · Content: Cwupid 1.2 94.4% / 99.0%, GPT-6 Sol 83.3% / 89.1%, Claude Opus 5.5 84.8% / 90.5%, Claude Haiku 4.5 83.8% / 88.9%.

PARA · Depth of field: Cwupid 1.2 95.2% / 99.1%, GPT-6 Sol 82.7% / 90.6%, Claude Opus 5.5 86.4% / 93.3%, Claude Haiku 4.5 87.4% / 93.0%.

PARA · Light: Cwupid 1.2 94.8% / 99.1%, GPT-6 Sol 87.1% / 92.5%, Claude Opus 5.5 84.2% / 90.2%, Claude Haiku 4.5 81.1% / 88.9%.

PARA · Image quality: Cwupid 1.2 97.5% / 99.9%, GPT-6 Sol 91.1% / 97.2%, Claude Opus 5.5 89.9% / 96.5%, Claude Haiku 4.5 90.0% / 95.9%.

AVA · Aesthetics: Cwupid 1.2 90.0% / 94.8%, GPT-6 Sol 80.8% / 87.8%, Claude Opus 5.5 82.3% / 86.7%, Claude Haiku 4.5 71.3% / 75.4%.

Likability results in text

Cwupid 1.1: within-page accuracy 70.7% (95% CI [65.9%, 75.1%]); clear-cut pairs 78.0%; mean within-page SRCC 0.533; 561 distinct scores; 0.0% tied pairs.

GPT-6 Sol: within-page accuracy 61.5% (95% CI [57.1%, 66.4%]); clear-cut pairs 66.8%; mean within-page SRCC 0.310; 88 distinct scores; 3.4% tied pairs.

Claude Opus 5.5: within-page accuracy 63.2% (95% CI [58.8%, 67.4%]); clear-cut pairs 68.3%; mean within-page SRCC 0.361; 44 distinct scores; 9.6% tied pairs.

Claude Haiku 4.5: within-page accuracy 56.9% (95% CI [51.8%, 61.6%]); clear-cut pairs 59.9%; mean within-page SRCC 0.211; 21 distinct scores; 27.3% tied pairs.

* Parameter counts for GPT-6 Sol and Claude Opus 5.5 have not been published. Size comparisons on this page assume each has more parameters than Moonshot AI's Kimi K3, the largest open-weight model, at 2.8 trillion. Kimi K3 alone is about 32,600 times the size of Cwupid 1.2, which has 85.8 million parameters.