For the photo-rating tasks, every request carried this system prompt, followed by the photograph and the task prompt. Each reply was constrained to a JSON schema with the named numeric fields.
You are taking part in a published research benchmark of photo-rating models. Judge only the photograph as an image. Reply with the requested JSON and nothing else.
CGFIQA-40k
V1 · GPT-6 Sol
This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).
Predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "score".
V2 · not selected
This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).
First, in "notes", briefly assess sharpness of the face, noise, compression artifacts, exposure and lighting on the face, resolution, and anything that obscures the face. Then predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "notes" followed by "score".
V3 · Claude Opus 5.5, Claude Haiku 4.5
This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).
For reference: below 0.3 means the face is badly degraded (heavy blur, noise, compression or very poor lighting); around 0.5 is noticeably imperfect but usable; above 0.7 is a clean, sharp, well-lit face.
Predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "score".
V4 · not selected
This is a face image from the CGFIQA-40k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).
Predict the mean opinion score as a number from 0 to 1 with three decimal places. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "score".
KonIQ-10k
V1 · Claude Opus 5.5, Claude Haiku 4.5
This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).
Predict the mean opinion score as a number from 0 to 100 with one decimal place. Return JSON with "score".
V2 · not selected
This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).
First, in "notes", briefly assess sharpness or blur, noise, compression artifacts, exposure, color cast, and how severe the overall distortion is. Then predict the mean opinion score as a number from 0 to 100 with one decimal place. Return JSON with "notes" followed by "score".
V3 · GPT-6 Sol
This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).
For reference: below 30 means heavily distorted (strong blur, noise or artifacts); around 60 is a typical everyday photo with minor flaws; above 80 is clean and sharp with no visible distortion.
Predict the mean opinion score as a number from 0 to 100 with one decimal place. Return JSON with "score".
V4 · not selected
This is an everyday photograph from the KonIQ-10k image-quality dataset. Crowd workers rated its technical image quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a 1 to 5 scale. The ratings were averaged and rescaled to a mean opinion score from 0 (worst) to 100 (best).
Predict the mean opinion score as a number from 0 to 100 with one decimal place. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "score".
PARA
V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5
Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.
Predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Return JSON with "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".
V2 · not selected
Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.
First, in "notes", briefly assess composition, lighting, color, focus and depth of field, how interesting the subject is, and technical execution. Then predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Return JSON with "notes" followed by "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".
V3 · not selected
Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.
For reference: around 2 means clearly weak on that dimension; around 3 is an ordinary, competent photograph; above 4 is excellent on that dimension.
Predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Return JSON with "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".
V4 · not selected
Human raters scored this photograph from 1 (very poor) to 5 (excellent) for overall aesthetics and for six attributes: color, composition, content (how interesting the subject matter is), depth of field, light, and image quality (technical quality such as sharpness, noise and exposure). Each score is the average across raters.
Predict the average rating for each of the seven, as numbers from 1 to 5 with two decimal places. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "aesthetic", "color", "composition", "content", "depth_of_field", "light", "quality".
AVA
V1 · GPT-6 Sol
This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.
Predict the mean rating as a number from 1 to 10 with one decimal place. Return JSON with "score".
V2 · Claude Opus 5.5, Claude Haiku 4.5
This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.
First, in "notes", briefly assess composition, lighting, color, focus and depth of field, how interesting the subject is, and technical execution. Then predict the mean rating as a number from 1 to 10 with one decimal place. Return JSON with "notes" followed by "score".
V3 · not selected
This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.
For reference: around 3 means most voters found it poor (badly exposed, out of focus or uninteresting); around 5.5 is typical for the contest; above 7 is exceptional work that stands out among skilled amateur photographers.
Predict the mean rating as a number from 1 to 10 with one decimal place. Return JSON with "score".
V4 · not selected
This photograph was entered in a DPChallenge.com photography contest. Around 200 contest members each rated its aesthetic quality from 1 (lowest) to 10 (highest). The score is the mean of their ratings.
Predict the mean rating as a number from 1 to 10 with one decimal place. Make fine distinctions. This photograph is one of many being rated on the same scale, so give the precise value that places it exactly where it belongs, using the full range of the scale. Avoid defaulting to round or habitual numbers: photographs that differ even slightly should receive different scores. Return JSON with "score".
FLIVE
V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5
This is an everyday photograph from the LIVE-FB (PaQ-2-PiQ) image-quality dataset. Crowd workers rated its overall technical picture quality, meaning distortions such as blur, noise, compression artifacts and poor exposure, not the subject or its artistic merit, on a continuous scale from 0 (worst) to 100 (best). The score is the mean rating.
Predict the mean rating as a number from 0 to 100 with one decimal place. Return JSON with "score".
TAD66K
V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5
This photograph is from the TAD66K theme-oriented aesthetics dataset. Raters judged its aesthetic quality against other photographs of the same theme (for example portraits, landscapes, food or animals) on a 1 to 10 scale. The score is the average rating.
Predict the average rating as a number from 1 to 10 with one decimal place. Return JSON with "score".
GFIQA-20k
V1 · GPT-6 Sol, Claude Opus 5.5, Claude Haiku 4.5
This is a face image from the GFIQA-20k face image-quality dataset. Human observers rated the perceptual quality of the face image, meaning sharpness, noise, compression, lighting and other degradations of the face, not the person's appearance or identity. The score is a mean opinion score normalised from 0 (worst) to 1 (best).
Predict the mean opinion score as a number from 0 to 1 with three decimal places. Return JSON with "score".