+ Pavel Klinger / Benchmark / October 2026

What gpt-image-2.5 draws when you describe everything, and which model can check it

300 images, three judge models, one human.

I make short factual explainers. Their pictures need background and atmosphere, but they must not add details that change the factual story. That raises two practical questions. How faithfully does an image model follow a long, specific prompt? And can another model check the result, or does a person have to look?

I ran a small benchmark to find out. The initial API calls cost $7.66.

The setup

Prompts. 15 prompts, each written in labelled sections (Scene / Subject / Details / Constraints), across three tests:

  • How much it keeps. A watermill yard with 8, 15 or 25 listed elements, such as “three flour sacks stacked beside the door”, “two white geese beside the well” and “a grey heron standing in the shallows”.
  • Specific details against the stereotype. Six scenes where the cliché is wrong: a 10th-century Italian market square (timber houses, a squat tower, no dome), a Bronze Age harbour with ox-hide-shaped copper ingots, an 1850s workshop lit by gas, a 1930s lab, a honey-fungus cluster on beech, a cumulonimbus over a flat plain.
  • Saying “no”. Three subjects with a strong default (domes on a hill town, red caps on mushrooms, a flywheel on a steam engine), each written twice: once with an explicit “No X”, once describing only what should be there.

Models. gpt-image-2.5 in two variants, sunburst and flare, 10 images per prompt each: 300 images.

Checklist. Every prompt has yes/no items (“Is any dome visible?”, “three flour sacks stacked beside the door”). An item passes only if it matches as written, including any count, arrangement or position it states. On top of that, one open question per image: did the model draw anything nobody asked for?

Readers. Three judge models answered the checklist for every image: gpt-6.1-sol, gpt-6-sol and gpt-6-luna. Claude produced an independent reference reading. Wherever any judge disagreed with the reference, I settled the question myself, in a small app that shows one large image and one question at a time without showing anyone’s answers. Then I ran a second round on every question where I had disagreed with the reference. For answers I did not adjudicate, I kept Claude’s reference reading. The scores below are therefore measured against a model reference corrected by human review, not against a dataset labelled by people from scratch.

A baseline. For detecting additions I also report a reader that never opens an image and always answers “yes”. A high accuracy score is not impressive if it cannot beat that reader.

What gpt-image-2.5 does

Image generation results
Measuresunburstflare
Listed elements drawn as written, 8 elements88.8%95.0%
Listed elements drawn as written, 15 elements94.0%96.0%
Listed elements drawn as written, 25 elements94.0%96.0%
Specific details drawn as written97.6%98.6%
Replaced a described scene with a stereotype or landmark0 of 600 of 60
Drew a forbidden thing0 of 500 of 50
Images with something nobody asked for (period scenes)54 of 6045 of 60
Unrequested things per image, average by test*3.2–5.32.0–4.2

*From the reference reading’s lists, which miss some small additions; the range runs from the lowest to the highest test average.

It keeps nearly all the listed details. Even at 25 listed elements, 94–96% of the items pass. The score does not fall as the list grows; the 8-element prompt scores lowest because one hard item (three sacks, stacked) weighs more in a short list. The misses are about how many and where: four or six sacks instead of three, three sacks side by side instead of stacked, a heron on a rock instead of in the water, a ladder against the wrong wall.

Storybook watermill yard with a water wheel, miller, donkey cart, well, geese, heron, scarecrow and a boy fishing
25 of 25. Every listed element is there. (flare)
Watermill with a miller at a red door and four flour sacks, one lying on three
Four sacks, not three. Asked for “three flour sacks stacked”. The count is the usual slip. (sunburst)

It respects “no”, here. No dome, red cap or flywheel appeared in any of the 100 images where the question was clear-cut. That held with “No X” in the prompt and without it, so the test shows that naming the forbidden thing did not backfire, not that it helped.

It follows specific details, but some unfamiliar ones still go wrong. It never swapped a described scene for a famous one. The ox-hide ingots were the weak spot: some images got the shape right but rendered the metal grey rather than copper-coloured.

Bronze Age quay with sewn plank boats and a stack of copper-coloured ingots
Copper ingots, on the quay. (flare)
Bronze Age quay with boats and a stack of grey ingots with flared corners
Right shape, wrong colour. Grey, not copper. (sunburst)

It fills gaps with its own assumptions. Almost every image contains things nobody asked for: hills, acorns, a railing, a second lantern. Most are harmless. Some can introduce anachronisms: one prompt asked for “a walled hill town seen from the valley road” and said nothing about the road. It became asphalt. If historical accuracy matters, check the background too. In this test, the road mattered as much as the buildings.

Walled hill town in a mountain valley with an asphalt road in the foreground
Nobody described the road. The model paved it. (flare)
Brown boletes in oak leaf litter with acorns at the lower left
Acorns nobody asked for. Plausible, but invented. (flare)

In this run, flare was slightly better at the same price. It drew more of what was asked and invented less, at the same cost per image. With 10 images per prompt that is a direction, not a proven difference.

Which model can check the images

Scores are measured against the corrected reference described above, over all 300 images; two ambiguous checklist items are excluded (see below).

Checklist items (positive = “it is visible”; 1,758 answers)

Judge checklist results
JudgePrecisionRecallSpecificity
gpt-6.1-sol98.7%99.1%93.3%
gpt-6-sol98.0%98.6%89.8%
gpt-6-luna96.6%99.5%81.7%

Most checklist items really are in the picture, so precision is high for everyone, including a reader who always says yes. Specificity matters here: when something is missing, does the judge notice? gpt-6-luna misses one absence in five.

Additions. Here I score one binary question per image: does it contain at least one unrequested addition? I do not score whether the judge found every addition, or whether each one it named was correct.

Judge addition detection results
JudgeAccuracyPrecisionRecall
gpt-6.1-sol98.3%100%98.2%
gpt-6-sol84.7%100%83.5%
gpt-6-luna60.7%100%57.7%
Always answers “yes”93.0%93.0%100%

None of the judges marked an image as containing additions when the reference marked it as containing none. But gpt-6-luna misses the presence of additions in about four out of ten images that contain them, and gpt-6-sol in about one in six. Since 93% of the images contain an addition, a reader who always answers “yes” is right 93.0% of the time, which is better than two of the three judges. Only gpt-6.1-sol beats it. Would more thinking help? I re-ran gpt-6.1-sol at medium instead of low effort over the same 300 images. On the checklist it changed nothing: it noticed a missing item 93.3% of the time either way. On additions it was slightly better, 99.0% against 98.3%: it named small details, such as grassy road verges, that the reference reading had missed.

I also looked at the 273 disputed questions I settled personally. They are a selection of hard cases, not a representative sample. On them Claude's reference reading agreed with my final answer 87.2% of the time and the best judge 84.6%. Claude read the checklist most literally (counts, “stacked”, “in the shallows”), and that is how I ended up reading it too. On additions, though, gpt-6.1-sol found more than Claude did. My first version of this comparison favoured Claude, because my second round only re-checked questions where I had disagreed with it. A later round re-checked the other side too, and I changed 6 of those 35 answers; the numbers here include it.

The human is not a gold standard either

In the second round I changed 30 of the 57 answers I was asked about again. These were disputed questions, so this is not a general human error rate. Some were misclicks. Most were questions that can be read two ways:

  • “A low, squat bell tower”: low compared to what? I flipped 14 of 16 answers.
  • “A flywheel”: is a solid crank disc a flywheel? I flipped all 5.
  • “Three flour sacks stacked beside the door”: I kept counting the sacks and forgetting “stacked”. In the end I asked the two questions separately: how many sacks, and does one lie on another.
  • Things so small you cannot tell what they are: a speck on the horizon, a dark shape in a window.

I excluded the tower and the crank-disc question from the scores. The lesson is about writing checklists. If the person settling the answer cannot decide, the judge cannot be scored on it. Spell out counts, arrangements and positions (“exactly three, one lying on the others”), and avoid relative words.

How this relates to public benchmarks

Public benchmarks already test whether an image model follows its prompt. GenEval (NeurIPS 2023) checks counts, colours and positions with an object detector on short prompts. DPG-Bench (2024) uses about a thousand long, dense prompts, 67 words on average, scored by a visual question-answering model. TIFA (ICCV 2023) and Davidsonian Scene Graph (ICLR 2024) turn a prompt into yes/no questions and let a model answer them, which is the checklist idea used here. So why measure again? First, none of them had results for the models I actually use. Second, and more important, their published scores rely on an automatic scorer, a detector or a question-answering model. Their authors also study scorer reliability; I wanted to check it for the particular models and prompts I use. Whether such a scorer can be trusted was one of my two questions, so here every judge is itself checked against answers a person settled, with an always-yes baseline beside it. The test is also closer to my use: prompts of up to 25 listed elements, in period scenes.

What I take from it

  1. Describe everything that matters. The model draws what you describe and fills the rest with its own assumptions.
  2. State counts and arrangements explicitly, and check them. That is where the model slips.
  3. “No X” did no harm here. Don’t drop it out of habit.
  4. For detecting whether an image contains additions, gpt-6.1-sol was the strongest of the three judges I tested.
  5. Put an always-yes baseline next to a judge’s accuracy. Without it, a judge scoring 84.7% can look fine while doing worse than one that never opens the image.

Limits. 15 prompts, one illustration style, 10 images per prompt and variant, one person writing the checklist and settling disagreements, and a shared model reference behind every score. A 98.3% here describes this sample, not a guarantee elsewhere.

Method notes

  • Generation: gpt-image-2.5 sunburst and flare, quality “low”, portrait 1088×1936 px, one image per call.
  • Judges: gpt-6.1-sol, gpt-6-sol and gpt-6-luna at reasoning effort “low”, each shown one image, the full prompt and the checklist, with this instruction: “You examine ONE generated image against a checklist. Answer only from what is visible in the pixels. For every checklist item return present=true only if the thing is clearly visible, otherwise false; never answer from the prompt’s wording. Then list in inventions every MATERIAL thing drawn (an object, person, animal, building, scenery feature or inscription) that the commissioned prompt did not ask for; if there is none return an empty list. Do not list style, lighting or the arrangement of asked-for things.”
  • Reference reading: Claude Opus 5.5 at medium effort, answering the same checklist and listing additions, one image at a time.
  • Cost: $7.66 of API calls: $1.71 for the 300 images and $5.95 for the judges (including $1.33 for one judge run that crashed and was repeated). The reference reading, follow-up reruns and my own time are not included.
  • One full prompt (the 8-element test):
Scene: A rural watermill yard on a river bank, late morning, in the style of a flat-colour storybook illustration. Subject: The mill yard with the mill building as the focal point; every required element below is drawn once, clearly visible and separate from the others. Details: 1. a timber water wheel on the mill’s left wall, turning in a stone race 2. a stone mill building with a slate roof 3. a red wooden door with black iron hinges 4. three flour sacks stacked beside the door 5. a donkey harnessed to a two-wheeled cart 6. a miller in a white floured apron standing at the door 7. a wooden footbridge crossing the river 8. a grey heron standing in the shallows Constraints: - Draw each required element exactly once. - No text, no lettering, no captions.

Data and code: every prompt and checklist, every judge’s answer, the reference reading, every question I settled and all 300 images are on GitHub at github.com/bigr/gpt-image-2.5-benchmark. One command recomputes every table in this article, without API keys.