GPT Image 2.5 quality settings: When should drafts and final assets use different configurations?

A concept draft and a delivery asset answer different questions. The draft asks whether the subject, framing, and visual direction work. The delivery asset also has to survive inspection at its intended display size. For GPT Image 2.5, that is a reason to make quality an explicit production-stage choice—not to equate the largest setting with approval.

The OpenAI image-generation guide, captured September 11, 2026, documents two identifiers: gpt-image-2.5-sunburst and gpt-image-2.5-flare. “GPT Image 2.5” below refers to that family, not a verified API alias named gpt-image-2.5. Keep the exact identifier in your configuration; the shared controls do not establish equivalent results.

The final-quality menu no longer stops at high

For both documented models, the quality options are low, medium, high, xhigh, max, and auto. Both default to auto, which chooses based on the prompt. The guide explicitly distinguishes the added xhigh and max options from earlier GPT Image models whose quality settings stop at high.

That changes the evaluation menu: a final-asset policy inherited from an earlier model may never compare the two additional options. But adding them to a comparison is not the same as making max the default for every final image.

The guide recommends low for quick drafts, then comparing higher settings for final assets to balance detail, latency, and cost. It does not establish a single best final setting or quantified savings from this workflow.

Stage Proposed quality policy Decision to make
Concept exploration Explicit low Is the visual direction worth developing?
Final-configuration evaluation Compare medium, high, xhigh, and max Which configuration meets the brief within measured operating constraints?
Delivery Use the selected explicit setting, then inspect the output Does this particular asset pass acceptance?

auto remains a documented option. If the goal is to compare named quality settings or maintain a stage-specific configuration, however, prompt-dependent selection is not the same policy as explicitly choosing one.

Compare configurations, not just attractive samples

Here is a proposed evaluation procedure, not a report of generated-image testing:

  1. Write the acceptance criteria before comparing settings. For a product scene, these might include a legible label, a clean silhouette, plausible material detail, and enough clear space for a later layout. Treat these as requirements to inspect, not promised model capabilities.
  2. Choose one exact model identifier. Start concept exploration at low; settle the direction and prompt before the final-quality comparison. If both Sunburst and Flare are candidates, run separate comparisons rather than merging their results.
  3. Hold the other controls fixed. Compare medium, high, xhigh, and max with the same prompt, dimensions, output format, background, and applicable compression setting. Quality is a rendering control; it is not a substitute for size or delivery-format decisions.
  4. Record more than the chosen image. For repeated outputs at each setting, retain the requested configuration, elapsed time, response usage, and pass/fail notes against the visual brief. Keeping inputs fixed makes the comparison more interpretable; it is not a guarantee of identical compositions.
  5. Choose against the requirements. If more than one configuration passes, use measured latency and usage to inform the tradeoff. If none passes, revisit the prompt, layout, or editing plan rather than treating another quality increase as an automatic repair.

The guide’s cost discussion is important here: the two models share token rates, but token consumption can differ by model and quality setting. Equal rates do not establish equal cost per image. It recommends explicit quality and size for estimation and the response’s usage for measurement. Older-model pricing tables are not a shortcut to GPT Image 2.5 quality-by-quality costs.

Keep visual approval separate from quality selection

The guide still warns about precise text placement and clarity, consistency of recurring characters or brand elements, and exact placement in layout-sensitive compositions. A higher quality label is not evidence that those issues have disappeared.

For a final asset, inspect lettering at the intended display size, compare recurring visual elements with approved references, and check the actual placement of subjects and negative space. Those checks should remain even after a configuration has performed well in evaluation.

The useful split is therefore low for deciding what to pursue, an explicit comparison including xhigh and max for deciding how to render, and visual review for deciding what to deliver. Which acceptance failure—text, consistency, composition, or material detail—would rule out an otherwise attractive result in your workflow?

I would separate the prompts used to choose the final configuration from the prompts used to check that choice. Repeating one settled prompt can tell you about that brief, but it leaves open whether the selected setting works across the kinds of assets you actually need.

The guide’s quality section gives this a concrete migration boundary: gpt-image-2.5-sunburst and gpt-image-2.5-flare add xhigh and max, while earlier GPT Image models stop at high. An evaluation inherited from those models therefore needs additional comparison groups, not just a renamed model column.

My proposed data-quality check would be:

  • Freeze a small evaluation set grouped by acceptance requirement—label-heavy images, recurring visual elements, and layout-sensitive scenes, for example. Keep some briefs out of prompt tuning and configuration selection.
  • For each exact model separately, compare the selected final setting with an explicit high baseline on those held-out briefs. If the selected setting is already high, compare it with the strongest alternative from the initial medium/high/xhigh/max comparison. “Strongest” here means against the written acceptance criteria, not the largest label.
  • Hide quality labels during visual scoring, and report accepted outputs divided by evaluated outputs for each brief group, with the counts visible. Retain rejected outputs too; a gallery of winners cannot show the failure rate.

This is a proposed validation method, not a measured result. It would support a narrower decision than “max is better”: whether an additional quality setting earns a place in a particular final-asset preset. If the held-out results disagree with the tuning results, I would keep that preset provisional rather than generalize from the attractive samples.