GPT-6 Image Encoding Fix: Sol and Luna Retest
OpenAI's September 25 image encoding fix improves visual tasks with GPT-6 Sol and GPT-6 Luna in the API and Codex. Here is a focused retest method for affected workflows.
The GPT-6 image encoding fix gives teams a reason to revisit failed image workflows with GPT-6 Sol and GPT-6 Luna. OpenAI’s September 25, 2026 API changelog identifies a problem in image processing and reports better visual-task results across the API and Codex, including computer use. Its practical recommendation is to repeat affected evaluations. Start with your saved failures, not a different model or a redesigned prompt.
What the GPT-6 image encoding fix changes
The affected model IDs are gpt-6-sol and gpt-6-luna. The official API changelog dates the correction to September 25. This article was first published on October 2; the fix itself is not an October release.
The Sol model documentation and Luna model documentation list image inputs and text outputs. That makes screenshot interpretation and extracting information from a supplied picture relevant tests. Image understanding is a different task from generating a new image. The dated fix names Sol and Luna, not GPT-6.1 Sol; keep those model IDs separate in your results.
For the broader model profiles, use our GPT-6 Sol page and GPT-6 Luna page. The question here is narrower: does a previously unsuccessful visual task now meet your application’s acceptance criteria?
Build a small retest set from real failures
The following test design is our editorial recommendation, not an OpenAI benchmark. Select cases with a known correct answer and preserve the original files. A useful set might include a screenshot with a missed navigation label, a document image with an incorrect field, and a chart whose legend was confused. Include a case that already worked so an improvement in one category does not hide a regression elsewhere.
Keep the first pass controlled:
- Reuse the same image bytes, prompt, model ID, and available tools. Save the image dimensions and the request date alongside the result.
- Record the previous model settings, including reasoning effort and any image-detail setting your integration used. Preserve supported settings rather than importing a configuration from another model.
- Define success before looking at the new answer. For field extraction, that could mean matching three reference values with no invented fourth value.
- Compare the new output against the reference answer and the saved old output. Repeat cases where a single lucky answer would be misleading.
- Record whether the failure was image interpretation, missing input, tool execution, or an application assertion. These need different fixes.
If the original image or request is missing, use the new run to measure current suitability. It cannot establish a clean before-and-after comparison. Give such a case its own label instead of combining it with preserved regressions.
Retest screenshot reading separately from frontend behavior
Consider a search interface with a Clear button. A proposed visual check asks the model to identify the button, read the current query, and describe the visible result state. Define the expected labels and card count from the fixture rather than accepting a plausible description.
Then test the interaction independently: enter a query with no matches, activate Clear, and check that the full list returns. A model correctly reading an empty input is not evidence that the application’s filter state recovered. Browser assertions can establish that second claim.
Our existing frontend QA example illustrates exactly this distinction. Its deliberately broken training fixture clears the input without refreshing the cards. That historical example used Astra and does not measure this Sol/Luna fix; it supplies a concrete acceptance pattern for a new retest.
For a Sol or Luna screenshot check, save the actual screenshot and compare the model’s observations with the fixture’s known state. For the interaction check, save the browser assertion. Keeping those records separate shows whether the improvement occurred in reading the screen or whether the application itself still needs a repair.
Use document and chart cases with checkable answers
For a document image, choose a short, non-sensitive sample with three fields you can verify manually. Ask for the fields by name and specify how to represent an unreadable value. Score each field separately. A fluent paragraph can conceal a wrong identifier, while a field-level result makes the mismatch visible.
For a chart, ask a bounded question such as which labeled series has the highest value at one marked point. Record the reference answer and whether the label, series, and point were each identified correctly. Avoid treating an attractive summary as proof that the plot was read accurately.
Keep two passes distinct. First repeat the original input to assess the previously affected workflow. Then, if needed, try a clearer crop or a larger version and record that as an input improvement. Changing both the image and the model conditions at once makes it harder to identify why the result changed.
Check computer-use execution in a controlled environment
Because OpenAI includes computer use in the affected visual tasks, teams with screenshot-driven agents should revisit saved failures there too. Start with a local fixture or test account, a bounded task, and no purchase, message sending, or irreversible action.
A useful proposed task is to locate a labeled control, open a test panel, and stop when the expected heading appears. Evaluate three stages: whether the agent identified the control, whether the action reached it, and whether the resulting state matched the task. Save the screen before and after the action.
An accurate screen description can still be followed by a failed action. Conversely, a successful click does not prove that every detail of the screen was understood. Keep observation, execution, and completion as separate checks, especially when deciding whether to restore an automated workflow.
Interpret better results without changing the whole stack
OpenAI’s vision guide still identifies difficulties with tiny or rotated text, some chart styles, precise spatial tasks, and object counts. Its image-detail behavior is model-specific. Do not transfer Astra’s resizing rules to Sol or Luna simply because they belong to the same family.
Use retest results to make a concrete decision. Restore a workflow when its required checks pass; keep a failing category under review; adjust image preparation only in a separately recorded pass. Preserve the model ID, input, date, settings, and observed outcome so the next comparison starts from evidence. The value of this fix is the opportunity to recover useful visual workflows with a repeatable test, rather than permanently rejecting a model based on an affected run.