GPT-6.1 Sol: Price, Context Window and Access
OpenAI launched GPT-6.1 Sol for complex coding, computer use and professional work, with lower cached-input pricing, a 1.05M context window and access through ChatGPT Work, Codex and the API.
GPT-6.1 Sol is now available as a point upgrade to GPT-6 Sol, with the exact API model ID gpt-6.1-sol. OpenAI lists a 1,050,000-token context window, up to 128,000 output tokens and five reasoning efforts: low, medium, high, xhigh and max; medium is the default, while none and minimal are not supported. Standard API prices are $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens and $10 per million output tokens. The model is rolling out in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, but OpenAI says it is not yet available in the ordinary Chat surface. Read OpenAI’s GPT-6.1 Sol announcement.
This is a different event from the September 22 GPT-6 Sol and Luna launch. That release established the original family, prices and access surfaces. The September 29 update introduces a new model ID and capability release, cuts Sol’s cached-input price from the earlier $0.20 listing to $0.10, and publishes the cache-write rate, long-context multipliers and processing-mode adjustments needed for an API cost estimate.
GPT-6.1 Sol pricing and long-context multipliers
At Standard speed, one million uncached input tokens cost $2, one million cached input tokens cost $0.10, one million cache-write tokens cost $2.50, and one million output tokens cost $10. Cached reads are therefore 5% of the uncached input rate, while cache writes cost 1.25 times the uncached input rate. Those four lines matter separately: repeatedly writing a new cache does not produce the same bill as reusing an existing prefix.
Requests containing more than 272,000 input tokens use a higher rate for the entire request, not only the tokens above that threshold. OpenAI lists a 2x multiplier for input and cache charges and a 1.5x multiplier for output. A long request with 300,000 uncached input tokens and 20,000 output tokens would therefore use an effective input rate of $4 per million and an output rate of $15 per million before tool-call fees or other charges. Teams using the 1.05M context window should model that threshold explicitly rather than multiplying every request by the headline $2/$10 prices.
Processing mode adds another layer. Fast mode costs 2x Standard. Batch and Flex are 50% below Standard. Regional processing adds 10% where it is available, and the model page says Fast mode is unavailable with EU data residency. The pricing calculator can provide a workload baseline, while the prompt-caching calculator is the better starting point when requests reuse a stable prefix. Both estimates should be adjusted for the published long-context and processing-mode rules.
What the gpt-6.1-sol API supports
OpenAI recommends the Responses API for tool calling. Chat Completions is supported without tool calling, so a migration that depends on tools should not assume both endpoints expose the same behavior. The model page lists streaming, function calling and structured outputs as supported, while fine-tuning is not supported. It also lists web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search among the tools available through the Responses API.
The published context limit is 1,050,000 tokens and the maximum output is 128,000 tokens. Text input and output and image input are supported; audio and video are not. The knowledge cutoff is April 30, 2026. These are model specifications, not a recommendation to fill every request to the limit. Large contexts can cross the 272K pricing boundary, increase review time and retain irrelevant material. A representative test should compare a compact retrieval path with a full-context path on the same task.
The model’s reasoning-effort range also makes a single benchmark number an incomplete routing rule. Start with medium, measure whether high, xhigh or max improves the cases that actually fail, and reserve the added reasoning work for those cases. The existing GPT-6 Sol guide provides a baseline workflow; teams should record the new model ID and effort setting separately so results from Sol and 6.1 Sol are not mixed.
Access through ChatGPT Work, Codex and the API
OpenAI says GPT-6.1 Sol is available starting September 29 to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex. It is not yet available in Chat, even for those plans. That product boundary is easy to miss: seeing the model in a Work or Codex selector does not mean it should appear in a normal Chat conversation, and API access remains a separate route using gpt-6.1-sol.
OpenAI also says GPT-6.1 Sol Ultrafast is coming in the next few days, with up to 8x faster token generation than Standard speed in Codex. Ultrafast is therefore announced but not part of the launch-day availability claim. The statement describes generation speed, not an 8x improvement in task completion time, which can also depend on tool latency, prompt processing and review.
For rollout verification, record the product surface, account plan, displayed model name and test date. For the API, save the returned model ID, endpoint, reasoning effort, input size, cache status and processing mode. This separates a staged product rollout from an API configuration error and makes later cost comparisons reproducible.
How to evaluate the upgrade without overreading benchmarks
OpenAI positions GPT-6.1 Sol near Astra on selected coding, computer-use and professional-work evaluations at lower cost. It reports gains over GPT-6 Sol on DeepSWE, AutomationBench, OSWorld, Terminal-Bench Science and a difficult factuality set. These are vendor-reported results produced under named evaluation settings. OpenAI notes that research or API evaluations can differ from production ChatGPT because system prompts, tools and effort settings differ, and that its difficult factuality conversations are not representative of typical use.
The practical test is narrower. Choose a saved set of repository tasks, document questions or multi-step workflows; run GPT-6 Sol and GPT-6.1 Sol with the same inputs and review rubric; then compare successful completion cost, factual corrections, tool failures and reviewer time. Use the Sol model page as the prior-generation reference and the model comparison tool to keep routing criteria visible.
GPT-6.1 Sol’s launch value is not a universal claim that it replaces Astra or wins every task. It is a more specific operating option: the same $2 input and $10 output headline as Sol, half the cached-input rate, a published cache-write price, explicit long-context and processing multipliers, and a new model ID with broader stated performance. That is enough to justify a controlled comparison before changing production routing.