Decisions API Beta: GPT-6 Luna Answers and Pricing
The Decisions API moves into public beta with GPT-6 Luna, typed classification answers and dedicated input-only billing. The new endpoint is separate from ordinary Luna requests.
The Decisions API entered beta on October 6, giving developers a dedicated GPT-6 Luna endpoint for turning text and images into bounded answers. The change moves beyond the limited preview described at DevDay: applications can now evaluate conditions, choose categories and score inputs through a published request contract. OpenAI reports roughly ten-times-faster answers than the Responses API; that is its comparison, not a benchmark we have reproduced. The important implementation detail is that this endpoint has its own input-only billing, rather than the full pricing structure of an ordinary Luna response. OpenAI’s API changelog records the beta release.
What changed in the Decisions API beta
Our DevDay availability report recorded the earlier limited preview. The new beta adds the operational information needed to evaluate an integration: the exact endpoint, supported input forms, answer handling and endpoint-specific charges. The GPT-6 Luna model page describes the model; this update concerns the separate API surface.
The dedicated route is POST /v1/decisions, and the supported model ID is gpt-6-luna. A request combines a model, shared input and questions; returned answers follow question order. The API reference accepts a text string or user messages containing text and inline images. Image inputs require data URLs, with at most 128 images per request. Files, audio, external image URLs, tool calls and non-user roles are outside this contract. A question can return a refusal instead of a scored answer.
For an existing integration, first inspect what the application sends, not just what it expects back. A Responses workflow carrying file references or tool history cannot be moved by changing the endpoint name. Keep the original input adapter until a separate Decisions adapter has been checked against this narrower contract.
Choose typed answers, not a free-form response
The Decisions guide describes three question types. A predicate estimates whether a condition holds. A choice selects a supplied category. A score evaluates ordered levels; its value is a probability-weighted average of zero-based level indices, so it can fall between levels. These are different contracts: departments have no natural ordering, while severity levels do.
For illustration, consider a release-note router. This is our proposed integration exercise, not a tested OpenAI example. Supply one synthetic note announcing that an export format will be removed next month. Define a predicate asking whether the note contains an explicit removal announcement, a choice selecting the affected product surface, and a score describing the expected disruption to a saved workflow. Keep the categories separate from the disruption rubric.
The application needs a rule for what happens next. In this exercise, a removal flag would create a review item, not automatically edit a customer-facing guide. A product label would choose the reviewer queue. A disruption score would sort that queue without deciding whether the release should be blocked. Those three actions have different consequences and should remain visible in the integration design.
Write the synthetic note and expected classifications before making a request. Include a second note that announces an addition rather than a removal, and a third that contains only vague timing. These contrasts test whether the chosen contract distinguishes the cases the router actually cares about. They also expose an overbroad question before it becomes a production routing rule.
Where Structured Outputs still fits
Decisions API typed answers do not replace every JSON-producing workflow. OpenAI’s Structured Outputs documentation describes generating objects that conform to a supplied JSON schema, including extraction into named fields and programmatically detectable refusals. That remains a different requirement from selecting a fixed category.
In the release-note exercise, routing can be the first stage. A later stage may need to extract an effective date, list affected features and draft a short explanation for a reviewer. Keep that richer object-generation task separate. This makes the interface between stages explicit: a label routes the work; an extracted record supplies the evidence a person will inspect.
The existing GPT-6 Luna guide covers structured batch workflows, and its examples should not be relabeled as Decisions runs. Compare the new endpoint against the same saved inputs, but record the API surface independently. Otherwise a change in transport or output contract can be mistaken for a change in model quality.
Decisions API pricing is endpoint-specific
OpenAI’s endpoint pricing section lists $0.10 per million input tokens for this endpoint, with no cache-read, cache-write or output-token charges. Regional processing premiums and long-context input multipliers still apply. Other GPT-6 Luna requests retain their applicable model and processing-tier rates; this is not a blanket removal of Luna output charges.
At the base input rate, one million input tokens would cost $0.10 before applicable premiums or multipliers. This arithmetic is a unit-cost illustration, not a measured workload estimate. For a migration comparison, record actual input usage alongside the previous workflow’s complete bill. A shorter answer contract does not by itself establish a particular monthly saving.
Our pricing guide remains the general model-pricing reference. Its tables and the site’s calculators have not been changed by this News publication to model the Decisions endpoint. Use the dedicated official rate when evaluating this beta.
Check regional and data-control requirements
The data-controls documentation lists Decisions with GPT-6 Luna and regional processing in the United States and Europe, including the EEA and Switzerland. Regional storage and regional processing are separate capabilities. OpenAI also requires approval and additional conditions for special retention controls; an endpoint’s eligibility does not mean every account already has them enabled.
Before using sensitive records, confirm the project’s actual region and retention configuration with the person responsible for that account. Keep synthetic fixtures for the first integration check. A successful classification response proves neither that a regional configuration is correct nor that an organization has completed its required agreements.
For the first rollout, preserve the old router and compare the new answers without taking external actions. Track wrong destinations, refusals, missing answers, latency and cost on the same fixture. Review disagreements before enabling a write path. The beta’s value is a narrower, published decision interface; whether it improves a particular application remains a workload-specific test.