GPT-6 Luna
Use GPT-6 Luna for focused, high-volume tasks. Check the model details, estimate your workload and try an example.
Model facts and API rates
gpt-6-luna · Last verified: · Official model reference
| Text tokens | USD per million (Standard) |
|---|---|
| Ordinary input | $0.1 |
| Cache reads | $0.01 |
| Cache writes | $0.125 |
| Billable output | $0.5 |
Context: 1,050,000 tokens · Maximum output: 128,000 tokens.
What GPT-6 Luna is for
GPT-6 Luna is OpenAI’s GPT-6 model for focused, high-volume tasks, released September 22, 2026. Its API ID is gpt-6-luna. Start with the published limits above, then try a small task that represents the work you want to repeat.
Give Luna a fixed output schema, allowed labels and a short set of examples. Keep each input record identifiable. Batch similar tasks, validate every response and send ambiguous cases to a review queue instead of silently forcing a label.
How to access GPT-6 Luna
Sol and Luna are available in Work and Codex for Plus, Pro, Business, Enterprise and Edu. Luna also has a Free/Go desktop route. The launch announcement does not offer Sol or Luna in ordinary Chat. The access guide has product, plan and client checks. For the API, use gpt-6-luna in a project with model access and billing enabled. The API setup guide provides a Responses request you can adapt.
Price your workload
The rates above come from the shared, dated model configuration. Calculate your actual usage with ordinary input, cache reads, writes and billable output. Reasoning tokens count as output. Requests above 272,000 input tokens use higher rates for the entire request, so check the total before adding a large document collection.
For recurring work, include the number of requests and expected retries. The prompt caching calculator helps estimate repeated prefixes. Cache reads and writes are categories within total input, not extra tokens to add again.
Tools, context and output
Use Responses for built-in tools and function calling. Text output, image input and structured outputs are supported. Native audio/video and fine-tuning are not supported by this model. The context window includes both input and output; reserve space for the answer rather than filling the whole window with source material.
Sol and Luna support Chat Completions function calling only with reasoning effort none. Reasoning effort and the Standard/Batch/Flex/Fast processing choice are different controls. EU data residency is available with Standard processing.
Try a task and inspect the output
Follow the practical guide, then inspect a worked example. Each example keeps its original model and run date. The prompt collection supplies reusable instructions and filled examples; skills package repeatable checks.
Compare with the rest of GPT-6
Open the Astra/Sol/Luna comparison to compare one workload across the family. It uses published rates and capabilities, without assigning a synthetic quality or speed score. For a migration from GPT-5, use the separate GPT-6 vs GPT-5 guide.
The release timeline explains when each family member launched. Related news records subsequent announcements; all GPT-6 models brings the current routes together.