Classification and routing
High-volume decisions such as categorization, intent routing, moderation support and structured tagging where per-request cost is critical.
Independent tool overview
GPT-6 Luna is OpenAI's most efficient model for focused, high-volume tasks, pairing GPT-6 reasoning and tools with a $0.10/$0.50 API rate.
Visit the official GPT-6 Luna site ↗
Overview
GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks. OpenAI introduced Luna alongside GPT-6 Sol as a faster and more affordable way to bring advances from GPT-6 Astra into production workloads that need to operate at scale.
Despite its budget positioning, Luna keeps the 1,050,000-token context window, 128,000-token maximum output, selectable reasoning effort and broad Responses API tool support documented for the new family. That makes it suitable for more than simple text completion, provided the workload is narrow enough for its capability tier.
The API model ID is gpt-6-luna. Its standard short-context price is $0.10 per million input tokens and $0.50 per million output tokens, with discounted cached input, Batch and Flex options for teams optimizing throughput and unit economics.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
High-volume decisions such as categorization, intent routing, moderation support and structured tagging where per-request cost is critical.
Turning documents, messages or images into structured fields, summaries and normalized records at scale.
Scheduled and asynchronous tasks that benefit from tool access but do not require a flagship model for every step.
Focused subtasks inside larger agent systems, with harder planning or coding stages routed to Sol or Astra only when needed.
Capabilities
Standard short-context pricing starts at $0.10 per million input tokens and $0.50 per million output tokens.
Supports none, low, medium, high, xhigh and max reasoning effort, with medium documented as the default.
Provides the same documented 1,050,000-token context capacity and 128,000-token maximum output as GPT-6 Sol.
Accepts text and image inputs while returning text, supporting focused visual extraction and multimodal routing tasks.
Supports web and file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.
Supports streaming, function calling and structured outputs for dependable integration into application workflows.
Cached input is priced at 10% of normal input, while Batch and Flex processing are priced at half the Standard rate.
Process
Step 1
Start with repeatable work that has clear inputs, outputs and quality criteria rather than the most ambiguous reasoning problem in the system.
Step 2
Use the Responses API when the workflow needs built-in tools, function calling or reasoning controls.
Step 3
Test none or low for simple routing and extraction, then raise effort only where measured quality improves enough to justify the added latency and tokens.
Step 4
Place stable instructions and reusable reference material first so repeated prefixes can benefit from the low cached-input rate.
Step 5
Route uncertain, high-impact or deeply agentic tasks to GPT-6 Sol or Astra instead of forcing Luna to handle every request.
Cost
GPT-6 Luna is priced for high-volume API work. The standard short-context rate is $0.10 per million input tokens and $0.50 per million output tokens. Requests above 272K input tokens receive higher long-context rates for the full request.
$0.10 input / $0.50 output per 1M tokens
The default API processing tier for short-context requests.
50% of Standard rates
The lowest-cost documented processing options for workloads that can accept batch or flexible scheduling.
2× the applicable rates
A higher-priced option for workloads that prioritize faster responses.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Choose GPT-6 Sol when complex coding and agentic workflows need more capability than Luna's focused high-volume tier.
Explore GPT-6 Sol →Consumer
Choose GPT-6 Astra for the hardest end-to-end work when maximum capability matters more than Luna's unit economics.
Explore GPT-6 Astra →Consumer
Compare Claude Fable 5.1 when evaluating another lower-cost general model for production writing, analysis and coding tasks.
Explore Claude Fable 5.1 →Questions
GPT-6 Luna is OpenAI's most efficient model for focused, high-volume tasks. It is the lowest-cost tier in the announced GPT-6 lineup.
Standard short-context API pricing is $0.10 per million input tokens, $0.01 per million cached input tokens, $0.125 per million cache-write tokens and $0.50 per million output tokens. Input above 272K tokens triggers higher long-context rates for the full request.
The documented context window is 1,050,000 tokens, with a maximum output of 128,000 tokens.
It supports none, low, medium, high, xhigh and max reasoning effort. Medium is the documented default.
Yes. Through the Responses API it supports web and file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search, as well as function calling.
Use Luna for focused, repeatable and high-volume tasks where cost is the main constraint. Use Sol for more complex coding and agentic work that needs stronger capability.
Use gpt-6-luna in OpenAI API requests. The Responses API is the recommended path when you need built-in tools and reasoning controls.
Bottom line
GPT-6 Luna is a compelling execution tier for large-scale classification, extraction, routing and focused automation. Its low price, large context window and complete tool surface make it unusually flexible for a budget model, but teams should route ambiguous, high-impact and deeply agentic tasks to Sol or Astra and verify quality with workload-specific evaluations.
Visit GPT-6 Luna website ↗
DeepSeek's, small, efficient new open model

Claude Opus 5.5 - Anthropic’s flagship model for coding, research, writing, and complex agentic work

Google Labs' opt-in app mining your Gmail and Photos for ideas

Apps SDK - Chat with and build apps directly in ChatGPT

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.