The Rundown AI homepage

Independent tool overview

GPT-6 Luna: low-cost GPT-6 for focused work at scale

GPT-6 Luna is OpenAI's most efficient model for focused, high-volume tasks, pairing GPT-6 reasoning and tools with a $0.10/$0.50 API rate.

Visit the official GPT-6 Luna site ↗
GPT-6 Luna product preview
Best for
Focused, high-volume tasks
Standard API price
$0.10 input / $0.50 output per 1M tokens
Context and output
1.05M context / 128K maximum output
API model ID
gpt-6-luna
Default reasoning
Medium

Overview

What GPT-6 Luna is

GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks. OpenAI introduced Luna alongside GPT-6 Sol as a faster and more affordable way to bring advances from GPT-6 Astra into production workloads that need to operate at scale.

Despite its budget positioning, Luna keeps the 1,050,000-token context window, 128,000-token maximum output, selectable reasoning effort and broad Responses API tool support documented for the new family. That makes it suitable for more than simple text completion, provided the workload is narrow enough for its capability tier.

The API model ID is gpt-6-luna. Its standard short-context price is $0.10 per million input tokens and $0.50 per million output tokens, with discounted cached input, Batch and Flex options for teams optimizing throughput and unit economics.

Use cases

Who GPT-6 Luna is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Classification and routing

High-volume decisions such as categorization, intent routing, moderation support and structured tagging where per-request cost is critical.

Extraction and transformation

Turning documents, messages or images into structured fields, summaries and normalized records at scale.

Background automation

Scheduled and asynchronous tasks that benefit from tool access but do not require a flagship model for every step.

Cost-sensitive agent steps

Focused subtasks inside larger agent systems, with harder planning or coding stages routed to Sol or Astra only when needed.

Capabilities

Core GPT-6 Luna features

1

Extremely low token pricing

Standard short-context pricing starts at $0.10 per million input tokens and $0.50 per million output tokens.

2

Six reasoning-effort levels

Supports none, low, medium, high, xhigh and max reasoning effort, with medium documented as the default.

3

1.05M-token context window

Provides the same documented 1,050,000-token context capacity and 128,000-token maximum output as GPT-6 Sol.

4

Text and image understanding

Accepts text and image inputs while returning text, supporting focused visual extraction and multimodal routing tasks.

5

Broad Responses API tool support

Supports web and file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.

6

Structured production outputs

Supports streaming, function calling and structured outputs for dependable integration into application workflows.

7

Caching and processing discounts

Cached input is priced at 10% of normal input, while Batch and Flex processing are priced at half the Standard rate.

Process

How the GPT-6 Luna workflow works

  1. Step 1

    Select a focused task

    Start with repeatable work that has clear inputs, outputs and quality criteria rather than the most ambiguous reasoning problem in the system.

  2. Step 2

    Call gpt-6-luna through Responses

    Use the Responses API when the workflow needs built-in tools, function calling or reasoning controls.

  3. Step 3

    Use the lowest effective reasoning level

    Test none or low for simple routing and extraction, then raise effort only where measured quality improves enough to justify the added latency and tokens.

  4. Step 4

    Design for prompt caching

    Place stable instructions and reusable reference material first so repeated prefixes can benefit from the low cached-input rate.

  5. Step 5

    Escalate hard cases

    Route uncertain, high-impact or deeply agentic tasks to GPT-6 Sol or Astra instead of forcing Luna to handle every request.

Cost

GPT-6 Luna pricing and free plan

GPT-6 Luna is priced for high-volume API work. The standard short-context rate is $0.10 per million input tokens and $0.50 per million output tokens. Requests above 272K input tokens receive higher long-context rates for the full request.

Standard processing

$0.10 input / $0.50 output per 1M tokens

The default API processing tier for short-context requests.

  • $0.01 per 1M cached input tokens
  • $0.125 per 1M cache-write tokens
  • Long-context pricing applies when input exceeds 272K tokens

Batch and Flex

50% of Standard rates

The lowest-cost documented processing options for workloads that can accept batch or flexible scheduling.

  • Well suited to bulk classification and extraction
  • Long-context multipliers still apply
  • Latency and delivery behavior depend on the selected processing mode

Fast mode

2× the applicable rates

A higher-priced option for workloads that prioritize faster responses.

  • Standard short-context equivalent is $0.20 input / $1 output
  • Regional processing can add a 10% premium
  • EU data residency is available only with Standard processing

Pricing checked . Check current pricing at the source ↗

Assessment

GPT-6 Luna strengths and limitations

Where it stands out

  • Provides GPT-6-era reasoning and tool support at a very low per-token price.
  • Keeps a 1.05M-token context window and 128K maximum output despite its budget positioning.
  • Supports the same six reasoning-effort settings documented for Sol.
  • Works well as an inexpensive execution tier inside a routed multi-model system.

What to consider

  • OpenAI positions Luna for focused tasks, not as the default choice for the hardest coding, research or end-to-end agent work.
  • Inputs above 272K tokens cost more for the full request, even though the maximum context window is much larger.
  • Audio and video input are not supported by the documented model interface.
  • Fine-tuning is not supported.
  • Chat Completions function calling is documented only when reasoning effort is set to none; use the Responses API for the fuller tool experience.

Compare

GPT-6 Luna alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

GPT-6 Sol

Choose GPT-6 Sol when complex coding and agentic workflows need more capability than Luna's focused high-volume tier.

Explore GPT-6 Sol

Consumer

GPT-6 Astra

Choose GPT-6 Astra for the hardest end-to-end work when maximum capability matters more than Luna's unit economics.

Explore GPT-6 Astra

Consumer

Claude Fable 5.1

Compare Claude Fable 5.1 when evaluating another lower-cost general model for production writing, analysis and coding tasks.

Explore Claude Fable 5.1

Questions

GPT-6 Luna FAQs

What is GPT-6 Luna?

GPT-6 Luna is OpenAI's most efficient model for focused, high-volume tasks. It is the lowest-cost tier in the announced GPT-6 lineup.

How much does GPT-6 Luna cost?

Standard short-context API pricing is $0.10 per million input tokens, $0.01 per million cached input tokens, $0.125 per million cache-write tokens and $0.50 per million output tokens. Input above 272K tokens triggers higher long-context rates for the full request.

What is the GPT-6 Luna context window?

The documented context window is 1,050,000 tokens, with a maximum output of 128,000 tokens.

What reasoning settings does GPT-6 Luna support?

It supports none, low, medium, high, xhigh and max reasoning effort. Medium is the documented default.

Can GPT-6 Luna use tools?

Yes. Through the Responses API it supports web and file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search, as well as function calling.

Should I use GPT-6 Luna or GPT-6 Sol?

Use Luna for focused, repeatable and high-volume tasks where cost is the main constraint. Use Sol for more complex coding and agentic work that needs stronger capability.

What is the GPT-6 Luna API model ID?

Use gpt-6-luna in OpenAI API requests. The Responses API is the recommended path when you need built-in tools and reasoning controls.

Bottom line

Our GPT-6 Luna verdict

GPT-6 Luna is a compelling execution tier for large-scale classification, extraction, routing and focused automation. Its low price, large context window and complete tool surface make it unusually flexible for a budget model, but teams should route ambiguous, high-impact and deeply agentic tasks to Sol or Astra and verify quality with workload-specific evaluations.

Visit GPT-6 Luna website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.