Balances useful reasoning quality with a lower input price for frequent use.
Back to models
Efficient reasoning Multimodal coding Agent-ready
OpenAI
GPT 5.6 Luna
Efficient GPT reasoning85% offInputImageTextOutputText
API formatOpenAI
An efficient GPT 5.6 tier for capable reasoning, coding, and multimodal assistant workflows.
gpt-5.6-lunaInput price$0.1467$1.00per million tokens
Output price$0.8802$6.00per million tokens
Context window400K tokensMaximum context per request
Max output128K
BillingPer token
HealthOperationalSuccess 100% · latency 3.3s · throughput 22.6
Model details
Capabilities, benchmarks, and typical use
GPT 5.6 Luna is the efficiency-oriented tier in the local Codex model family. The checked new-api configuration defines a model ratio of 0.5, a six-times completion ratio, and cache-creation ratio of 1.25; these resolve to the original prices shown here at the adapter's USD-per-million-token baseline. It accepts text and image inputs and produces text.
Combines code and image understanding for practical development workflows.
Supports tool calling, structured output, and context caching.

