OpenAI Decisions API Beta: 10x Faster Typed Answers, $0.10 Input

October 9, 2026 • news
OpenAIModel ReleaseProduction AI

OpenAI launched the Decisions API in public beta on October 6, 2026, making gpt-6-luna available as a non-generative decision endpoint that returns typed probabilities, choices, and scores rather than prose. The API is hosted by OpenAI only, through POST /v1/decisions, with no open-weight or self-hosted variant. OpenAI expects general availability in the coming weeks.

A request has three fields: model, input, and questions. Input can be a text string or user messages containing input_text and inline base64 input_image parts. Each question has a unique name, a type, instructions, and optional choices or levels. The response returns an answers array keyed by question name, and refusals are possible.

OpenAI supports three question types:

  • Predicate: evaluates a boolean condition and returns a probability from 0 to 1, such as whether a product photo shows a crack, tear, or dent.
  • Choice: picks one value from developer-supplied options, returning the selected choice, per-option probabilities, and a confidence field.
  • Score: rates input against zero-indexed ordered levels and returns a probability-weighted average. OpenAI's severity example uses level probabilities of 0.1, 0.7, and 0.2 to produce a score of 1.1, plus an overall confidence metric.

SDK minimums are Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, and Java 4.78.0.

Pricing is input-only: $0.10 per million input tokens, with no charge for output tokens, cache reads, or cache writes. That undercuts the $0.20 input price for the previous-generation GPT-5.6 Luna in AI Mastery's pricing data, though long-context multipliers still apply. The GPT-6 Luna model card lists a 1,050,000-token context window and bills prompts above 272,000 tokens at 2x the base input rate. The model's parameter count is not disclosed.

For speed, OpenAI claims about 10x faster responses than the Responses API, but has not published independent accuracy or calibration data for the endpoint. DevDay coverage reported by MarkTechPost put a single decision near 150 milliseconds against roughly 1.6 seconds for a standard gpt-6-luna generation call.

Compliance coverage includes Zero Data Retention and HIPAA for eligible customers, with data residency in the United States and Europe (EEA plus Switzerland). Voice interactions can drive decisions through client delegation with the Live API.

OpenAI draws an explicit boundary around the endpoint. The Decisions API is for probabilities, choices, and scores only. Custom JSON schema fill or explanation generation belongs on Structured Outputs, and workloads that need tool calls with arguments should use function calling.

The closest competitor is TypeSafe Jev, a System One model in early access since September 15, 2026. TypeSafe returns typed values with calibrated probabilities across up to 255 choices. It prices input at $0.042 per million tokens with free output, making OpenAI's base rate about 2.4x higher, and reports 70–500 ms end-to-end latency from the US West Coast. OpenAI's advantages remain image input support, compliance options, and open beta access.

Feature OpenAI Decisions API TypeSafe Jev 1.13 GPT-6 Luna (Responses API)
Release Status Beta (Oct 6, 2026) Early Access (Sep 15, 2026) GA (Sep 22, 2026)
Underlying Model gpt-6-luna Jev (System One) gpt-6-luna
Input Support Text, inline base64 images Text and structured state Text, images
Output Format Probability, choice, or score Typed values with probabilities Generated text (JSON)
Input Price (per 1M tokens) $0.10 $0.042 $0.10
Output Price (per 1M tokens) $0.00 $0.00 $0.50

AI Mastery analysis

The Decisions API is the clearest evidence yet that Bounded Outputs Are Replacing Prose in Production AI. Developers have long forced generative models into classifier roles, then wrapped them in parsing logic to recover a boolean or category from a stream of text. Removing generation from the request path changes both cost and latency behavior at the API boundary.

The score type matters beyond convenience. A probability-weighted average such as 1.1 exposes uncertainty that a hard “Level 1” label would hide, letting applications route borderline cases to human review on explicit confidence thresholds. The trade-off is auditability: with no generated chain-of-thought, debugging a wrong decision means evaluating the model's calibration instead of reading its reasoning.

OpenAI's documentation advises developers to build calibration tests from their own labeled examples, since no endpoint accuracy data has been published. For now, the launch is a production routing primitive with an instruction manual, not an independently benchmarked model.

Sources

Frequently asked questions

How much does the OpenAI Decisions API cost per million tokens?

OpenAI prices the Decisions API at $0.10 per million input tokens, with no charge for output tokens, cache reads, or cache writes. Long-context multipliers still apply: GPT-6 Luna prompts above 272,000 tokens are billed at 2x the base input rate.

Does the OpenAI Decisions API run on self-hosted hardware?

No. It runs only on OpenAI's hosted infrastructure through POST /v1/decisions, with no open-weight or self-hosting option available. Compliance coverage includes Zero Data Retention and HIPAA for eligible customers.

What is the difference between the Decisions API and Structured Outputs?

Decisions returns only probabilities, choices, and scores. Structured Outputs is for filling custom JSON schemas or writing explanations, while function calling is used when a model must request a tool call with arguments.

How fast is the OpenAI Decisions API compared to the Responses API?

OpenAI claims the Decisions API is about 10x faster than the Responses API, though no independent benchmark has been published. MarkTechPost DevDay coverage put a single decision near 150 milliseconds versus about 1.6 seconds for a regular GPT-6 Luna generation call.

Free interactive tools for the decisions this piece raises.

Related Reading