OpenAI Puts Decisions API in Beta, Returning Probabilities and Scores From GPT-6 Luna at $0.10 per Million Input Tokens
OpenAI's Decisions API, in beta since October 6, returns probabilities, choices and scores from gpt-6-luna; OpenAI lists input at $0.10 per 1M tokens.
Editor's Note ·
- Clarification:
- The article says a "similar interface appeared the next day on a third-party platform" (Perplexity's Decider V1.1 27B). The only source for that paragraph is an OpenRouter search-results listing that shows October 7, 2026 as OpenRouter's listing date; it does not state when Perplexity released the model, and its "same API contract" wording refers to Perplexity's earlier Decider V1 27B, not to OpenAI's Decisions API. The similarity to OpenAI's interface and the one-day timing are the article's characterization, not statements by OpenRouter or Perplexity.
Overview
OpenAI released a Decisions API in beta on October 6, according to the OpenAI API changelog, which describes it as a way to “Turn text and images into typed answers 10x faster than the Responses API.” The speed figure is OpenAI’s own claim; the pages reviewed for this article do not give a baseline or methodology for it. Per the Decisions guide, gpt-6-luna is the only model currently available, and requests go to a dedicated POST /v1/decisions endpoint.
What the API returns
According to the guide, a request has three parts: a model, an input that can be a text string or user messages containing text and images, and a list of questions. Each question has a type that determines the answer:
- predicate returns a
probability, an estimate from 0 to 1 that a condition is true. - choice returns one of the values the developer supplied, along with per-option probabilities and a separate
confidencefield. - score rates the input against ordered levels and returns the probability-weighted average of the level indices, so the result can fall between levels. In the guide’s illustrative example, probabilities of 0.1, 0.7 and 0.2 across three levels produce a score of 1.1.
The guide lists classification, routing and prioritization as intended uses and says to use labeled examples from the application to set thresholds. It points developers to Structured Outputs with the Responses API when they need an object that follows their own JSON schema, such as extracted fields or a written explanation.
Several constraints appear in the same guide. Images must be inline base64 data URLs; hosted HTTP or HTTPS image URLs and file_id inputs are not supported by this endpoint. Independent questions can share one request, but for decisions that depend on an earlier answer the guide says to send separate requests. The minimum SDK versions it lists include Python 3.26.0 and JavaScript 7.30.0.
Pricing and availability
The guide states that with gpt-6-luna, input costs $0.10 per 1M tokens, and that there are no cache-read, cache-write or output-token charges. It adds that regional processing premiums and long-context input pricing multipliers apply. OpenAI’s GPT-6 Luna model page lists text-token prices of $0.1 per 1M input tokens and $0.5 per 1M output tokens for the model, so the Decisions rate matches the listed input price while dropping output charges. The model page also lists a 1,050,000-token context window and describes Luna as “our most efficient model for focused, high-volume tasks.”
The guide says the API supports Zero Data Retention and HIPAA use for eligible customers, with data residency and regional processing supported in the United States and Europe (EEA + Switzerland). On timing, it says: “The Decisions API is in public beta, and we expect to GA in the coming weeks.”
Early tooling and a similar interface elsewhere
In a post dated October 8, Vercel describes an experimental decision API in its AI SDK, which requires ai 7.0.128 or later, and shows GPT-6 Luna Decisions on AI Gateway under the model ID openai/gpt-6-luna-decisions. Vercel notes that “Gateway’s OpenAI-compatible Decisions endpoint currently accepts text only,” so image workflows require direct OpenAI access. The post presents its seven use cases, such as checking incident handoffs or screenshots, as designs to test, saying: “These are workflow designs to evaluate against your own examples.”
A similar interface appeared the next day on a third-party platform. A listing on OpenRouter dated October 7 describes Perplexity’s Decider V1.1 27B as a new checkpoint of Perplexity’s decision model, “succeeding Decider V1 27B with the same API contract.” According to that listing, it reads content passed as state and returns typed, probabilistic answers to named questions, handles up to 128 questions about the same content in one request, has a 262K-token context window, and is priced at $0.02 per million input tokens and $0 per million output tokens.
What We Don’t Know
- Accuracy and calibration. None of the pages reviewed report accuracy or calibration results for the returned probabilities, and the 10x speed comparison does not specify what workload it was measured on.
- General availability. OpenAI gives only “the coming weeks” for GA, and the pages reviewed do not say whether pricing or the model list will change at that point.
- Perplexity’s own documentation. The Decider details here come from OpenRouter’s listing; Perplexity’s own materials were not reviewed, and none of the sources links the two companies’ releases.