What is OpenAI's Decisions API? A Summary of the Differences with Jev and Clef

What is OpenAI's Decisions API? A Summary of the Differences with Jev and Clef

OpenAI released a new "Decisions API," so I organized the differences from existing Jev and Clef, and actually tried and compared them.
2026.10.07

This page has been translated by machine translation. View original

Hello, I'm Kema.

OpenAI has released a "Decisions API" that, rather than generating text, returns only the probability for each question.
It uses the same decision model as TypeSafe's Jev and Cloudflare's Clef, but there are some differences in the request and response formats.

In this article, I've summarized an overview of the Decisions API and its differences from Jev and Clef, then actually sent the same questions I tried in my previous Clef article to compare the responses.

1. What is the Decisions API

The Decisions API is an API that, when given input data (text or images) and typed questions, returns a probability for each question.
The endpoint is POST /v1/decisions, and the only available model at this time is gpt-6-luna.

Three question types are available:

  • predicate: A yes/no question. Returns the probability of yes (probability)

  • choice: A question that selects one option from defined choices. Returns the selected choice, probabilities for each choice, and a confidence score (confidence)

  • score: A question that evaluates on ordered levels. Returns a probability-weighted score, probabilities for each level, and a confidence score

Charges apply only to input tokens, at $0.10 per 1 million tokens.

Input costs $0.10 per 1M tokens. You pay only for input tokens: there are no cache-read, cache-write, or output-token charges.

Source: Decisions API | OpenAI API

2. Differences from Jev and Clef

2.1 Comparison Table

Item Decisions API Jev Clef
Provider OpenAI TypeSafe Cloudflare (Workers AI)
Input Text/JSON and images Text/JSON only Text/JSON and images
How input is passed input (a string, or messages combining input_text and input_image) state state and images
How questions are passed Array of questions. Each question has a name Dictionary of questions Dictionary of questions (keys are question names)
yes/no type predicate noul noul
How choices/levels are written choices (value and description), levels (label and description) criteria criteria
How answers are returned Array of answers (accessed by name) Dictionary of answers Dictionary of answers
Per-choice probabilities Array like [{"value": "technical", "probability": 1.0}] Dictionary like {"technical": 1.0} Dictionary like {"technical": 0.8088}
Level names label in each element of probabilities Separate legend Separate legend
Probability decimal places 2 decimal places 2 decimal places 4 decimal places
usage input_tokens, output_tokens, total_tokens with breakdown. (output_tokens is always 0) input_tokens, output_tokens. ("A value also appears in output_tokens") input_tokens, output_tokens. (output_tokens is always 0)
Price (per 1M input tokens) $0.10 $0.042 $0.24 (flash is $0.09)

Jev and Clef can be called with a common request format, but the Decisions API has a structure where questions are passed as an array and results are also returned as an array.
Therefore, if you reuse code written for Jev or Clef, you will need to rewrite the request construction and response parsing logic.

Also, images are not referenced by URL; instead, they must be converted to a base64 data URL and specified in input_image.
In that case, rather than a single string, you pass an array of messages to input, and submit the request alongside input_text (comments below are for explanation purposes).

"input": [{
  "role": "user",
  "content": [
    // Instruction: Inspect the product in this photo.
    {"type": "input_text", "text": "Inspect the product in this photo."},
    {"type": "input_image", "image_url": "data:image/png;base64,iVBORw0KGgo..."}
  ]
}]

2.2 How the Same Question Is Written and Returned

Here is the request for asking just one question — "Is this inquiry urgent?" — and the actual response that came back.

Jev/Clef (same except for model)

{
  "model": "jev-latest",
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "urgent": {"type": "noul", "instructions": "Is this support request urgent?"}
  }
}
{
  "model": "jev-1.13.0",
  "answers": {
    "urgent": {
      "type": "noul",
      "noul": 0.97
    }
  },
  "usage": {
    "input_tokens": 284,
    "output_tokens": 20
  }
}

Clef also returns answers in the same shape, but in the REST API the whole thing is wrapped in result.

Decisions API

{
  "model": "gpt-6-luna",
  "input": "Checkout has been failing for every customer for the last hour.",
  "questions": [
    {"type": "predicate", "name": "urgent", "instructions": "Is this support request urgent?"}
  ]
}
{
  "model": "gpt-6-luna",
  "answers": [
    {
      "type": "predicate",
      "name": "urgent",
      "probability": 0.99
    }
  ],
  "usage": {
    "input_tokens": 166,
    "input_tokens_details": {
      "cached_tokens": 0,
      "cache_write_tokens": 0
    },
    "output_tokens": 0,
    "output_tokens_details": {
      "reasoning_tokens": 0
    },
    "total_tokens": 166
  }
}

3. Trying the Official Guide Samples

First, I ran the choice and score samples from the official guide exactly as written, 5 times each.

The comments inside the requests are for explanation purposes and were not included in the actual requests.

3.1 choice (Responsible Department)

curl https://api.openai.com/v1/decisions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    // Input: I was charged twice for my order.
    "input": "I was charged twice for my order.",
    "questions": [{
      "type": "choice",
      "name": "department",
      // Which department should handle this complaint?
      "instructions": "Which department should handle this complaint?",
      "choices": [
        {"value": "billing", "description": "Payments, invoices, and refunds."},
        {"value": "technical", "description": "Problems using the product."},
        {"value": "shipping", "description": "Delivery and tracking."},
        {"value": "other", "description": "Requests outside these categories."}
      ]
    }]
  }'
{
  "model": "gpt-6-luna",
  "answers": [
    {
      "type": "choice",
      "name": "department",
      "choice": "billing",
      "probabilities": [
        {
          "value": "billing",
          "probability": 1.0
        },
        {
          "value": "technical",
          "probability": 0.0
        },
        {
          "value": "shipping",
          "probability": 0.0
        },
        {
          "value": "other",
          "probability": 0.0
        }
      ],
      "confidence": 1.0
    }
  ],
  "usage": {
    "input_tokens": 145,
    "input_tokens_details": {
      "cached_tokens": 0,
      "cache_write_tokens": 0
    },
    "output_tokens": 0,
    "output_tokens_details": {
      "reasoning_tokens": 0
    },
    "total_tokens": 145
  }
}

All 5 runs returned billing 1.0 with confidence 1.0 (median execution time 0.273 seconds (273ms), input tokens 145).
The official guide's example response shows billing 0.95 with confidence 0.93, but in practice it came back pegged at 1.0.

3.2 score (Severity)

curl https://api.openai.com/v1/decisions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    // Input: Export fails in Safari but works in Chrome.
    "input": "Export fails in Safari but works in Chrome.",
    "questions": [{
      "type": "score",
      "name": "severity",
      // How severe is this issue?
      "instructions": "How severe is this issue?",
      "levels": [
        {"label": "Cosmetic", "description": "Appearance only; no lost functionality."},
        {"label": "Workaround available", "description": "A task fails, but another way works."},
        {"label": "Fully blocked", "description": "A task fails with no workaround."}
      ]
    }]
  }'
{
  "model": "gpt-6-luna",
  "answers": [
    {
      "type": "score",
      "name": "severity",
      "score": 0.97,
      "probabilities": [
        {
          "value": 0,
          "label": "Cosmetic",
          "probability": 0.03
        },
        {
          "value": 1,
          "label": "Workaround available",
          "probability": 0.97
        },
        {
          "value": 2,
          "label": "Fully blocked",
          "probability": 0.0
        }
      ],
      "confidence": 0.96
    }
  ],
  "usage": {
    "input_tokens": 146,
    "input_tokens_details": {
      "cached_tokens": 0,
      "cache_write_tokens": 0
    },
    "output_tokens": 0,
    "output_tokens_details": {
      "reasoning_tokens": 0
    },
    "total_tokens": 146
  }
}

All 5 runs returned score 0.97 with confidence 0.96 (median execution time 0.256 seconds (256ms), input tokens 146).
The official guide's example response shows score 1.1 with confidence 0.55, but in practice 0.97 was concentrated on "Workaround available", skewing toward one side.

4. Sending the Same Questions as Jev and Clef

Next, I rewrote the state and question content from my Clef article into the Decisions API format and sent 5 requests in the same way.
The state assumes a support inquiry that "checkout has been failing for all customers for the past hour," and the 3 question types are combined into a single request.

This time I tried omitting description from the request, but it did not cause an error and a normal response was returned.

Item Jev clef clef-flash Decisions API
Urgent (probability of yes) 0.96–0.97 0.9906 0.9551 0.96
Responsible team (probability of technical) technical (0.99–1.0) technical (0.8088) technical (0.9355) technical (1.0)
Severity (0–3) 2.99 2.9573 2.7182 2.97
Variation across 5 runs Fluctuated at 2 decimal places Identical all 5 runs Identical all 5 runs Identical all 5 runs
Input tokens 402 346 346 406
Execution time (median) 0.238 seconds (238ms) 0.478 seconds (478ms) 0.294 seconds (294ms) 0.255 seconds (255ms)
Price (per request, at 150 yen per dollar) $0.0000169 (approx. 0.0025 yen) $0.0000830 (approx. 0.0125 yen) $0.0000311 (approx. 0.0047 yen) $0.0000406 (approx. 0.0061 yen)

※ The Jev and Clef figures are actual measurements taken when I wrote the Clef article (October 4, 2026).
The verdict conclusion was consistent across all 4 models, and the Decisions API returned probabilities to 2 decimal places, the same as Jev.

5. Summary

The Decisions API is a decision model that, like Clef, accepts not only text but also images as input, and in terms of pricing it is less expensive than clef and close to the level of clef-flash.
In my local runs, there was also no variation between trials, and it operated at a response speed comparable to Jev and Clef.

However, the request and response structure is not compatible with Jev or Clef.
Since questions and answers are exchanged in array format, migrating from existing code will require modifications to the construction and parsing logic.

If cost is the top priority for text-only decisions, Jev is a strong option; if you want to include images in the input while balancing cost and speed, the Decisions API looks like a compelling choice.

https://dev.classmethod.jp/articles/cloudflare-clef-overview/


AI白書2026 配布中

クラスメソッドが独自に行なったAI診断調査をもとに、企業のAI活用の現在地を調査レポートとしてまとめました。企業規模別の活用度傾向に加え、規模を超えてAI活用を進める企業に共通する取り組みまで、自社の現在地を捉えるためのヒントにぜひ。

AI白書2026

無料でダウンロードする

Share this article

DevelopersIO 2026