What is Cloudflare's Clef? A Summary of the Differences with Jev
This page has been translated by machine translation. View original
Hi, I'm Kema.
On October 1, 2026, a decision model called Clef was released on Cloudflare Workers AI.
It belongs to the same family as TypeSafe's Jev, and the APIs are compatible.
In this article, I've summarized an overview of Clef, its differences from Jev, and the results of sending the same questions to both models and comparing their responses.
1. What is Clef
Clef is the first model trained in-house by the Cloudflare Workers AI team, and on October 1, 2026, the following two models were released.
| Model | Positioning | Base Model | Price (per 1M input tokens) |
|---|---|---|---|
| Clef | Larger model prioritizing accuracy | Qwen3.8-27B | $0.24 |
| Clef-flash | Smaller model prioritizing speed | Qwen3.5-9B | $0.09 |
The prices are as listed on each model page (Models: clef | Cloudflare Workers AI docs, Models: clef-flash | Cloudflare Workers AI docs).
The weights for both are published on Hugging Face (Cloudflare/clef, Cloudflare/clef-flash) under the Apache 2.0 license, and you can also run them on your own GPU.
Like TypeSafe's Jev, Clef is a decision model that does not generate text but instead returns probabilities for choices per question.
It takes a state (state) and typed questions as input, and answers questions such as "Is this inquiry urgent?" or "Which team should handle this?" with probabilities.
There are three types of questions.
-
noul: A yes/no question. Returns the probability of yes. -
choice: Selects one from a defined set of choices. Returns the chosen option, the probability for each choice, and a confidence score. -
score: Evaluates using ordered levels. Returns a probability-weighted score and the probability for each level.
2. Differences from Jev
| Item | Clef | Jev |
|---|---|---|
| Provider | Cloudflare (Workers AI) | TypeSafe |
| Input | Text/JSON and images (up to 4 per request) | Text/JSON only |
| Number of questions (per request) | Up to 64 | No stated limit |
| Context | 64K tokens | 64K per request, 32K for state and longest question |
| Price (per 1M input tokens) | $0.24 (flash is $0.09) | $0.042 |
| Weights | Public (Apache 2.0) | Not public |
| API | System One compatible | System One |
The model page overview states that video is also supported, but the API parameters only include images, and there is no documentation on how to pass video.
The biggest difference is that images can be used as input.
Since Jev cannot accept images, handling document images requires first converting them to text via OCR, but Clef allows you to pass images directly as inputs for decision-making.
On the other hand, Jev is cheaper — Clef costs approximately 6 times as much as Jev.
Images are not passed as URLs but are embedded in the images field of the request.
In my testing, I confirmed that PNG images encoded as base64 and sent in the following data URL format worked successfully.
"images": ["data:image/png;base64,iVBORw0KGgo..."]
The model page lists the following limits.
| Item | Limit |
|---|---|
| Number of images | Up to 4 per request |
| Formats | PNG, JPEG, WebP |
| Size per image | Up to 4 MiB and 16 megapixels |
| Total images | Up to 8 MiB decoded |
| Entire request | Up to 13 MiB |
| Specifying by URL | Not allowed |
The original text on the model page explains that images are placed before the state and are subject to the above limits.
Optional embedded PNG, JPEG, or WebP images placed before the state (max 4; 4 MiB and 16 megapixels each, 8 MiB total decoded; whole request body max 13 MiB). Remote URLs are not accepted.
Source: Models: clef | Cloudflare Workers AI docs
The APIs are compatible, and when switching from Jev, only three things need to change: the endpoint, the API key, and model.
Let's send the same question to both Jev and Clef.
When sending to Jev:
curl https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": { "urgent": { "type": "noul", "instructions": "Is this support request urgent?" } }
}'
{"model":"jev-1.13.0","answers":{"urgent":{"type":"noul","noul":0.97}},"usage":{"input_tokens":284,"output_tokens":20}}
When sending to Clef ($CLOUDFLARE_AUTH_TOKEN is a Cloudflare API token with Workers AI permissions):
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": { "urgent": { "type": "noul", "instructions": "Is this support request urgent?" } }
}'
{"result":{"model":"clef","answers":{"urgent":{"type":"noul","noul":0.9917}},"usage":{"input_tokens":154,"output_tokens":0}},"success":true,"errors":[],"messages":[]}
The request content is completely identical except for the model specification, but the response structure differs.
While Jev returns answers directly at the root level, Clef called via the REST API wraps everything inside result and returns success and errors at the same level.
When reusing code written for Jev with the REST API, you will need to add processing to extract result before accessing answers.
3. Sending the Same Questions
I sent the same state and questions shown in the Clef model page sample to Jev (jev-1.13.0), clef, and clef-flash five times each.
The state is a support inquiry stating that "checkout has been failing for every customer for the last hour."
In response to this, three questions — one for each type — are sent together in a single request.
The request is as follows (Japanese comments are for explanation purposes and are not included in the actual JSON sent).
{
"model": "clef", // For Jev: "jev-latest", for clef-flash: "clef-flash"
// Input for decision-making (state): Checkout has been failing for every customer for the last hour.
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": {
"type": "noul", // yes/no question. Returns the probability of yes.
"instructions": "Is this support request urgent?" // Is this inquiry urgent?
},
"team": {
"type": "choice", // Select one from choices
"instructions": "Which team should handle this request?", // Which team should handle this?
"criteria": {
"billing": "Payments, invoices, and refunds", // Payments, billing, refunds
"technical": "Outages, errors, and configuration", // Outages, errors, configuration
"sales": "Plans and upgrades" // Plans, upgrades
}
},
"severity": {
"type": "score", // Evaluate using ordered levels
"instructions": "How severe is the customer impact?", // How severe is the impact on customers?
// 0: No impact, 1: Minor, 2: Major, 3: Critical
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}
| Item | Jev | clef | clef-flash |
|---|---|---|---|
| Urgent? (probability of yes) | 0.96–0.97 | 0.9906 | 0.9551 |
| Responsible team (probability of technical) | technical (0.99–1.0) | technical (0.8088) | technical (0.9355) |
| Severity (0–3) | 2.99 | 2.9573 | 2.7182 |
| Variation across 5 runs | Fluctuated at the 2nd decimal place | Identical across all 5 runs | Identical across all 5 runs |
| Input tokens | 402 | 346 | 346 |
| Execution time (median) | 0.238 seconds (238ms) | 0.478 seconds (478ms) | 0.294 seconds (294ms) |
| Cost (per request, at 150 JPY per USD) | $0.0000169 (approx. 0.0025 JPY) | $0.0000830 (approx. 0.0125 JPY) | $0.0000311 (approx. 0.0047 JPY) |
All three models produced the same decision results, but there were differences in the probability values and number of decimal places.
Jev returns probabilities to two decimal places, and 0.96 and 0.97 alternated across the five trials.
clef and clef-flash return probabilities to four decimal places, and the values were completely identical across all five runs.
In my testing environment, Jev had the shortest execution time.
However, since execution time varies depending on timing, network conditions, and server load, a simple comparison cannot be made from such a small number of trials. Execution times fell within a range of roughly 0.2 to 0.5 seconds, a level where the difference would be barely perceptible to humans.
Note that the model-side processing times (median) published by Cloudflare are: clef 209.3ms, clef-flash 38.8ms, and Jev 524.1ms (Changelog | Cloudflare Docs).
4. Summary
Clef is a decision model that can be used with the same API as Jev, with the main differences being that it accepts image input and its weights are publicly available.
For text-only decisions, the more affordable Jev is a good fit; for cases where you want to include images such as documents or screenshots as inputs, Clef is the better choice.
When using Clef, it's a good idea to start with clef-flash to check speed and accuracy, and then upgrade to clef if needed.