Creuto is now an OpenAI Select Partner Read More
Cloudflare Clef returns a probability per option in one forward pass. The specs disagree on context length — 64K or 16,384 — and that one bites.

Cloudflare Clef is an open-weight decision model: instead of generating text you then have to parse, it returns one probability for every allowed option of every typed question you ask, in a single forward pass. Cloudflare released it on 1 October 2026 under Apache-2.0. Before you design anything around it, check the context length — Cloudflare's own two sources give different numbers.
There are two models. Clef at 27 billion parameters and Clef-flash at 9 billion, served on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, with weights on Hugging Face. The model card states that Clef is post-trained from Qwen/Qwen3.8-27B; the launch post names Qwen 3.5-9B for the smaller one. Both are Apache 2.0 licensed, which means you can download them and run them yourself.
Cloudflare's changelog and launch post both advertise a 64K token context window, and the Workers AI model page lists the context window as 65,536 tokens. The Hugging Face model card describes something different: max_length defaults to 16,384 tokens, configurable per call through encode_record, with a separate max_state_tokens parameter to bound the input state.
Both are primary sources and they are not saying the same thing. The most likely reading is that 64K is the architectural ceiling and 16,384 is the default the library applies when you do not override it — which is the worse of the two possibilities, because a ceiling you cannot reach is an annoyance while a default you did not notice is a silent bug. Nothing in either document says so explicitly, so treat this as our reading rather than Cloudflare's statement.
The failure mode is specific and quiet. A decision model has no prose output to inspect. If your record — the state plus the questions — runs past the limit in force, the truncated part simply stops influencing the answer, and you get a confident probability derived from an input that is missing its tail. There is no error, no refusal, no half-finished sentence to tip you off. Accuracy drops on exactly the long inputs you added the model to handle.
Three things to do before you size a prompt budget: set max_length explicitly rather than inheriting it, log the token count of every record you encode so truncation shows up as a metric rather than a mystery, and test one deliberately oversized record end to end to see what the model returns when the limit bites.
The model card describes a backbone — Qwen3.8-27B with a vision encoder — plus a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly. The output is one logit per allowed option per question. That is the whole mechanism, and it explains both the speed and the constraints.
Because nothing is generated token by token, there is no decode loop. Cloudflare describes the inference path as a prefill-only pass that then scores the valid schema choices in parallel, which is why its published median latency is 209.3 ms for Clef and 38.8 ms for Clef-flash against 524.1 ms for Jev. Those figures are Cloudflare's own, measured by the company now selling against Jev, and no third party had reproduced them as of 2 October 2026.
Two architectural details matter more than the latency. First, scoring all options of all questions jointly means the questions in one request can share evidence: an answer about urgency can draw on the same hidden states as an answer about category, which is not true of independent classifier calls. Cloudflare documents 1 to 64 questions per request, with ids restricted to letters, digits, underscore, dot and hyphen up to 100 characters. Second, Cloudflare says it trained with label-smoothed cross-entropy and a Brier loss term specifically to calibrate the probabilities, which is the difference between a score you can threshold and a score that merely ranks.
The input state is not limited to a prompt string. The model card lists text, JSON, images as PIL objects and video as frame arrays; the Workers AI page documents up to four images per request at 4 MiB each in PNG, JPEG or WebP. So a single request can carry a structured order record, the customer's free-text complaint and the photo they attached, and ask eight typed questions about all three at once. For anyone who has built a classification pipeline out of three separate services, that collapse is the interesting part of this release — more interesting than the benchmark table.
The reason to care about any of this is the layer it deletes. A text-generating model asked to classify something returns a string, and around that string every team builds the same scaffolding. If you have shipped this, you will recognise the list.
| What you build around a text model | What a decision model needs |
|---|---|
| JSON-schema prompt instructions and few-shot examples | The question definition itself |
| A parser, plus a repair path for malformed JSON | Nothing — the response shape is fixed |
| Validation that the returned label is in your enum | Nothing — options are the only outputs scorable |
| A retry budget and backoff for when it is not | Nothing to retry |
| A self-reported confidence number of unknown meaning | A calibrated probability per option |
| Prompt-injection surface in the instruction text | A smaller one: options are not generated |
The question types are the same three shapes the category settled on: noul for true/false, choice for named options with descriptions, and score for ordered ones. We wrote those up with examples in Choice, Score and Noul with examples, and the same reasoning about where this beats an LLM is in when a decision model beats an LLM in your software. The practical difference in a codebase is usually a few hundred lines of defensive parsing deleted, and a set of thresholds introduced in their place — which is work, not magic. Our notes on calling a decision model from a Node app and getting typed answers cover the shape of that integration.
The honest caveat: removing the parser does not remove the need to measure. A calibrated probability is still only calibrated on the distribution it was trained and evaluated on. Cloudflare's published CLINC150+OOS score of 97.43 for Clef sits beside 66.77 for Clef-flash on the same out-of-scope benchmark, which is a 30-point spread between two models from the same release on the task of knowing when none of your options apply. Pick the wrong one of the two and the layer you deleted comes back as an incident.
Not yet, not on your own. Alongside the models Cloudflare announced a reinforcement-learning fine-tuning service, and it is worth being precise about its status: it starts as a hands-on engagement with its forward-deployed engineering team, with self-serve described as a later step and a design-partner signup in the meantime. Cloudflare names the pieces it assembles — AI Gateway to capture the data, Workers AI to run rollouts, Containers as the RL sandbox, and a new Trainer component — but publishes no price and no general-availability date as of 2 October 2026.
For most teams that makes fine-tuning a decision model a 2027 question rather than a 2026 one. The useful preparation is the same either way, and you can start it today: capture every decision your current model makes along with the outcome, because a labelled set of your own traffic is the input to any fine-tune, and it is also what tells you whether you need one. Weights under Apache-2.0 mean a self-hosted fine-tune is legally open to you regardless of what Cloudflare ships; the constraint is GPUs and labelled data, not licence. If you want to run the base model locally first, the patterns in running decision models on your own hardware still apply.
A decision model is the wrong tool whenever the answer is not a choice from a set you can write down in advance. That rules out more than it sounds like.
Where it does fit — routing, triage, moderation, eligibility, scoring, tool selection, anything where the set of valid answers is finite and the volume is real — it fits well enough that the parse-validate-retry layer starts looking like something you only built because the model could not answer the question directly. Whether to move an existing Jev integration onto Clef is a separate decision with its own costs, and we work through it in our companion comparison of Clef against Jev.
If you are deciding where a decision model belongs in a system you already run, that scoping is what our AI engineering practice does: find the decisions that are genuinely closed-set, measure the current error rate, and only then pick the model. The measurement is the part people skip, and it is the part that decides whether any of this pays.
Cloudflare Clef is an open-weight decision model released on 1 October 2026 under Apache-2.0. Rather than generating text, Clef returns a calibrated probability for every allowed option of every typed question in a request, scoring all of them in a single forward pass on Workers AI or your own hardware.
The weights for both Clef and Clef-flash are published on Hugging Face under the Apache 2.0 licence, so you may self-host, modify and fine-tune them. The hosted Workers AI service is a separate commercial product priced per million input tokens, and the training data is not published.
Cloudflare's changelog and Workers AI model page give 64K tokens, roughly 65,536, while the Hugging Face model card states a default max_length of 16,384. Set max_length explicitly in your own code rather than inheriting the default, because truncation in a decision model produces no error.
The Hugging Face model card states that Clef is post-trained from Qwen/Qwen3.8-27B, with a vision encoder in the backbone and a transformer scoring head on top. Cloudflare's launch post names Qwen 3.5-9B as the base for the smaller Clef-flash model.
Not self-serve yet. Cloudflare's reinforcement-learning fine-tuning service starts as a hands-on engagement with its forward-deployed engineering team, with no published price or general-availability date as of 2 October 2026. Self-hosted fine-tuning is permitted by the Apache-2.0 licence if you have the GPUs.
Avoid a decision model when the answer is not a choice from a list you can write down in advance: open-ended extraction, any output a user reads, taxonomies that change faster than your deploys, or decisions where a human reviewer needs to see the reasoning rather than a probability.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand