Creuto is now an OpenAI Select Partner Read More

AI & Machine Learning

Cloudflare Clef: what typed probabilities remove

Cloudflare Clef returns a probability per option in one forward pass. The specs disagree on context length — 64K or 16,384 — and that one bites.

Cloudflare Clef: what typed probabilities remove

Cloudflare Clef is an open-weight decision model: instead of generating text you then have to parse, it returns one probability for every allowed option of every typed question you ask, in a single forward pass. Cloudflare released it on 1 October 2026 under Apache-2.0. Before you design anything around it, check the context length — Cloudflare's own two sources give different numbers.

There are two models. Clef at 27 billion parameters and Clef-flash at 9 billion, served on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, with weights on Hugging Face. The model card states that Clef is post-trained from Qwen/Qwen3.8-27B; the launch post names Qwen 3.5-9B for the smaller one. Both are Apache 2.0 licensed, which means you can download them and run them yourself.

The Cloudflare Clef context length contradiction, and which number to budget against

Cloudflare's changelog and launch post both advertise a 64K token context window, and the Workers AI model page lists the context window as 65,536 tokens. The Hugging Face model card describes something different: max_length defaults to 16,384 tokens, configurable per call through encode_record, with a separate max_state_tokens parameter to bound the input state.

Both are primary sources and they are not saying the same thing. The most likely reading is that 64K is the architectural ceiling and 16,384 is the default the library applies when you do not override it — which is the worse of the two possibilities, because a ceiling you cannot reach is an annoyance while a default you did not notice is a silent bug. Nothing in either document says so explicitly, so treat this as our reading rather than Cloudflare's statement.

The failure mode is specific and quiet. A decision model has no prose output to inspect. If your record — the state plus the questions — runs past the limit in force, the truncated part simply stops influencing the answer, and you get a confident probability derived from an input that is missing its tail. There is no error, no refusal, no half-finished sentence to tip you off. Accuracy drops on exactly the long inputs you added the model to handle.

Three things to do before you size a prompt budget: set max_length explicitly rather than inheriting it, log the token count of every record you encode so truncation shows up as a metric rather than a mystery, and test one deliberately oversized record end to end to see what the model returns when the limit bites.

What the architecture actually does

The model card describes a backbone — Qwen3.8-27B with a vision encoder — plus a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly. The output is one logit per allowed option per question. That is the whole mechanism, and it explains both the speed and the constraints.

Because nothing is generated token by token, there is no decode loop. Cloudflare describes the inference path as a prefill-only pass that then scores the valid schema choices in parallel, which is why its published median latency is 209.3 ms for Clef and 38.8 ms for Clef-flash against 524.1 ms for Jev. Those figures are Cloudflare's own, measured by the company now selling against Jev, and no third party had reproduced them as of 2 October 2026.

Two architectural details matter more than the latency. First, scoring all options of all questions jointly means the questions in one request can share evidence: an answer about urgency can draw on the same hidden states as an answer about category, which is not true of independent classifier calls. Cloudflare documents 1 to 64 questions per request, with ids restricted to letters, digits, underscore, dot and hyphen up to 100 characters. Second, Cloudflare says it trained with label-smoothed cross-entropy and a Brier loss term specifically to calibrate the probabilities, which is the difference between a score you can threshold and a score that merely ranks.

Text, JSON, images and video in the same record

The input state is not limited to a prompt string. The model card lists text, JSON, images as PIL objects and video as frame arrays; the Workers AI page documents up to four images per request at 4 MiB each in PNG, JPEG or WebP. So a single request can carry a structured order record, the customer's free-text complaint and the photo they attached, and ask eight typed questions about all three at once. For anyone who has built a classification pipeline out of three separate services, that collapse is the interesting part of this release — more interesting than the benchmark table.

What typed probabilities remove from your code

The reason to care about any of this is the layer it deletes. A text-generating model asked to classify something returns a string, and around that string every team builds the same scaffolding. If you have shipped this, you will recognise the list.

What you build around a text modelWhat a decision model needs
JSON-schema prompt instructions and few-shot examplesThe question definition itself
A parser, plus a repair path for malformed JSONNothing — the response shape is fixed
Validation that the returned label is in your enumNothing — options are the only outputs scorable
A retry budget and backoff for when it is notNothing to retry
A self-reported confidence number of unknown meaningA calibrated probability per option
Prompt-injection surface in the instruction textA smaller one: options are not generated

The question types are the same three shapes the category settled on: noul for true/false, choice for named options with descriptions, and score for ordered ones. We wrote those up with examples in Choice, Score and Noul with examples, and the same reasoning about where this beats an LLM is in when a decision model beats an LLM in your software. The practical difference in a codebase is usually a few hundred lines of defensive parsing deleted, and a set of thresholds introduced in their place — which is work, not magic. Our notes on calling a decision model from a Node app and getting typed answers cover the shape of that integration.

The honest caveat: removing the parser does not remove the need to measure. A calibrated probability is still only calibrated on the distribution it was trained and evaluated on. Cloudflare's published CLINC150+OOS score of 97.43 for Clef sits beside 66.77 for Clef-flash on the same out-of-scope benchmark, which is a 30-point spread between two models from the same release on the task of knowing when none of your options apply. Pick the wrong one of the two and the layer you deleted comes back as an incident.

Can you fine-tune Clef on your own data?

Not yet, not on your own. Alongside the models Cloudflare announced a reinforcement-learning fine-tuning service, and it is worth being precise about its status: it starts as a hands-on engagement with its forward-deployed engineering team, with self-serve described as a later step and a design-partner signup in the meantime. Cloudflare names the pieces it assembles — AI Gateway to capture the data, Workers AI to run rollouts, Containers as the RL sandbox, and a new Trainer component — but publishes no price and no general-availability date as of 2 October 2026.

For most teams that makes fine-tuning a decision model a 2027 question rather than a 2026 one. The useful preparation is the same either way, and you can start it today: capture every decision your current model makes along with the outcome, because a labelled set of your own traffic is the input to any fine-tune, and it is also what tells you whether you need one. Weights under Apache-2.0 mean a self-hosted fine-tune is legally open to you regardless of what Cloudflare ships; the constraint is GPUs and labelled data, not licence. If you want to run the base model locally first, the patterns in running decision models on your own hardware still apply.

When a text model is still the right choice

A decision model is the wrong tool whenever the answer is not a choice from a set you can write down in advance. That rules out more than it sounds like.

  • Open-ended extraction. Pulling an address, a date range or a part number out of prose has no enumerable option list. Scoring options cannot invent a string.
  • Anything the user reads. Summaries, replies, explanations, draft copy. There is no text output here at all.
  • Taxonomies that change weekly. Every option must exist at request time; if your categories churn faster than your deploys, a text model with a prompt you can edit is more honest about that.
  • Chained reasoning you need to inspect. You get probabilities, not a rationale. When a human reviewer must see why, a model that explains itself is worth the parsing tax.
  • Very low volume. If you make a hundred decisions a day, the engineering to introduce a second model class costs more than the parsing it removes.

Where it does fit — routing, triage, moderation, eligibility, scoring, tool selection, anything where the set of valid answers is finite and the volume is real — it fits well enough that the parse-validate-retry layer starts looking like something you only built because the model could not answer the question directly. Whether to move an existing Jev integration onto Clef is a separate decision with its own costs, and we work through it in our companion comparison of Clef against Jev.

If you are deciding where a decision model belongs in a system you already run, that scoping is what our AI engineering practice does: find the decisions that are genuinely closed-set, measure the current error rate, and only then pick the model. The measurement is the part people skip, and it is the part that decides whether any of this pays.

Frequently asked questions

Cloudflare Clef is an open-weight decision model released on 1 October 2026 under Apache-2.0. Rather than generating text, Clef returns a calibrated probability for every allowed option of every typed question in a request, scoring all of them in a single forward pass on Workers AI or your own hardware.

The weights for both Clef and Clef-flash are published on Hugging Face under the Apache 2.0 licence, so you may self-host, modify and fine-tune them. The hosted Workers AI service is a separate commercial product priced per million input tokens, and the training data is not published.

Cloudflare's changelog and Workers AI model page give 64K tokens, roughly 65,536, while the Hugging Face model card states a default max_length of 16,384. Set max_length explicitly in your own code rather than inheriting the default, because truncation in a decision model produces no error.

The Hugging Face model card states that Clef is post-trained from Qwen/Qwen3.8-27B, with a vision encoder in the backbone and a transformer scoring head on top. Cloudflare's launch post names Qwen 3.5-9B as the base for the smaller Clef-flash model.

Not self-serve yet. Cloudflare's reinforcement-learning fine-tuning service starts as a hands-on engagement with its forward-deployed engineering team, with no published price or general-availability date as of 2 October 2026. Self-hosted fine-tuning is permitted by the Apache-2.0 licence if you have the GPUs.

Avoid a decision model when the answer is not a choice from a list you can write down in advance: open-ended extraction, any output a user reads, taxonomies that change faster than your deploys, or decisions where a human reviewer needs to see the reasoning rather than a probability.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

2 Oct 2026

·

9 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved