Creuto is now an OpenAI Select Partner Read More
Clef vs Jev: Cloudflare's open 27B model answers in 209ms against Jev's 524ms — but it published the test. What migration really costs.

Clef vs Jev comes down to three things, and the benchmark table is not one of them: whether you can realistically run a 27-billion-parameter model yourself, whether the benchmarks Cloudflare chose resemble the decisions your product actually makes, and what switching costs you in recalibration rather than in code. Every head-to-head figure published so far comes from Cloudflare, which now sells a competing product.
Cloudflare released Clef and Clef-flash on 1 October 2026 under the Apache 2.0 licence, positioned explicitly against Jev. If you integrated Jev in the past year — and a lot of teams did, some of them on our recommendation — you now have a migration question. Here is what changed, and what it does not change.
Two models: Clef at 27B and Clef-flash at 9B, both open-weight, both served on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The model card states that Clef is post-trained from Qwen/Qwen3.8-27B, and Cloudflare's launch post names Qwen 3.5-9B as the smaller model's base. Weights are on Hugging Face as sharded BF16 safetensors, so you can download and run them.
The numbers below are the ones doing the persuading. They are all from Cloudflare's own launch post and changelog — there is no independent reproduction of any of them as of 2 October 2026, and the party that ran the test is the party that benefits from the result. Read the table as a vendor claim, not as a measurement.
| Cloudflare's published figure | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 95.75 |
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS, macro-F1 | 97.43 | 66.77 | 89.27 |
| Median latency | 209.3 ms | 38.8 ms | 524.1 ms |
| p95 latency | 238.6 ms | 122.4 ms | 536.0 ms |
Cloudflare's changelog claims its models score highest on 7 of 10 decision benchmarks, and that Clef-flash is 13x faster than Jev at median latency. Our own verification of the launch material also found Clef ahead of Jev in three of four workflow evaluations. Nothing in that set is implausible. All of it is self-reported.
Counting won rows is the wrong way to read a vendor benchmark sheet, because the vendor picked the rows. Three questions survive that objection.
Look at the CLINC150+OOS row again. Clef scores 97.43 and Jev 89.27, but Clef-flash scores 66.77 — far below the incumbent. That benchmark measures out-of-scope detection: knowing when the input matches none of the options you offered. If your product routes support tickets, parses whatever a user typed into a free-text box, or sits in front of a tool-calling agent, out-of-scope detection is most of your real failure budget, and the cheap, fast model is the one that fails it.
BANKING77 and BFCL point the other way. If your decisions are a closed set of well-separated intents, or structured function-call arguments, Clef-flash at 38.8 ms median looks like the interesting model and the 27B is overkill. The benchmark that matters is whichever one is shaped like your questions, and for most products in our experience that is the out-of-scope one, because production traffic does not respect your enum.
This is the same trap we wrote about in when not to trust a confidence score: a model's probability is only useful if it is calibrated on inputs like yours. Cloudflare says it trained Clef with label-smoothed cross-entropy and a Brier loss term specifically for calibration, which is a good sign and not a substitute for measuring it on your data.
The Apache-2.0 licence is the real headline for anyone who has to keep decisions inside their own network. 27 billion parameters in BF16 is roughly 54 GB of weights before any KV cache, which means a multi-GPU box or a single 80 GB accelerator, plus a vision encoder in the same process. Clef-flash at 9B is the version that fits ordinary hardware. If self-hosting is the whole reason you are interested, read what you can and cannot self-host first — the constraints we described there apply with more force to a model five times the size.
If you are not self-hosting, the economics run the opposite way to what "open source beats proprietary" suggests. Cloudflare's pricing table lists Clef at $0.240 per million input tokens and Clef-flash at $0.090, both input-only — there is no output rate, for the same reason TypeSafe charges none: a decision model generates no text to bill. Jev is published at $42 per billion input tokens, which is $0.042 per million. So Clef costs 5.7 times the incumbent per input token, and Clef-flash 2.1 times.
Two multiples make a sharper decision than one premium. Clef buys you the accuracy — 97.43 on out-of-scope detection against Jev's 89.27 — at 5.7x. Clef-flash gives most of that money back at 2.1x and answers in 38.8 ms, but scores 66.77 on that same benchmark, well below the model you would be leaving. Paying 5.7x for accuracy you can measure on your own traffic is a defensible decision. Paying 2.1x for a regression in out-of-scope handling is not, and that is the trap in reading the latency column first. We have written about what $0.042 per million tokens really costs if you want the per-decision arithmetic for the incumbent side.
Self-hosting is where the direction flips: fixed GPU cost, unlimited decisions, no per-token line item, and all the operational work that implies. That trade is the decision, not the latency table.
Read the p95 column too, because that is the number your users feel. Clef's p95 of 238.6 ms sits 14% above its median, which is a tight distribution. Clef-flash's p95 of 122.4 ms is more than three times its 38.8 ms median, so the model that looks thirteen times faster on average is also the burstiest of the three. Jev's p95 of 536.0 ms is barely above its median — slower, but the most predictable of the three, which matters if you size timeouts rather than averages.
Less code than you expect, and more evaluation than you planned for. Cloudflare describes Clef as Jev-API compatible, and the Workers AI documentation presents its image support as an extension to the System One API, so the request shape you already build is largely the request shape it accepts. Concretely, a migration looks like this:
env.AI.run() or the REST endpoint /ai/run, or against your own inference server if you self-host. The question payload — noul, choice and score types — carries over; if those names are unfamiliar, we covered them in Choice, Score and Noul with examples.0.85 you tuned against Jev's probabilities is meaningless against a differently-trained head. Every auto-approve cutoff, every escalate-to-human rule, every retry trigger needs to be re-derived from a labelled sample of your own traffic.In the systems we build, that recalibration pass — not the integration — is what sets the timeline. Budget weeks of labelled comparison, not an afternoon.
Jev is proprietary and served through TypeSafe's API, so your fallback if the terms change is a rewrite. Clef's weights are Apache-2.0 and already downloaded, which means even a team running it on Workers AI today holds a permanent escape route: the same model, on your own hardware, at fixed cost. That optionality is worth more than any single benchmark row and it is the part of the launch that is genuinely hard to argue with.
The counter-argument deserves its strongest form. Running Clef on Workers AI couples you to Cloudflare's platform as surely as Jev couples you to TypeSafe — the binding, the gateway, the edge. Cloudflare is also launching a reinforcement-learning fine-tuning service for these models, which initially runs as a hands-on engagement with its forward-deployed engineering team rather than self-serve, with no published price and no general-availability date as of 2 October 2026. A tuned model you cannot take elsewhere is a new dependency wearing an open-source licence. We go into that platform in our separate explainer on Cloudflare's Clef models.
There is a third option that neither vendor will raise: keep the hosted model for the long tail and run a small local model for the hot path. We described that pattern in running Jev-style decision models on your hardware, and Clef-flash at 9B makes it more practical than it was a month ago.
Stay if your current decisions are accurate enough and 524 ms is inside your budget. Latency only matters where it is the constraint; if the decision happens inside a request that already waits 900 ms on a database, halving the model's share changes nothing a user can feel.
Stay if your out-of-scope behaviour is tuned and working. You would be trading a known calibration for an unknown one in exchange for benchmark rows you cannot reproduce.
Stay if you cannot self-host and your volume is high, because the per-token premium on Workers AI compounds with every decision.
Move, or at least evaluate seriously, if the weights being Apache-2.0 solves a real problem — data residency, air-gapped deployment, a vendor you are not allowed to depend on — or if your decision sits on the critical path of something interactive where 300 ms of saved latency is visible. Those are good reasons. "It won seven of ten benchmarks the challenger selected" is not one.
If you want this evaluated against your own traffic rather than against Cloudflare's, that is the kind of work our AI engineering practice does: build the labelled sample, shadow both models, and decide on your numbers. The honest first step costs a week and tells you whether the rest of the migration is worth starting.
Migrating from Jev to Clef is worth it if Apache-2.0 weights solve a real constraint such as data residency or air-gapped deployment, or if saved latency is visible to users. If your thresholds are tuned and 524 ms fits your budget, staying on Jev is the cheaper decision.
Cloudflare publishes Clef at 209.3 ms median latency against Jev at 524.1 ms, and Clef-flash at 38.8 ms. Those figures come from Cloudflare, which competes with Jev, and no independent reproduction existed as of 2 October 2026. Treat them as vendor claims.
Clef and Clef-flash weights are released under Apache-2.0 on Hugging Face, so self-hosting is free of licence cost. Running them on Workers AI is not: Cloudflare lists Clef at $0.240 and Clef-flash at $0.090 per million input tokens, against TypeSafe's published $0.042 per million for Jev.
Yes, within reason. Clef is 27 billion parameters in BF16, roughly 54 GB of weights before any KV cache, so it needs a single 80 GB accelerator or a multi-GPU host. Clef-flash at 9 billion parameters is the version that fits ordinary inference hardware.
Recalibration, not integration. Cloudflare describes Clef as Jev-API compatible, so the request shape mostly carries over. Every probability threshold you tuned against Jev must be re-derived from a labelled sample of your own traffic, because a different model's probabilities are not interchangeable.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand