Creuto is now an OpenAI Select Partner Read More
The Jeeves decision model reports 0.889 accuracy against Jev's 0.857 - and loses on knowledge questions. What reasoning buys, and what it costs.

PostHog's Jeeves decision model reports 0.889 accuracy on held-out test data where Jev scores 0.857 and Kev-9B scores 0.822 — and loses to Jev on knowledge questions by a wider margin than it wins by. Both facts come from the same README. If you are evaluating the Jeeves decision model as a Jev replacement, the second one decides more cases than the first.
Jeeves is a 9B model built on Qwen3.5-9B with LoRA and a pointer head, described by its authors as a reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO. Unlike a Jev-style classifier that answers from the prompt alone, it writes a reasoning chain per question first, then decides. Every number below is the project's own, measured by the project, and published with the training code.
A decision model takes a state — a support ticket, a document, a chunk of JSON — and a set of typed questions, and returns a calibrated probability for each option rather than prose you then have to parse. Jeeves supports the same three question types in one request as Jev does: yes/no (noul), multiple choice (choice) and rating (score), through a Jev-compatible endpoint at /v1/systemone. If you already know what Jev is and how a System One model differs from an LLM, the request shape will look familiar, because the Python SDK is written as a drop-in replacement for typesafe-sdk.
The difference is the middle step. The README states the problem it set out to solve directly: Jev-like models give calibrated decision probabilities but at low accuracy, so a lot of pipelines fall back to a reasoning model anyway. Jeeves puts the reasoning inside the classifier. Mechanically, the prompt loads state and questions into the Qwen chat template, the model rolls out a reasoning chain, and after the closing think token the questions are repeated before a pointer head scores each option. The authors report that ablations found both choices load-bearing: substituting plain text markers for the rare Qwen tokens hurt performance, and not repeating the questions after the reasoning block hurt it too.
On this evidence, yes, but unevenly. The cleanest comparison is the same checkpoint with and without thinking: 0.840 with reasoning against 0.804 without, on a 2,962-item test split. That is 3.6 points for the entire reasoning apparatus — real, reproducible from the published code, and smaller than the headline gap over Jev suggests.
All figures below are from the Jeeves README, measured with thinking on, greedy decoding and a 2,560-token cap. The Jev and Kev-9B columns are the numbers Kev publishes, not independent re-runs.
| Benchmark | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain and held-out) | 0.822 | 0.857 | 0.889 |
| JevBench overall, 231 public items | 0.715 | 0.866 | 0.935 |
| JevBench hard, 111 public items | 0.451 | 0.730 | 0.865 |
| Transfer overall (MMLU-Pro and buried state) | 0.579 | 0.800 | 0.746 |
| MMLU | 0.738 | 0.900 | 0.793 |
| MMLU-Pro, 10-way | 0.515 | 0.840 | 0.739 |
| JevBench expected calibration error | — | 0.049 | 0.037 |
The bottom four rows are the interesting ones. Jeeves trails Jev on knowledge questions by 10.7 points on MMLU and 10.1 on MMLU-Pro, and its own README lists that as the first limitation. Reasoning helps a model apply a policy it has been given; it does not give a 9B model facts it never learned. The Kev-9B JevBench figures are marked in the README as Kev-8B on Qwen3, because no Kev-9B JevBench result is published, and the authors note that the Kev and Jev comparisons outside JevBench use different items from the same sources. That is unusually candid for a project benchmark table, and it is also a reason not to treat the gaps as precise.
Calibration is the quiet win. Expected calibration error on the public JevBench items falls from 0.049 to 0.037, and on the "unknowable" probe — questions answered at p ≥ 0.9 that should not have been answered at all, where lower is better — Jeeves scores 0.055 against Jev's 0.090 and Kev-9B's 0.000. If you route on thresholds, that matters more than a point of accuracy, and it is the same argument we made about the difference between confidence and probability and when not to trust either.
This is where the trade becomes a product decision. On one H100, Jeeves answers in about 0.3 seconds with thinking off and a 3.3 second median with it on — roughly eleven times slower at the median. The tail is worse: 17.1 seconds at p90 with full chains, which the README lists as a limitation in its own right.
| Setting | Accuracy | Mean reasoning tokens | Median / p90 latency |
|---|---|---|---|
| Full thinking | 0.825 | 1,138 | 3.3 s / 17.1 s |
max_think 768, nothink_threshold 0.9 | 0.806 | 344 | 2.0 s / 5.6 s |
| No thinking | 0.775 | 0 | about 0.3 s |
Those are 325 dev questions, so treat them as a shape rather than a specification. The shape is the useful part: the middle row buys back two thirds of the accuracy gain for a third of the tokens and a p90 that drops from 17 seconds to under 6. nothink_threshold is the mechanism — the model answers straight away when its no-think confidence is high enough and only reasons when it is not, which is exactly the escalation policy most teams were building by hand around a Jev-style model and an LLM fallback.
Cost follows latency, because you are paying for GPU seconds. Full thinking averages 1,138 reasoning tokens per request; the truncated setting averages 344. The block-4 diffusion drafter, adapted from Orthrus to support Qwen3.5's Gated DeltaNet layers, moves single-question chain decoding from 109 to 176 tokens per second, and about 960 tokens per second in total with eight questions batched. Block 8 is faster for one question at a time; block 4 is the serving default because it holds up under batching.
Yes, and that is most of the point. The fused weights, pointer head, fitted temperature and both drafters are on Hugging Face; you download them, serve them with python -m inference.serve, and send Jev-format requests to a local port with no API key. The full training pipeline is published too — SFT over 19,126 questions from 12 public datasets plus synthetic policy data, then CISPO reinforcement learning over 9,992 questions, then a single temperature fitted on dev.
Check the licensing before you plan around it, because there are two. The repository's LICENSE is MIT. The weights on Hugging Face are Apache-2.0, inherited from Qwen3.5-9B, and the README notes that each public dataset used to build the training data stays under its own licence. Both are permissive, but "MIT" alone is not an accurate description of what you are deploying. The hardware requirement is firmer: CUDA, and Hopper specifically for the FP8 kernel. We covered the general version of this calculation in what you can and cannot self-host among the Jev clones, and Jeeves lands on the genuinely self-hostable side of it.
Do not adopt these numbers. Adopt the method. A benchmark table published by the project that built the model tells you the model works on the items its authors chose; it tells you nothing about your tickets, your taxonomy or your edge cases. The README itself hands you the tools to check — test.py, jevbench.py and calibrate.py are all in the repository, along with the prep scripts that rebuild the datasets at pinned revisions.
In the classification systems we build, the sequence that settles this in a week is short. Pull 300 to 500 real, labelled items from production, weighted the way your traffic actually is rather than balanced. Run Jev and Jeeves over the same set through the same Jev-format request. Compare accuracy, but also compare expected calibration error and the rate of confident wrong answers at your routing threshold, because that is the number that generates escalations. Then run Jeeves again with max_think at 768 and nothink_threshold at 0.9 and see how much of the gain survives at a latency your product can absorb.
Two caveats the authors state plainly and we would not discover for you: evaluation is English-only, and no language consistency reward was included, so the reasoning chains are not reliably interpretable. If your plan was to show users the model's reasoning as an explanation, test that assumption first.
Jeeves is the wrong choice if your decisions turn on world knowledge rather than on a policy you supply — the MMLU gap is the evidence, and it points the other way. It is the wrong choice if your p99 budget is under a second and you cannot batch, because even the truncated setting has a 5.6 second p90. It is the wrong choice if you have no CUDA hardware and no appetite to rent it, since a Jev-compatible API you do not operate is cheaper than a GPU you do.
It is a strong choice when the decision is a policy application, when calibration matters more than raw accuracy, when the data cannot leave your infrastructure, and when you already have labelled examples to check it against. That last condition is not optional. If you want help building the evaluation set before the model choice rather than after it, that is what our AI engineering practice does. Figures cited are as published on 29 September 2026; the project acknowledges Kev as its inspiration.
Jeeves is a 9B open-weight decision model from PostHog, built on Qwen3.5-9B with LoRA and a pointer head. It writes a reasoning chain before returning a calibrated probability for each option, and answers yes/no, multiple-choice and rating questions through a Jev-compatible API endpoint.
On PostHog's own measurements it helps by a modest margin. The same Jeeves checkpoint scores 0.840 with reasoning and 0.804 without on a 2,962-item test split. Reasoning improves policy application and calibration more than it improves world knowledge, where Jeeves still trails Jev.
Yes. The fused weights, pointer head, fitted temperature and both speculative-decoding drafters are published on Hugging Face, and the inference server accepts Jev-format requests with no API key. You need a CUDA GPU, and Hopper specifically if you want the FP8 kernel.
PostHog reports about 0.3 seconds per request without thinking and a 3.3 second median with it on one H100, with a 17.1 second p90 on full chains. Capping max_think at 768 with a nothink_threshold of 0.9 brings the median to 2.0 seconds and the p90 to 5.6.
Two apply. The Jeeves code repository is MIT licensed, while the weights published on Hugging Face carry Apache-2.0, inherited from the Qwen3.5-9B base model. Each public dataset used to rebuild the training data stays under its own licence, so check those before redistributing data.
No. Every figure comes from PostHog's own README and evaluation scripts, and the Jev and Kev-9B columns are numbers Kev publishes rather than independent re-runs. The authors note that comparisons outside JevBench use different items from the same sources. Re-run the evaluation on your own labelled data.
On knowledge-heavy benchmarks. Jeeves scores 0.793 on MMLU against Jev's 0.900, and 0.739 on 10-way MMLU-Pro against Jev's 0.840, which drags its transfer overall score to 0.746 against Jev's 0.800. Its own README lists this as the first limitation.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand