Creuto is now an OpenAI Select Partner Read More

AI & Machine Learning

Strands Decider 2B: the decision model that fits on a laptop

Strands decider 2B is 1.9B parameters, Apache-2.0 and runs on a laptop. What AWS claims, what the model card contradicts, and when to use it.

Strands Decider 2B: the decision model that fits on a laptop

Strands decider 2B is an open-source decision model AWS released on 1 October 2026 under the Apache-2.0 licence: 1.9 billion parameters, no ability to generate text at all, and a median of roughly 115 ms per decision on an RTX 3090. At that size it is the first of the three decision models now on the market that you can run on an ordinary laptop.

The surprising part is what AWS does not claim for it. Secondary coverage has the model ranking first on JevBench among models that publish a full training recipe. AWS's own announcement claims something much narrower — "3rd of 33 in the 2B class, and 1st of 30 excluding the just-over-2B models" — and the repository's own benchmark page places the model 50th of 89 ranked systems overall. On the same page, AWS names a competitor on the same base model that beats it by eight tasks.

What strands decider 2B actually is, and what AWS claims for it

The release is three artefacts: weights on Hugging Face, code and training recipe on GitHub, and a pip package. Here is what the primary sources say, as of 3 October 2026.

MeasurementFigureWhere it is published
Parameters1.9B (AWS's blog rounds to "2 billion")Repository README
LicenceApache-2.0, weights and codeModel card and repository
JevBench v1 public accuracy0.723 (167 of 231 tasks)Model card and repository
Tier split (easy / standard / hard)1.000 / 0.875 / 0.505Repository benchmark page
Latency, RTX 3090, median / p95115 ms / 299 msRepository README
Latency, M3 Pro, warm median153 ms under 300 tokens, 234 ms over all tasksRepository results page
Board position, 25 September 202650th of 89 systems; 3rd of 33 at 2BRepository benchmark page

The accuracy figure is the one thing every source agrees on. The Hugging Face model card records 167 of 231 tasks, and so does the GitHub repository. Both are AWS's own measurement against JevBench, run locally rather than by the benchmark's maintainers — a point the repository makes itself, noting that its run is not the published leaderboard composite because speed, cost and the sealed held-out items are missing from it.

Every head-to-head number in this market is published by a party with an interest in the result. AWS ran JevBench against its own model; Cloudflare published the Clef figures we covered when it released its typed-probability models; TypeSafe publishes Jev's. None has been independently reproduced. The useful response is to notice which claims are checkable — a licence, a parameter count and the contents of a git repository are, in a way a benchmark score is not.

Removing the language-modelling head is a different bet from Clef's

Most decision models put a classification head over the backbone's final hidden states. AWS does something else. The architecture document describes discarding the language-modelling head entirely and replacing it with a pointer head of about a million parameters that takes the hidden state at the <answer> position as a query and the hidden state at the last token of each option's line as keys, then scores one against the other. One forward pass. No decoding loop.

Three things follow, and they are the practical argument for the design. Because the head holds no per-option parameters, nothing in the model can learn that the first option is usually the right one. Nothing caps how many options a question may carry — the schema's limit of 255 is the only ceiling. And the label set is defined by the request rather than baked into the weights, so adding a category means editing a prompt, not retraining a classifier.

The costs are stated too, which is unusual and worth credit. Reversing the option list still moves the distribution by about 0.016, because options attend to earlier options under causal attention, and eliminating that would need per-option attention isolation the release does not implement. AWS also notes a bidirectional encoder would pool better in principle, and that it kept attention causal because the conversion needs retraining it could not afford on a single 3090.

What the published material does not support is a claim that the pointer head is better than a head over final hidden states. There is no controlled comparison between the two designs anywhere in the release. The nearest thing is a competitor, decider-2b, built on the identical Qwen3.5-2B-Base torso with a trained readout instead: its board entry scores 164 tasks to AWS's 167, but the repository records its newest release at 175 on AWS's own harness — eight tasks ahead. Our reading is that the architecture is a defensible bet with a clear mechanical rationale, and that the evidence it is the winning bet does not yet exist.

1.9B is the story: what fits on a laptop, and what does not

Size is where this release separates from the other two, and the arithmetic is blunt. Strands decider's base weights are a 4.5 GB download, per the repository's inference notes. Clef, the 27-billion-parameter model Cloudflare published the same week, needs roughly 54 GB in BF16 before any KV cache — that is simply 27 billion parameters at two bytes each — which means an 80 GB accelerator or a multi-GPU box. Clef-flash at 9B is the version that fits ordinary hardware, and Jev publishes no weights at all.

Three consequences matter when you are choosing.

  • Per-request cost goes to zero. Cloudflare lists Clef at $0.240 per million input tokens and Clef-flash at $0.090 on Workers AI, against Jev's $0.042 — figures we worked through in detail in what $0.042 per million tokens really costs. A model on your own hardware has no per-token line item at all. It has a fixed cost and unlimited decisions, which is a different kind of budget and a different kind of operational work.
  • Air-gapped deployment becomes possible. Apache-2.0 weights you can download are the only answer when the data cannot leave your network, which is the constraint we kept running into when we looked at what you can and cannot self-host in this category. Clef's licence permits it; its memory footprint makes it a procurement exercise. Strands decider makes it a laptop.
  • You can retrain it. AWS publishes the recipe and puts the cost at about 11 hours on one RTX 3090, or roughly an hour and ten minutes on eight H100s. That is an afternoon, not a capital request.

The latency without a GPU is where expectations need adjusting. The headline 115 ms median is CUDA. On an M3 Pro the repository's measurements give a warm median of 153 ms for tasks under 300 tokens and 234 ms across all tasks — comfortably interactive. But the first request at each new input length pays a one-off compile: 310 ms for short tasks, 1,644 ms for tasks of 300 to 1,000 tokens. Longer tasks stay slow even warm, at 2,172 ms for 1,000 to 2,500 tokens against 228 ms on the 3090. The Mac is fine for short, repetitive decisions and poor for long documents, which is a usable rule.

One design detail makes the economics better than the per-decision latency suggests: the state is encoded once and its cache broadcast across every question in the request. On the repository's measurement, eight questions over a 2,000-token state take 369 ms through the shared prefix against 1,964 ms batched. If you are asking one model six things about the same ticket, you pay for the ticket once.

Calibration matters more than accuracy, and the two primaries disagree about it

A decision model that is confidently wrong is worse than one that abstains, because you built an auto-approve threshold on top of it. That makes the Brier score and the expected calibration error the figures to read first — and here the model card and the repository do not agree.

SourceWindowBrier / ECE
Hugging Face model card4096 tokens0.348 / 0.050
Hugging Face model card3072 tokens0.349 / 0.056
GitHub repository README3072 tokens (pre-registered)0.342 / 0.052
GitHub benchmark page4096 tokens0.3416 / 0.0507

The gap is small — about 0.006 of Brier — but both documents describe the same checkpoint, and they also disagree on accuracy at the larger window: the model card records 167 of 231 tasks at both 3072 and 4096, while the repository says 168 at 4096 and calls that flip noise rather than a gain. We treat the model card as authoritative, because it is the artefact you download, and would quote 0.348 Brier and 0.050 ECE. Anyone citing 0.342 is reading a real AWS document too, so name both.

What those numbers mean for a threshold is more useful than the numbers themselves. The repository breaks the calibration down by confidence band: of the 53 answers the model gave at 0.9 confidence or above, all were correct, at a mean claimed confidence of 0.957. The middle band overclaims — 0.707 correct against a mean claim of 0.736. So a high threshold behaves conservatively and a mid threshold does not, which is the shape you want but not the shape you should assume.

AWS says so itself. The model card lists as a limitation that calibration is one temperature per question type, fitted on held-out short classification tasks, and that you should measure on your own traffic before trusting a threshold. That is the same warning we gave about Jev in when not to trust a confidence score, and it is the single most important sentence in the release. A confidence number earned on one distribution tells you very little about a cutoff on another.

Open weights and an open recipe are different claims

This is the most defensible thing in the story, and it is the part the coverage blurs. Clef and strands decider are both Apache-2.0 open-weight models. Only one of them ships the recipe.

What AWS publishes alongside the weights: the training scripts and configs, a data inventory naming every source with its revision, licence and role, the stage timings and data hashes from the reference run, the evaluation harness, and a research log in which every run since v9 states its predictions and failure conditions before training and appends the outcome without editing what came before. The repository notes that most runs missed their bar, that five were promoted by the maintainers' decision anyway, and that the record says so. It also discloses a one-line patch it made to the benchmark harness to run on Windows.

That is a materially stronger claim than open weights, and it is checkable by anyone with a git client — which is why it is worth more than a benchmark position. But it is not the claim that "first on JevBench among models publishing a full recipe" makes. The ranking AWS publishes is first of 30 at 2B parameters or fewer, excluding three models just over the line; counting them, third of 33. The "full recipe" framing appears to come from a comparison table in MarkTechPost's launch coverage, which lists recipe publication as a separate column rather than as the basis of the ranking. Two true facts, joined into a claim neither supports.

Where each of the three actually fits

The honest summary is that three vendors now sell the same idea at three sizes, and the size is the decision. Strands decider is the right choice when the decisions are short and repetitive, when the data cannot leave your network, or when you want to retrain on your own rubric — AWS's own documentation says score and noul transfer poorly to rubrics unlike the training mix, and the recipe is the remedy it offers. Clef is the right choice when accuracy on out-of-scope detection is the failure budget and you have somewhere to put 54 GB of weights, which we worked through in which decision model to build on. Jev remains the right choice when you want none of this operational surface and the hosted price is fine.

Strands decider is the wrong choice in two specific cases, both of which AWS states. It is wrong for long, multi-step documents: the hard tier of JevBench sits at 0.505, barely clear of chance on the long-policy family, and the repository's own window sweep shows that feeding the model the complete document does not fix it — the model reads the extra text and cannot use it. And it is wrong wherever you need deliberation: AWS says the single-pass design makes it "significantly worse at solving complex problems than reasoning models", and the top of the JevBench board is reasoning models at 0.996.

If you are already running a decision model in production, the work this release creates is not an integration. It is a shadow run: serve strands decider alongside what you have, on your own traffic, for long enough to re-derive the thresholds. A decision model produces no prose to eyeball, so a silent accuracy regression looks exactly like normal operation. That evaluation pass is what sets the timeline on the AI engineering work we do around these models, and a 4.5 GB model you can run on the laptop you already own makes it cheaper to start than it has ever been.

Frequently asked questions

Strands decider is an open-source decision model AWS released on 1 October 2026. It picks between options, rates on a scale and returns a calibrated confidence, but cannot generate text: its language-modelling head is replaced by a roughly one-million-parameter pointer head that scores each option in a single forward pass.

Strands decider is published under the Apache-2.0 licence, with weights on Hugging Face and code on GitHub. AWS also publishes the training scripts, the data inventory and the evaluation harness, which is a stronger claim than open weights alone and is verifiable from the repository.

Strands decider 2B has 1.9 billion parameters and its base weights are a 4.5 GB download, so it serves on a CPU, a consumer GPU or an Apple silicon Mac. AWS measures a warm median of 153 ms per short decision on an M3 Pro.

Strands decider 2B scores 0.723 on the JevBench v1 public set, answering 167 of 231 tasks correctly. By tier that is 1.000 on easy, 0.875 on standard and 0.505 on hard tasks, measured by AWS against the benchmark's own harness rather than by the benchmark's maintainers.

Strands decider can replace Jev where the decisions are short and repetitive or the data cannot leave your network, since it runs locally with no per-token cost. It is a poor replacement for long, multi-step documents, where AWS's own hard-tier score of 0.505 sits barely clear of chance.

JevBench is an independent, MIT-licensed benchmark for Jev-class typed decision models, with 231 public tasks across 18 families. Its board ranked 89 systems as of 25 September 2026, and AWS places its own model 50th overall and third of 33 in the 2B class.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

3 Oct 2026

·

11 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved