AWS Open Sources a Fast Decision Model for AI Agents

AWS Strands Labs announced Strands Decider 2B on October 1, 2026, offering an open source model built for one focused job: making decisions inside AI agent workflows. With 1.9 billion parameters, it chooses between options or rates an answer on a scale instead of writing a long response.
The model runs on a CPU, a consumer GPU, or an Apple silicon Mac. Its weights are available on Hugging Face under Apache-2.0, and the weights, code, data list, and training recipe are published under the same license. Strands Decider 2B is free to download and use, with no hosted API or per-call fee.
A narrow model for quick agent decisions
Decision models, also called System One models, became a category after TypeSafe AI launched Jev last month. TypeSafe introduced Jev on September 15 and charges $0.042 per million input tokens. Jev takes application state and typed questions, then returns choices, scores, or yes/no probabilities through an API.
Strands Decider takes a similar focused approach, but its question format is built around three types: choice, noul, and score. Every answer comes from the allowed options and includes a confidence value. That design gives an agent a direct result instead of asking a language model to produce, interpret, and check a written answer.
The model is not meant to replace a reasoning model for complex problems. AWS says it performs worse on those tasks and is unsuited for coding, chat, or summarization. Its purpose is narrower: select among known options or assign a score when an agent needs a fast, structured decision.
The model’s architecture supports that goal. It replaces the language-modelling head of Qwen3.5-2B-Base with a small pointer head of about 1 million parameters. One forward pass produces the result, so the system does not use a decoding loop. The torso uses a rank-16 LoRA, while the head runs in fp32.
Speed, accuracy, and local access
The bundled installation offers a command-line interface and an HTTP server. Running pip install strands-decider provides both, while the bundled server binds to 127.0.0.1 with no authentication. No hosted inference provider serves the model yet, so users work with the local package rather than sending requests to a hosted endpoint.
AWS measured benchmarks and latency on JevBench and hardware such as an RTX 3090 and an M3 Pro. On an RTX 3090, Strands Decider records a 115-millisecond median latency and 299 milliseconds at p95. On an M3 Pro, its warm median latency is 153 milliseconds for inputs under 300 tokens.
AWS says the model can make local decisions in tens of milliseconds on short tasks and in under 100 milliseconds on some hardware and inputs. The company did not provide a general per-token or per-request operating-cost estimate for Strands Decider, but the local setup removes a hosted per-call fee.
On the v1.4.2 board dated September 25, Strands Decider 2B ranked third of 33 models in the 2B class. It ranked first of 30 when three models just over 2B were excluded. On JevBench public accuracy, it scored 0.723 with an ECE of 0.052.
AWS also measures accuracy and calibration using JevBench. Its v19 results show roughly 72% accuracy and a Brier score of roughly 0.35. Those measurements give users two ways to judge the model: whether its choices are correct and whether its confidence matches its performance.
How it compares with Jev and other open models
Jev remains a hosted alternative. TypeSafe reports end-to-end Jev responses ranging from roughly 70 to 500 milliseconds, while AWS lists Jev latency at a 106-millisecond median and 296 milliseconds at p95 on an RTX 3090. Jev costs $0.042 per million input tokens, while Strands Decider has no hosted API or per-call charge.
TypeSafe has not released Jev’s weights or full training recipe. Jev’s approach can also be influenced by adversarial text in an agent’s input. By contrast, Strands Decider publishes its weights, code, data list, and recipe under Apache-2.0, giving users access to the pieces needed to inspect and run the model locally.
Other work is exploring related ways to speed up agents. Researchers at Stanford and Nvidia developed open CLM-8B, which can cache representations of reusable actions. Mapika’s decider-2b v11 scores 175 of 231 on the Strands harness, with eight tasks ahead of Strands Decider 2B.
AWS plans to publish version 20 alongside the code, training data, and scripts. For developers building agents around fixed choices, scores, or confidence values, that release could make the model easier to test and adapt. The main tradeoff remains clear: Strands Decider is built for fast, structured decisions, not open-ended reasoning or conversation.
Based on




