AWS Open-Sources Strands Decider 2B for AI Agent Workflows

Amazon Web Services has released Strands Decider 2B, an open-source decision model designed to choose between predefined options and report a confidence score. The model is based on the Qen3.5-2B language-model backbone, is small enough to run locally, and was released through AWS Strands Labs, which develops tools and protocols for deploying AI agents.
The release arrived in the same week that OpenAI announced a similar offering. It reflects growing interest among AI developers in models aimed at computer automation rather than broad, frontier-scale text generation. Instead of producing prose, Strands Decider is intended to make calibrated choices within a closed set of answers.
A focused component for agent workflows
Amazon distinguished engineer Marc Brooker initiated the project after encountering Jev, the decision model developed by TypeSafe. His initial implementation briefly reached the top of the Jevbench ranking for models of its size, after which AWS engineers refined it for public release.
Brooker said customer conversations highlighted a gap in agentic workflows: not every step needs the capability or expense of a fully featured large language model. A decision model can address a narrower question—what the next action should be given the current state—while offering lower latency, potentially lower cost and confidence information for the choice it makes.
The distinction is important for systems that already constrain the available actions. A model that selects among known options can be assessed on both the correctness of its selection and the calibration of its confidence, rather than on the open-ended quality of generated text.
Open models join a growing category
TypeSafe named Jev after economist William Stanley Jevons, invoking the idea that lower costs can increase demand for a resource. Since TypeSafe introduced the concept, researchers have produced dozens of comparable models, indicating broad interest in inexpensive, fast decision components.
Brooker cautioned that speed alone is not the goal. He described a balance between improving accuracy and calibration on decision tasks and preserving language understanding and general knowledge, the capabilities that make a model useful across different contexts.
TypeSafe founder and CEO Diogo Almeida has also argued that reproducing the architecture is not equivalent to making such models genuinely capable. For businesses, the practical implication is to test small decision models on bounded workflow steps, measuring accuracy, confidence calibration, latency and cost before replacing a general-purpose model.

