{"slug": "running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-rag", "title": "Running AWS Strands Decider 2B Locally: A Complete Setup Guide for AI Routing & Multi-RAG Systems", "summary": "A developer documented a step-by-step process for installing and running AWS's Strands Decider 2B, a lightweight decision model for routing, classification, scoring, and agent orchestration, locally on Windows via WSL2. The guide covers setting up a Python virtual environment, installing the strands-decider package, serving the StrandsAgents/strands-decider-2B-hobson-v19 model on CPU, and calling its REST endpoint with choice, yes/no, and score decision formats. A sample choice query returned a \"PLM\" selection with 0.94 confidence, illustrating how a dedicated decision model can handle routing so an LLM focuses on reasoning and generation.", "body_md": "As GenAI applications become more sophisticated, one challenge continues to surface:\n\nHow do we make reliable decisions before invoking an LLM?\n\nFor example:\n\nTraditionally, we let an LLM make these decisions.\n\nRecently, AWS introduced **Strands Decider 2B**, a lightweight decision model designed specifically for routing, classification, scoring, and orchestrating agent workflows.\n\nUnlike traditional LLMs, Strands Decider doesn't generate arbitrary text. Instead, it selects from predefined options and provides confidence scores, making it ideal for Agentic AI and Multi-RAG systems.\n\nIn this article, I'll walk through how I installed and tested Strands Decider 2B locally on Windows using WSL2.\n\nA common architecture today looks like this:\n\nThe problem?\n\nThe LLM is responsible for both:\n\nA better approach is:\n\nNow the LLM focuses on reasoning and generation, while the decision model handles routing and orchestration.\n\nFor this walkthrough I used:\n\nOpen PowerShell:\n\n```\nwsl -l -v\n```\n\nExample output:\n\n```\nNAME      STATE    VERSION\nUbuntu    Running  2\n```\n\nLaunch Ubuntu:\n\n```\nwsl -d Ubuntu\nmkdir -p /mnt/c/GENAI/strands\n\ncd /mnt/c/GENAI/strands\n```\n\nUpdate Ubuntu:\n\n```\nsudo apt update\n```\n\nInstall required dependencies:\n\n```\nsudo apt install -y \\\n python3 \\\n python3-pip \\\n python3-venv \\\n python3-dev \\\n build-essential \\\n gcc \\\n g++\n```\n\nCreate the environment:\n\n```\npython3 -m venv .venv\n```\n\nActivate it:\n\n```\nsource .venv/bin/activate\n```\n\nUpgrade pip:\n\n```\npip install --upgrade pip setuptools wheel\npip install strands-decider\n```\n\nVerify installation:\n\n```\nstrands-decider --help\n```\n\nInstall Hugging Face Hub:\n\n```\npip install huggingface_hub\n```\n\nList available Strands models:\n\n``` python\npython -c \"from huggingface_hub import list_models; [print(m.id) for m in list_models(search='strands')]\"\n```\n\nThe model used in this guide:\n\n```\nStrandsAgents/strands-decider-2B-hobson-v19\n```\n\nLaunch the model:\n\n```\nstrands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --device cpu\n```\n\nExpected output:\n\n```\nApplication startup complete.\nUvicorn running on http://127.0.0.1:8000\n```\n\nThe first startup downloads and caches the model automatically.\n\nOpen in your browser:\n\n```\nhttp://127.0.0.1:8000/docs\n```\n\nOr:\n\n```\ncurl http://127.0.0.1:8000/openapi.json\n```\n\nStrands Decider supports three decision formats.\n\nChoose one option from a list.\n\n```\n{\n  \"type\": \"choice\",\n  \"instructions\": \"Select the best datasource.\",\n  \"criteria\": {\n    \"PLM\": \"Engineering changes and parts\",\n    \"JIRA\": \"Issue tracking system\",\n    \"CONFLUENCE\": \"Documentation repository\",\n    \"UNKNOWN\": \"No suitable source\"\n  }\n}\n```\n\nYes / No decision.\n\n```\n{\n  \"type\": \"noul\",\n  \"instructions\": \"Determine whether this statement is true.\"\n}\n```\n\nRate against an ordered scale.\n\n```\n{\n  \"type\": \"score\",\n  \"instructions\": \"Rate the sentiment.\",\n  \"criteria\": [\n    \"Very Negative\",\n    \"Negative\",\n    \"Neutral\",\n    \"Positive\",\n    \"Very Positive\"\n  ]\n}\n```\n\nCreate a file called:\n\n```\ndecider_demo.py\npython\nimport requests\nimport json\n\npayload = {\n    \"state\": \"User wants ECO information\",\n    \"questions\": {\n        \"datasource\": {\n            \"type\": \"choice\",\n            \"instructions\": \"Select the most appropriate datasource.\",\n            \"criteria\": {\n                \"PLM\": \"Engineering changes and parts\",\n                \"JIRA\": \"Issue tracking system\",\n                \"CONFLUENCE\": \"Documentation repository\",\n                \"UNKNOWN\": \"No suitable source\"\n            }\n        }\n    }\n}\n\nresponse = requests.post(\n    \"http://127.0.0.1:8000/v1/systemone\",\n    json=payload\n)\n\nprint(json.dumps(response.json(), indent=2))\n```\n\nRun:\n\n```\npython decider_demo.py\n```\n\nSample output:\n\n```\n{\n  \"answers\": {\n    \"datasource\": {\n      \"choice\": \"PLM\",\n      \"confidence\": 0.94\n    }\n  }\n}\n```\n\nOne use case I was particularly interested in was reducing hallucinations across multiple RAG systems.\n\nRouting logic becomes simple:\n\n```\ndecision = response[\"answers\"][\"datasource\"][\"choice\"]\n\nif decision == \"PLM\":\n    plm_rag.search(query)\n\nelif decision == \"JIRA\":\n    jira_rag.search(query)\n\nelif decision == \"CONFLUENCE\":\n    confluence_rag.search(query)\n\nelse:\n    print(\"No reliable datasource identified.\")\n```\n\nInstead of asking an LLM to guess which datasource to use, the decision model handles routing first.\n\nI see strong potential in the following scenarios:\n\n✅ Multi-RAG orchestration\n\n✅ Agent tool selection\n\n✅ Engineering Change workflows\n\n✅ PLM assistants\n\n✅ SharePoint routing\n\n✅ Confluence routing\n\n✅ Jira ticket management\n\n✅ Intent classification\n\n✅ Confidence-based validation\n\n✅ Hallucination reduction\n\nOne of the biggest lessons I've learned building GenAI applications is:\n\nNot every problem requires text generation.\n\nDecision-making and text generation are fundamentally different tasks.\n\nUsing a decision model before retrieval and generation creates a much cleaner architecture:\n\n```\n Decision Model\n      ↓\n  Retrieval\n      ↓\n     LLM\n```\n\nFor enterprise AI systems, agentic workflows, and multi-RAG architectures, this pattern improves reliability, control, and observability.\n\nIf you're building AI agents today, I highly recommend experimenting with decision models as part of your architecture.\n\n**Repository:** [https://github.com/ujjwalbsoni/strands-decider-end-to-end](https://github.com/ujjwalbsoni/strands-decider-end-to-end)\n\n**Ujjwalkumar Soni**\n\nPassionate about AI Agents, RAG Architectures, Knowledge Management, and Enterprise GenAI Solutions.\n\nLet's connect and share ideas around Agentic AI and next-generation enterprise applications.", "url": "https://wpnews.pro/news/running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-rag", "canonical_source": "https://dev.to/ujjwalbsoni/running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-multi-rag-systems-249c", "published_at": "2026-10-03 02:33:20+00:00", "updated_at": "2026-10-03 02:37:42.593433+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "ai-infrastructure"], "entities": ["AWS", "Strands Decider 2B", "StrandsAgents/strands-decider-2B-hobson-v19", "Hugging Face", "WSL2", "Ubuntu"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-rag", "markdown": "https://wpnews.pro/news/running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-rag.md", "text": "https://wpnews.pro/news/running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-rag.txt", "jsonld": "https://wpnews.pro/news/running-aws-strands-decider-2b-locally-a-complete-setup-guide-for-ai-routing-rag.jsonld"}}