{"slug": "architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents", "title": "Architectural Breakdown: How to use the OpenAI Decisions API with Strands Agents", "summary": "A developer describes replacing a naive Strands agent loop — which sent full conversation history and every tool definition to gpt-4o on each iteration, causing unbounded token growth, no confidence checks, and no concurrency limits — with a \"decision gate\" built on the OpenAI Decisions API that classifies one option from a bounded list of 2 to 20 and returns a confidence score. The gate adds an AsyncMutex semaphore capped at 3 concurrent calls, retry handling, and a circuit breaker that opens after 5 failures and resets after 30 seconds. The rewrite followed an incident in which the agent consumed 4.2 GB of RAM on an 8 GB instance and entered a retry spiral that cost 47 minutes of uptime.", "body_md": "\n\n```\n![Architecture Diagram](https://image.pollinations.ai/prompt/high+performance+cloud+systems+How+to+use+the+OpenAI+Decision+round+2?width=800&height=400&nologo=true)\n\n# Your Strands Agent Is a Garbage Disposal, Here's How to Stop It\n\nIt was 3:14 AM on a Tuesday. Our Strands agent chewed through 4.2 GB of RAM on an 8 GB instance and started swapping. The root cause wasn't exotic. Nobody used a decision gate. Every loop step sent the entire conversation history through `gpt-4o` to pick one tool.\n\nA customer ticket escalated. The agent entered a retry spiral. We lost 47 minutes of uptime while an SRE team pulled the plug.\n\nThis is how we stopped it.\n\n## The Naive Pattern (What Everyone Writes First)\n\nYou feed the full conversation history plus every tool definition into a chat completion endpoint on every single iteration. Sounds reasonable until you do the math on a 50-step task:\n```\n\ntypescript\n\n// What your senior engineer wrote because \"it worked locally\"\n\nasync function agentLoop(context: string, tools: Tool[]): Promise {\n\n  const response = await fetch('[https://api.openai.com/v1/chat/completions](https://api.openai.com/v1/chat/completions)', {\n\n    method: 'POST',\n\n    headers: {\n\n      'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`,\n\n      'Content-Type': 'application/json',\n\n    },\n\n    body: JSON.stringify({\n\n      model: 'gpt-4o',\n\n      messages: [\n\n        { role: 'system', content: `Available tools: ${tools.map(t => t.name).join(', ')}` },\n\n        ...contextMessages, // Grows without bound. O(n^2) token cost.\n\n      ],\n\n      response_format: { type: 'json_object' },\n\n    }),\n\n  });\n\nconst decision = (await response.json()).choices[0].message.content;\n\n  // No confidence check. No schema validation. No concurrency limit.\n\n  // Model might return \"call_search_tool\" instead of \"search_tool\".\n\n  // Your runtime crashes. Your user gets a broken experience.\n\n}\n\n```\nThree things go wrong, guaranteed:\n\n**Context compounding.** Each iteration appends history. Token cost compounds multiplicatively. By step 20, you're paying for 210 cumulative context sizes. By step 50, you're setting money on fire.\n\n**No confidence threshold.** The model picks whatever it feels like. Sometimes randomly. Sometimes selecting a tool that doesn't exist. Null pointer exceptions at 3 AM instead of clean failures at build time.\n\n**Zero concurrency control.** Ten requests fire at once. API rate limits hit. Retry logic loops forever because nobody built one. The process eats memory and never recovers.\n\n## The Decision Gate (What Actually Works)\n\nA decision problem is classification, not generation. You don't need a model that writes essays to pick one option from a bounded list. You need a model that classifies.\n\nThe OpenAI Decisions API accepts a problem statement, 2 to 20 options, and returns a selected option with a confidence score. No freeform text. No schema drift. No context window that grows until it kills you.\n```\n\ntypescript\n\nimport { createConnection } from 'node:https';\n\ninterface DecisionResult {\n\n  selected: string;\n\n  confidence: number;\n\n  reasoning?: string;\n\n}\n\nclass DecisionGate {\n\n  private readonly semaphore = new AsyncMutex(3);\n\n  private failureCount = 0;\n\n  private circuitOpen = false;\n\n  private circuitOpenAt = 0;\n\nasync decide(problem: string, options: string[]): Promise {\n\n    if (options.length < 2 || options.length > 20) {\n\n      throw new Error(`Decisions API requires 2-20 options, got ${options.length}`);\n\n    }\n\n```\nif (this.circuitOpen) {\n  if (Date.now() - this.circuitOpenAt > 30_000) {\n    this.circuitOpen = false;\n    this.failureCount = 0;\n  } else {\n    throw new Error('Circuit open. Decisions API unavailable');\n  }\n}\n\nawait this.semaphore.acquire();\ntry {\n  const result = await this.withRetry(problem, options);\n  this.failureCount = 0;\n  return result;\n} catch (e) {\n  this.failureCount++;\n  if (this.failureCount >= 5) {\n    this.circuitOpen = true;\n    this.circuitOpenAt = Date.now();\n  }\n  throw e;\n}\n```\n\n}\n\nprivate async withRetry(\n\n    problem: string,\n\n    options: string[]\n\n  ): Promise {\n\n    for (let attempt = 0; attempt <= 2; attempt++) {\n\n      try {\n\n        const resp = await fetch('[https://api.openai.com/v1/decisions](https://api.openai.com/v1/decisions)', {\n\n          method: 'POST',\n\n          headers: {\n\n            'Content-Type': 'application/json',\n\n            Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,\n\n            'Connection': 'keep-alive',\n\n          },\n\n          body: JSON.stringify({ model: 'o3', problem, options }),\n\n        });\n\n```\n    if (!resp.ok) {\n      if (resp.status === 429 && attempt < 2) {\n        await this.backoff(attempt);\n        continue;\n      }\n      throw new Error(`Decisions API returned ${resp.status}`);\n    }\n\n    const data = await resp.json();\n    if (data.decision.confidence < 0 || data.decision.confidence > 1) {\n      throw new Error(`Invalid confidence: ${data.decision.confidence}`);\n    }\n    return data.decision;\n  } catch (e) {\n    if (attempt === 2) throw e;\n    await this.backoff(attempt);\n  }\n}\n```\n\nprivate backoff(attempt: number): Promise {\n\n    const base = 250 * Math.pow(2, attempt);\n\n    const jitter = Math.random() * 200;\n\n    return new Promise(r => setTimeout(r, base + jitter));\n\n  }\n\n}\n\nclass AsyncMutex {\n\n  private acquired = 0;\n\n  private readonly queue: Array<() => void> = [];\n\nconstructor(private readonly limit: number) {}\n\nasync acquire(): Promise {\n\n    if (this.acquired < this.limit) {\n\n      this.acquired++;\n\n      return;\n\n    }\n\n    return new Promise(resolve => this.queue.push(resolve));\n\n  }\n\nrelease(): void {\n\n    if (this.queue.length > 0) {\n\n      this.queue.shift()!();\n\n    } else {\n\n      this.acquired--;\n\n    }\n\n  }\n\n}\n\n```\nThe `AsyncMutex` matters. A counter-based semaphore has a TOCTOU race where two requests can observe the same count and both proceed past the limit. The promise-queue version defers acquisition until a slot is genuinely available. No races. No silent overflows.\n\n## The Router: Because Confidence Is a Number, Not a Suggestion\n```\n\ntypescript\n\nclass DecisionRouter {\n\n  constructor(\n\n    private readonly gate: DecisionGate,\n\n    private readonly confidenceThreshold = 0.65\n\n  ) {}\n\nasync select(problem: string, options: string[]): Promise {\n\n    const result = await this.gate.decide(problem, options);\n\n```\nif (result.confidence < this.confidenceThreshold) {\n  throw new Error(\n    `Confidence ${result.confidence.toFixed(2)} below threshold. ` +\n    `Selected: \"${result.selected}\". Reasoning: ${result.reasoning}`\n  );\n}\n\nif (!options.includes(result.selected)) {\n  throw new Error(`Selected unknown option: ${result.selected}`);\n}\n\nreturn result;\n```\n\n}\n\n}\n\n```\nBelow 0.6, the model is guessing. Above 0.8, you escalate everything and your humans drown in tickets. We landed on 0.65 after two weeks of production logging. Adjust for your own failure patterns.\n\n## The Agent Loop: Bounded or Broken\n```\n\ntypescript\n\nclass BoundedAgentLoop {\n\n  private readonly stepQueue: Array<{ problem: string; options: string[] }> = [];\n\nconstructor(\n\n    private readonly router: DecisionRouter,\n\n    private readonly maxSteps = 50,\n\n    private readonly maxQueueDepth = 100\n\n  ) {}\n\nenqueue(problem: string, options: string[]): void {\n\n    if (this.stepQueue.length >= this.maxQueueDepth) {\n\n      throw new Error('Queue full. Apply backpressure upstream.');\n\n    }\n\n    this.stepQueue.push({ problem, options });\n\n  }\n\nasync run(\n\n    taskRunner: (problem: string, options: string[]) => Promise\n\n  ): Promise {\n\n    let steps = 0;\n\n```\nwhile (this.stepQueue.length > 0 && steps < this.maxSteps) {\n  const { problem, options } = this.stepQueue.shift()!;\n  const result = await this.router.select(problem, options);\n  await taskRunner(problem, options);\n  steps++;\n}\nThe queue caps at 100. Prevents memory growth when upstream producers enqueue faster than the loop drains. The step cap at 50 prevents infinite loops where the gate keeps picking the same tool with high confidence but the tool returns garbage every time.\n\n## What This Looks Like Under Load\n\n| Component | Peak RAM | Notes |\n|---|---|---|\n| DecisionGate (idle) | ~12 MB | TLS pool, no in-flight requests |\n| AsyncMutex state | <1 MB | Promise queue, bounded by semaphore |\n| BoundedAgentLoop (100 tasks) | ~24 MB | Queue array, hard-capped |\n| **Total steady-state** | **~55 MB** | Well within 8 GB |\n| **Max burst (3 concurrent)** | ~120 MB | Three in-flight payloads |\n\nThe naive pattern burns 4 GB in 11 minutes. This stays under 120 MB under burst conditions. That's the difference between a system that runs and one that requires a 3 AM restart.\n\nFor teams who want a production-ready scaffold that includes this pattern out of the box, [ShipMVP's rapid development stack](https://www.shipmvp.tech) ships with the decision gate already wired into their Strands templates. Real production builds, not demo code.\n\nHere's the question I still think about: when should a decision gate route low-confidence cases to a smaller model for clarification instead of hard-failing? We chose the hard path because silent escalation creates undetectable degradation. But I've seen teams run a cheap fallback model before throwing the error. Different tradeoff. What would you do?\n```\n\n", "url": "https://wpnews.pro/news/architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents", "canonical_source": "https://dev.to/agenticstack/architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents-h85", "published_at": "2026-10-08 00:08:28+00:00", "updated_at": "2026-10-08 00:16:57.458786+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "developer-tools"], "entities": ["OpenAI", "Strands Agents", "OpenAI Decisions API", "gpt-4o"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents", "markdown": "https://wpnews.pro/news/architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents.md", "text": "https://wpnews.pro/news/architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents.txt", "jsonld": "https://wpnews.pro/news/architectural-breakdown-how-to-use-the-openai-decisions-api-with-strands-agents.jsonld"}}