{"slug": "from-playground-to-caching-learning-amazon-bedrock", "title": "From Playground to Caching: Learning Amazon Bedrock", "summary": "A developer working through the Formação AWS course built a ticket-clerk chatbot on Amazon Bedrock using Anthropic's Claude Sonnet 4.5, progressing from a single Bash script to an interactive Node.js version with streaming and prompt caching. The developer hit a ThrottlingException caused by the Cross-region model inference tokens per minute quota, which was inactive by default, and resolved it by requesting a Service Quotas increase to 5,000,000 tokens per minute. The account also documents Playground sampling parameters (Temperature, Top P, Top K) and System Prompt effects for tuning customer-facing responses.", "body_md": "This post is part of my notes from [Formação AWS](https://pages.formacaoaws.com.br/), a course by Henrylle Maia. It covers a class from Desafio Labs (one of the course's sections) that focuses on Amazon Bedrock.\n\nIn this class, I start with the Playground to get a feel for how Bedrock behaves, and then move on to scripts to learn how to actually use the Bedrock API.\n\nTo learn by doing, I created a simple chatbot that works like a clerk selling tickets for an AWS Course. It starts as a single Bash script that sends one message and prints the response, then I make it interactive so I can actually chat with it. From there, I move it to Node so I can stream the answers as they come in, and finally add a cache checkpoint so the same System Prompt doesn't get billed at full price on every message.\n\n**Amazon Bedrock** is a fully managed AWS service that provides API access to multiple foundation models. There are several models to choose from, but for this project I'm using Anthropic's Claude Sonnet 4.5.\n\nWhen I selected my preferred model in the Playground and tried to send a prompt, I received the following error message: `ThrottlingException: \"Too many tokens per day, please wait before trying again.\"`\n\nDigging into it, I discovered that there is a quota called **Cross-region model inference tokens per minute for Anthropic Claude Sonnet 4.5 V1**. It limits the combined input and output tokens that can be sent to that model per minute through the cross-region inference profile, across every API: Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.\n\n**Cross-region inference** is a Bedrock feature that routes requests across multiple AWS Regions instead of keeping them in just one, so if one Region is busy, another can still take the request. The routing is automatic, but the throughput is still limited per model and per Region, and that limit is what I was hitting.\n\nTo fix this, I went to the Service Quotas console, found that quota for Claude Sonnet 4.5, and requested an increase to 5,000,000 tokens per minute.\n\nOnce the request showed up as Case Closed, I could finally start testing in the Playground.\n\n**Note:** It looks like every model comes with its own quota, and none of them are active on the account by default, so each one has to be requested separately. This isn't very well documented or explained by AWS. I only found out how it works by poking around Service Quotas myself.\n\nThe **Playground** is where I can test the model's responses directly in the AWS console and easily adjust parameters like Temperature, Top P and Top K. These parameters help tune the kind of response I get from the model.\n\n**Temperature** is a value between 0 and 1. At 0, the model sticks to less creative options and returns a cold, mechanical response. At 1, it leans toward more creative options and may hallucinate more.\n\n**Top P** limits the model's vocabulary. A value of 0.5 means it only chooses from the pool of words whose probabilities add up to 50%, while 0.95 means it chooses from the pool of words that add up to 95%. So a higher value makes the response more unpredictable, as it can choose a word that had a probability of 5%, for example.\n\n**Top K** works a bit differently: it limits the model to the K most probable words. The higher the value, the bigger that pool of words gets for each prediction.\n\nFor the tests below, I kept the default value of Top K and set the temperature to 0.7, which seems to work well for customer relations. The responses aren't rigid and cold, but they also aren't so creative that the model starts to hallucinate.\n\nBesides these sampling parameters, the Playground also lets me set a **System Prompt**, which gives the model the context it needs to generate the response for any user query. To see how much difference it makes, I ran two tests.\n\nFor the first test, I used a very simple System Prompt, just \"Act like a ticket clerk\". Then I asked for tickets for the game on Sunday, and the model gave a generic response: it asked which game I meant, how many tickets I needed, my seating preference and my budget.\n\nFor the second test, I used a longer, more detailed System Prompt. I told the model that it sells tickets for an AWS Training Course with AI, and that it should talk ONLY about the regular and VIP tickets, including the price and what's included in each one.\n\nThis time, when I asked for tickets for the game on Sunday, the model didn't fall for it. It clarified that it only sells tickets for the AWS Training Course, not for a game, and presented the two ticket options instead. When I said I was interested in the VIP ticket, it went on to explain what's included in it.\n\nSo a detailed System Prompt makes a big difference in keeping the model focused on the task I actually want it to do.\n\nMoving from the console to code, I wrote a simple Bash script. It sets the System Prompt, sends a single user message, and prints the model's response along with the token metrics.\n\n```\n# V1 - Basic Chatbot with Converse API\n# Demonstrates only the use of System Prompt\n\n#Use inference profile (cross-region) - prefix \"us.\"\nMODEL_ID=\"us.anthropic.claude-sonnet-4-5-20250929-v1:0\"\nREGION=\"us-east-1\"\n\n# System Prompt: behavior + context\nBEHAVIOR=\"Act like a ticket clerk.\"\nCONTEXT=\"You will sell my AWS Training Course with AI.\nTalk ONLY about the regular and VIP tickets.\nRegular ticket: \\$15\nVIP ticket: \\$30\nRegular: Live Event + Recording available for 2 days\nVIP: Same from regular + extra class + Recording available for 7 days.\"\n\nSYSTEM_PROMPT=\"$BEHAVIOR. $CONTEXT\"\n\nUSER_PROMPT=\"I want to buy a ticket for the event\"\n\necho \"🤖 Ticket Selling Chatbot - V1 Basic (Converse API)\"\necho \"===================================================\"\necho \"\"\necho \"📝 System Prompt: \"\necho \"$SYSTEM_PROMPT\"\necho \"\"\necho \"⏳ Sending to Bedrock...\"\necho \"\"\n\nSYSTEM_JSON=$(jq -n --arg text \"$SYSTEM_PROMPT\" '[{\"text\": $text}]')\n\n# Calling Bedrock with Converse API\nRESPONSE=$(aws bedrock-runtime converse \\\n    --region $REGION \\\n    --model-id $MODEL_ID \\\n    --system \"$SYSTEM_JSON\" \\\n    --messages '[{\"role\":\"user\",\"content\":[{\"text\":\"'\"$USER_PROMPT\"'\"}]}]' \\\n    --inference-config '{\"maxTokens\":512,\"temperature\":0.7}' \\\n    --output json)\n\nASSISTANT_RESPONSE=$(echo $RESPONSE | jq -r '.output.message.content[0].text')\n\necho \"💳 ASSISTANT:\"\necho \"$ASSISTANT_RESPONSE\"\necho \"\"\n\n# Metrics\nINPUT_TOKENS=$(echo $RESPONSE | jq -r '.usage.inputTokens')\nOUTPUT_TOKENS=$(echo $RESPONSE | jq -r '.usage.outputTokens')\nTOTAL_TOKENS=$(echo $RESPONSE | jq -r '.usage.totalTokens')\n\necho \"📊 Metrics\"\necho \" Input Tokens: $INPUT_TOKENS\"\necho \" Output Tokens: $OUTPUT_TOKENS\"\necho \" Total Tokens: $TOTAL_TOKENS\"\n```\n\nAs you can see in the image below, the response followed the System Prompt and used a total of only 243 tokens.\n\nFor something closer to a real chatbot, I modified the script to keep accepting user prompts until the user quits. The System Prompt is also read from a file now, though I didn't change its contents for this part.\n\n```\n#V2 - Chatbot with Converse API in Interactive Loop\n...\n\n# Read System Prompt from external file\nSCRIPT_DIR=\"$(cd \"$(dirname \"${BASH_SOURCE[0]}\")\" && pwd)\"\nSYSTEM_PROMPT=$(cat \"$SCRIPT_DIR/system_prompt.txt\")\nSYSTEM_JSON=$(jq -n --arg text \"$SYSTEM_PROMPT\" '[{text: $text}]')\n\n# Initialize empty messages history\nMESSAGES_JSON='[]'\n\necho \"🤖 Ticket Selling Chatbot - V2 Interactive Loop\"\necho \"===================================================\"\necho \"\"\necho \"📄 System prompt carregado de: system-prompt.txt\"\necho \"\"\necho \"💡 Type 'quit' to exit\"\necho \"\"\n\nwhile true; do\n    echo -n \"👤 You: \"\n    read USER_INPUT\n\n    if [[ \"$USER_INPUT\" == \"quit\" ]]; then\n        echo \"👋 Bye!\"\n        break\n    fi\n\n    MESSAGES_JSON=$(echo \"$MESSAGES_JSON\" | jq --arg text \"$USER_INPUT\" '. += [{role: \"user\", content: [{text: $text}]}]')\n\n    # Calling Bedrock with Converse API\n    RESPONSE=$(aws bedrock-runtime converse \\\n        --region $REGION \\\n        --model-id $MODEL_ID \\\n        --system \"$SYSTEM_JSON\" \\\n        --messages \"$MESSAGES_JSON\" \\\n        --inference-config '{\"maxTokens\":512,\"temperature\":0.7}' \\\n        --output json)\n\n    ASSISTANT_RESPONSE=$(echo $RESPONSE | jq -r '.output.message.content[0].text')\n\n    MESSAGES_JSON=$(echo \"$MESSAGES_JSON\" | jq --arg text \"$ASSISTANT_RESPONSE\" '. += [{role: \"assistant\", content: [{text: $text}]}]')\n\n    # Show response\n    echo \"🤖 ASSISTANT:\"\n    echo \"$ASSISTANT_RESPONSE\"\n    echo \"\"\n\ndone\n```\n\nThe interactive version allows the user to ask follow-up questions within the same context.\n\nThen I moved the project to JavaScript and ran it with Node. The AWS SDK for JavaScript lets me call the streaming version of the Converse API.\n\n```\n// V3 - Streaming with SDK Node.js\n// Version with code streaming for real time responses\n\nimport { readFileSync } from \"node:fs\";\nimport { fileURLToPath } from \"node:url\";\nimport { dirname, join } from \"node:path\";\nimport readline from \"node:readline/promises\";\nimport { stdin as input, stdout as output } from \"node:process\";\nimport {\n  BedrockRuntimeClient,\n  ConverseStreamCommand,\n} from \"@aws-sdk/client-bedrock-runtime\";\n\n// Use inference profile (cross-region) - prefix \"us.\"\nconst MODEL_ID = \"us.anthropic.claude-sonnet-4-5-20250929-v1:0\";\nconst REGION = \"us-east-1\";\n\n// Read System Prompt from external file\nconst SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));\nconst SYSTEM_PROMPT = readFileSync(join(SCRIPT_DIR, \"complete_prompt.txt\"), \"utf-8\");\nconst SYSTEM_JSON = [{ text: SYSTEM_PROMPT }];\n\n// Initialize empty messages history\nlet messages = [];\n\nconst client = new BedrockRuntimeClient({\n  region: REGION,\n  maxAttempts: 5,\n  retryMode: \"adaptive\",\n});\n\nconst rl = readline.createInterface({ input, output });\n\nconsole.log(\"🤖 Ticket Selling Chatbot - V3 Streaming\");\nconsole.log(\"===================================================\");\nconsole.log(\"\");\nconsole.log(\"📄 System prompt carregado de: complete_prompt.txt\");\nconsole.log(\"\");\nconsole.log(\"💡 Type 'quit' to exit\");\nconsole.log(\"\");\n\nwhile (true) {\n  const userInput = await rl.question(\"👤 You: \");\n\n  if (userInput === \"quit\") {\n    console.log(\"👋 Bye!\");\n    break;\n  }\n\n  messages.push({ role: \"user\", content: [{ text: userInput }] });\n\n  // Calling Bedrock with Converse API\n  const response = await client.send(\n    new ConverseStreamCommand({\n      modelId: MODEL_ID,\n      system: SYSTEM_JSON,\n      messages,\n      inferenceConfig: { maxTokens: 512, temperature: 0.7 },\n    })\n  );\n\n  process.stdout.write(\"\\n💳 Assistant:\\n\");\n\n  let assistantText = \"\";\n  for await (const event of response.stream) {\n    if (event.contentBlockDelta?.delta?.text) {\n      const chunk = event.contentBlockDelta.delta.text;\n      process.stdout.write(chunk);\n      assistantText += chunk;\n    }\n  }\n\n  messages.push({ role: \"assistant\", content: [{ text: assistantText }] });\n  console.log(\"\\n\");\n}\n\nrl.close();\n```\n\nA streaming response makes the chatbot feel more natural. It feels like the bot is actually typing the answer (really fast) instead of the whole thing just showing up on the screen at once.\n\nNext, I added a cache checkpoint to the System Prompt to see how caching actually affects the token cost.\n\nCaching works by writing content to the cache once, at a **cache checkpoint** (in this case, right after the System Prompt). On the following requests, the model can read that content from the cache instead of processing it again as new input. Each cache checkpoint has a **Time To Live (TTL)**: if no request reuses it within that window, it expires, and the next request has to write it again.\n\nThere are two TTL options available, 5 minutes and 1 hour, and I tested both.\n\n```\n// V4 - Streaming with Cache Point on System Prompt (5m)\n...\n\nwhile (true) {\n  ...\n\n  // Calling Bedrock with Converse API\n  const response = await client.send(\n    new ConverseStreamCommand({\n      modelId: MODEL_ID,\n      system: [\n        { text: SYSTEM_PROMPT },\n        { cachePoint: { type: \"default\", ttl: \"5m\" }},\n      ],\n      messages,\n      inferenceConfig: { maxTokens: 512, temperature: 0.7 },\n    })\n  );\n\n  ...\n  let usage = {};\n\n  for await (const event of response.stream) {\n    ...\n    if (event.metadata?.usage) {\n      usage = event.metadata.usage;\n    }\n  }\n\n  ...\n\n  const output = usage.outputTokens ?? 0;\n  const cacheRead = usage.cacheReadInputTokens ?? 0;\n  const cacheWrite = usage.cacheWriteInputTokens ?? 0;\n  const inputNoCache = usage.inputTokens ?? 0;\n  const input = inputNoCache + cacheRead;\n\n  console.log(\"\\n\");\n  console.log(\"┌──────────────────────────────────────────────────┐\");\n  console.log(\"│                 📊  Token Usage                  │\");\n  console.log(\"├──────────────────────────────────────────────────┤\");\n  console.log(`│ Input total:        ${String(input).padStart(8)}  tokens             │`);\n  console.log(`│   ├ Cache read:     ${String(cacheRead).padStart(8)}  (10% of price)     │`);\n  console.log(`│   ├ Cache write:    ${String(cacheWrite).padStart(8)}  (125% of price)    │`);\n  console.log(`│   └ No cache:       ${String(inputNoCache).padStart(8)}  (full price)       │`);\n  console.log(`│ Output:             ${String(output).padStart(8)}  tokens             │`);\n  console.log(\"└──────────────────────────────────────────────────┘\");\n}\n```\n\nThe cache checkpoint goes right after the System Prompt in the `system` array, with a `ttl` field. For the 1-hour test, I used the same script and just changed the `ttl` to `\"1h\"`.\n\nWith the 5-minute cache, the first request had nothing to read yet, so it wrote the System Prompt to the cache. That write costs 125% of the normal price.\n\nOn the next request, still inside the 5-minute window, the model read the System Prompt from the cache instead of writing it again. That cache read costs only 10% of the normal price.\n\nThe 1-hour cache write costs more, 200% of the normal price, but it lasts a full hour instead of just 5 minutes.\n\nJust like with the 5-minute cache, reading from the 1-hour cache costs 10% of the normal price.\n\nSo the tradeoff is simple: a longer cache costs more to write, but it pays off when messages come in more than 5 minutes apart, since the 5-minute cache would have expired and been written again at full write price.\n\nThat's as far as I got in this class: the quota is sorted, the Playground is explored, and the ticket chatbot now streams its answers and reuses its System Prompt instead of paying full price for it on every message. More to come.", "url": "https://wpnews.pro/news/from-playground-to-caching-learning-amazon-bedrock", "canonical_source": "https://dev.to/pribeiro/from-playground-to-caching-learning-amazon-bedrock-3g4e", "published_at": "2026-09-28 05:10:36+00:00", "updated_at": "2026-09-28 05:17:50.330405+00:00", "lang": "en", "topics": ["ai-products", "large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["Amazon Bedrock", "AWS", "Anthropic", "Claude Sonnet 4.5", "Formação AWS", "Henrylle Maia", "Service Quotas", "Desafio Labs"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-playground-to-caching-learning-amazon-bedrock", "markdown": "https://wpnews.pro/news/from-playground-to-caching-learning-amazon-bedrock.md", "text": "https://wpnews.pro/news/from-playground-to-caching-learning-amazon-bedrock.txt", "jsonld": "https://wpnews.pro/news/from-playground-to-caching-learning-amazon-bedrock.jsonld"}}