cd /news/ai-products/from-playground-to-caching-learning-… · home › topics › ai-products › article
[ARTICLE · art-140788] src=dev.to ↗ pub= topic=ai-products verified=true sentiment=· neutral

From Playground to Caching: Learning Amazon Bedrock

A developer working through the Formação AWS course built a ticket-clerk chatbot on Amazon Bedrock using Anthropic's Claude Sonnet 4.5, progressing from a single Bash script to an interactive Node.js version with streaming and prompt caching. The developer hit a ThrottlingException caused by the Cross-region model inference tokens per minute quota, which was inactive by default, and resolved it by requesting a Service Quotas increase to 5,000,000 tokens per minute. The account also documents Playground sampling parameters (Temperature, Top P, Top K) and System Prompt effects for tuning customer-facing responses.

by read11 min views1 publishedSep 28, 2026

This post is part of my notes from Formação AWS, a course by Henrylle Maia. It covers a class from Desafio Labs (one of the course's sections) that focuses on Amazon Bedrock.

In this class, I start with the Playground to get a feel for how Bedrock behaves, and then move on to scripts to learn how to actually use the Bedrock API.

To learn by doing, I created a simple chatbot that works like a clerk selling tickets for an AWS Course. It starts as a single Bash script that sends one message and prints the response, then I make it interactive so I can actually chat with it. From there, I move it to Node so I can stream the answers as they come in, and finally add a cache checkpoint so the same System Prompt doesn't get billed at full price on every message.

Amazon Bedrock is a fully managed AWS service that provides API access to multiple foundation models. There are several models to choose from, but for this project I'm using Anthropic's Claude Sonnet 4.5.

When I selected my preferred model in the Playground and tried to send a prompt, I received the following error message: ThrottlingException: "Too many tokens per day, please wait before trying again."

Digging into it, I discovered that there is a quota called Cross-region model inference tokens per minute for Anthropic Claude Sonnet 4.5 V1. It limits the combined input and output tokens that can be sent to that model per minute through the cross-region inference profile, across every API: Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.

Cross-region inference is a Bedrock feature that routes requests across multiple AWS Regions instead of keeping them in just one, so if one Region is busy, another can still take the request. The routing is automatic, but the throughput is still limited per model and per Region, and that limit is what I was hitting.

To fix this, I went to the Service Quotas console, found that quota for Claude Sonnet 4.5, and requested an increase to 5,000,000 tokens per minute.

Once the request showed up as Case Closed, I could finally start testing in the Playground.

Note: It looks like every model comes with its own quota, and none of them are active on the account by default, so each one has to be requested separately. This isn't very well documented or explained by AWS. I only found out how it works by poking around Service Quotas myself.

The Playground is where I can test the model's responses directly in the AWS console and easily adjust parameters like Temperature, Top P and Top K. These parameters help tune the kind of response I get from the model.

Temperature is a value between 0 and 1. At 0, the model sticks to less creative options and returns a cold, mechanical response. At 1, it leans toward more creative options and may hallucinate more.

Top P limits the model's vocabulary. A value of 0.5 means it only chooses from the pool of words whose probabilities add up to 50%, while 0.95 means it chooses from the pool of words that add up to 95%. So a higher value makes the response more unpredictable, as it can choose a word that had a probability of 5%, for example.

Top K works a bit differently: it limits the model to the K most probable words. The higher the value, the bigger that pool of words gets for each prediction.

For the tests below, I kept the default value of Top K and set the temperature to 0.7, which seems to work well for customer relations. The responses aren't rigid and cold, but they also aren't so creative that the model starts to hallucinate.

Besides these sampling parameters, the Playground also lets me set a System Prompt, which gives the model the context it needs to generate the response for any user query. To see how much difference it makes, I ran two tests.

For the first test, I used a very simple System Prompt, just "Act like a ticket clerk". Then I asked for tickets for the game on Sunday, and the model gave a generic response: it asked which game I meant, how many tickets I needed, my seating preference and my budget.

For the second test, I used a longer, more detailed System Prompt. I told the model that it sells tickets for an AWS Training Course with AI, and that it should talk ONLY about the regular and VIP tickets, including the price and what's included in each one.

This time, when I asked for tickets for the game on Sunday, the model didn't fall for it. It clarified that it only sells tickets for the AWS Training Course, not for a game, and presented the two ticket options instead. When I said I was interested in the VIP ticket, it went on to explain what's included in it.

So a detailed System Prompt makes a big difference in keeping the model focused on the task I actually want it to do.

Moving from the console to code, I wrote a simple Bash script. It sets the System Prompt, sends a single user message, and prints the model's response along with the token metrics.


#Use inference profile (cross-region) - prefix "us."
MODEL_ID="us.anthropic.claude-sonnet-4-5-20250929-v1:0"
REGION="us-east-1"

BEHAVIOR="Act like a ticket clerk."
CONTEXT="You will sell my AWS Training Course with AI.
Talk ONLY about the regular and VIP tickets.
Regular ticket: \$15
VIP ticket: \$30
Regular: Live Event + Recording available for 2 days
VIP: Same from regular + extra class + Recording available for 7 days."

SYSTEM_PROMPT="$BEHAVIOR. $CONTEXT"

USER_PROMPT="I want to buy a ticket for the event"

echo "🤖 Ticket Selling Chatbot - V1 Basic (Converse API)"
echo "==================================================="
echo ""
echo "📝 System Prompt: "
echo "$SYSTEM_PROMPT"
echo ""
echo "⏳ Sending to Bedrock..."
echo ""

SYSTEM_JSON=$(jq -n --arg text "$SYSTEM_PROMPT" '[{"text": $text}]')

RESPONSE=$(aws bedrock-runtime converse \
    --region $REGION \
    --model-id $MODEL_ID \
    --system "$SYSTEM_JSON" \
    --messages '[{"role":"user","content":[{"text":"'"$USER_PROMPT"'"}]}]' \
    --inference-config '{"maxTokens":512,"temperature":0.7}' \
    --output json)

ASSISTANT_RESPONSE=$(echo $RESPONSE | jq -r '.output.message.content[0].text')

echo "💳 ASSISTANT:"
echo "$ASSISTANT_RESPONSE"
echo ""

INPUT_TOKENS=$(echo $RESPONSE | jq -r '.usage.inputTokens')
OUTPUT_TOKENS=$(echo $RESPONSE | jq -r '.usage.outputTokens')
TOTAL_TOKENS=$(echo $RESPONSE | jq -r '.usage.totalTokens')

echo "📊 Metrics"
echo " Input Tokens: $INPUT_TOKENS"
echo " Output Tokens: $OUTPUT_TOKENS"
echo " Total Tokens: $TOTAL_TOKENS"

As you can see in the image below, the response followed the System Prompt and used a total of only 243 tokens.

For something closer to a real chatbot, I modified the script to keep accepting user prompts until the user quits. The System Prompt is also read from a file now, though I didn't change its contents for this part.

#V2 - Chatbot with Converse API in Interactive Loop
...

SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
SYSTEM_PROMPT=$(cat "$SCRIPT_DIR/system_prompt.txt")
SYSTEM_JSON=$(jq -n --arg text "$SYSTEM_PROMPT" '[{text: $text}]')

MESSAGES_JSON='[]'

echo "🤖 Ticket Selling Chatbot - V2 Interactive Loop"
echo "==================================================="
echo ""
echo "📄 System prompt carregado de: system-prompt.txt"
echo ""
echo "💡 Type 'quit' to exit"
echo ""

while true; do
    echo -n "👤 You: "
    read USER_INPUT

    if [[ "$USER_INPUT" == "quit" ]]; then
        echo "👋 Bye!"
        break
    fi

    MESSAGES_JSON=$(echo "$MESSAGES_JSON" | jq --arg text "$USER_INPUT" '. += [{role: "user", content: [{text: $text}]}]')

    RESPONSE=$(aws bedrock-runtime converse \
        --region $REGION \
        --model-id $MODEL_ID \
        --system "$SYSTEM_JSON" \
        --messages "$MESSAGES_JSON" \
        --inference-config '{"maxTokens":512,"temperature":0.7}' \
        --output json)

    ASSISTANT_RESPONSE=$(echo $RESPONSE | jq -r '.output.message.content[0].text')

    MESSAGES_JSON=$(echo "$MESSAGES_JSON" | jq --arg text "$ASSISTANT_RESPONSE" '. += [{role: "assistant", content: [{text: $text}]}]')

    echo "🤖 ASSISTANT:"
    echo "$ASSISTANT_RESPONSE"
    echo ""

done

The interactive version allows the user to ask follow-up questions within the same context.

Then I moved the project to JavaScript and ran it with Node. The AWS SDK for JavaScript lets me call the streaming version of the Converse API.

// V3 - Streaming with SDK Node.js
// Version with code streaming for real time responses

import { readFileSync } from "node:fs";
import { fileURLToPath } from "node:url";
import { dirname, join } from "node:path";
import readline from "node:readline/promises";
import { stdin as input, stdout as output } from "node:process";
import {
  BedrockRuntimeClient,
  ConverseStreamCommand,
} from "@aws-sdk/client-bedrock-runtime";

// Use inference profile (cross-region) - prefix "us."
const MODEL_ID = "us.anthropic.claude-sonnet-4-5-20250929-v1:0";
const REGION = "us-east-1";

// Read System Prompt from external file
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));
const SYSTEM_PROMPT = readFileSync(join(SCRIPT_DIR, "complete_prompt.txt"), "utf-8");
const SYSTEM_JSON = [{ text: SYSTEM_PROMPT }];

// Initialize empty messages history
let messages = [];

const client = new BedrockRuntimeClient({
  region: REGION,
  maxAttempts: 5,
  retryMode: "adaptive",
});

const rl = readline.createInterface({ input, output });

console.log("🤖 Ticket Selling Chatbot - V3 Streaming");
console.log("===================================================");
console.log("");
console.log("📄 System prompt carregado de: complete_prompt.txt");
console.log("");
console.log("💡 Type 'quit' to exit");
console.log("");

while (true) {
  const userInput = await rl.question("👤 You: ");

  if (userInput === "quit") {
    console.log("👋 Bye!");
    break;
  }

  messages.push({ role: "user", content: [{ text: userInput }] });

  // Calling Bedrock with Converse API
  const response = await client.send(
    new ConverseStreamCommand({
      modelId: MODEL_ID,
      system: SYSTEM_JSON,
      messages,
      inferenceConfig: { maxTokens: 512, temperature: 0.7 },
    })
  );

  process.stdout.write("\n💳 Assistant:\n");

  let assistantText = "";
  for await (const event of response.stream) {
    if (event.contentBlockDelta?.delta?.text) {
      const chunk = event.contentBlockDelta.delta.text;
      process.stdout.write(chunk);
      assistantText += chunk;
    }
  }

  messages.push({ role: "assistant", content: [{ text: assistantText }] });
  console.log("\n");
}

rl.close();

A streaming response makes the chatbot feel more natural. It feels like the bot is actually typing the answer (really fast) instead of the whole thing just showing up on the screen at once.

Next, I added a cache checkpoint to the System Prompt to see how caching actually affects the token cost.

Caching works by writing content to the cache once, at a cache checkpoint (in this case, right after the System Prompt). On the following requests, the model can read that content from the cache instead of processing it again as new input. Each cache checkpoint has a Time To Live (TTL): if no request reuses it within that window, it expires, and the next request has to write it again.

There are two TTL options available, 5 minutes and 1 hour, and I tested both.

// V4 - Streaming with Cache Point on System Prompt (5m)
...

while (true) {
  ...

  // Calling Bedrock with Converse API
  const response = await client.send(
    new ConverseStreamCommand({
      modelId: MODEL_ID,
      system: [
        { text: SYSTEM_PROMPT },
        { cachePoint: { type: "default", ttl: "5m" }},
      ],
      messages,
      inferenceConfig: { maxTokens: 512, temperature: 0.7 },
    })
  );

  ...
  let usage = {};

  for await (const event of response.stream) {
    ...
    if (event.metadata?.usage) {
      usage = event.metadata.usage;
    }
  }

  ...

  const output = usage.outputTokens ?? 0;
  const cacheRead = usage.cacheReadInputTokens ?? 0;
  const cacheWrite = usage.cacheWriteInputTokens ?? 0;
  const inputNoCache = usage.inputTokens ?? 0;
  const input = inputNoCache + cacheRead;

  console.log("\n");
  console.log("┌──────────────────────────────────────────────────┐");
  console.log("│                 📊  Token Usage                  │");
  console.log("├──────────────────────────────────────────────────┤");
  console.log(`│ Input total:        ${String(input).padStart(8)}  tokens             │`);
  console.log(`│   ├ Cache read:     ${String(cacheRead).padStart(8)}  (10% of price)     │`);
  console.log(`│   ├ Cache write:    ${String(cacheWrite).padStart(8)}  (125% of price)    │`);
  console.log(`│   └ No cache:       ${String(inputNoCache).padStart(8)}  (full price)       │`);
  console.log(`│ Output:             ${String(output).padStart(8)}  tokens             │`);
  console.log("└──────────────────────────────────────────────────┘");
}

The cache checkpoint goes right after the System Prompt in the system array, with a ttl field. For the 1-hour test, I used the same script and just changed the ttl to "1h".

With the 5-minute cache, the first request had nothing to read yet, so it wrote the System Prompt to the cache. That write costs 125% of the normal price.

On the next request, still inside the 5-minute window, the model read the System Prompt from the cache instead of writing it again. That cache read costs only 10% of the normal price.

The 1-hour cache write costs more, 200% of the normal price, but it lasts a full hour instead of just 5 minutes.

Just like with the 5-minute cache, reading from the 1-hour cache costs 10% of the normal price.

So the tradeoff is simple: a longer cache costs more to write, but it pays off when messages come in more than 5 minutes apart, since the 5-minute cache would have expired and been written again at full write price.

That's as far as I got in this class: the quota is sorted, the Playground is explored, and the ticket chatbot now streams its answers and reuses its System Prompt instead of paying full price for it on every message. More to come.

── more in #ai-products 4 stories · sorted by recency
── more on @amazon bedrock 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/from-playground-to-c…] indexed:0 read:11min 2026-09-28 · —