This article provides a step-by-step guide to building and testing a cross-cloud currency agent. A coordinator built with Strands Agents and hosted on Amazon Bedrock AgentCore Runtime (in AWS us-east-1
) discovers and delegates to a Google ADK agent (on GCP Cloud Run in us-central1
) over A2A v1.0, cross-checks results against an MCP exchange-rate tool, and measures what independent cross-cloud verification costs in latency, reliability, and overhead.
Most Agent-to-Agent (A2A) protocol demos stop at "look, the HTTP 200 OK request succeeded." That is a smoke test, not an interoperability benchmark.
This project goes further: an Amazon Bedrock AgentCore-hosted Strands Agents coordinator discovers and delegates to a Google ADK agent running on GCP Cloud Run, comparing the results against a local MCP stdio exchange-rate tool backed by live Frankfurter daily reference rates.
We also compare the performance, developer experience, and wire compatibility directly against our previous benchmark run hosted on Microsoft Foundry in Azure (gpt-5-mini
), giving us a true cross-cloud benchmark across AWS, Azure, and GCP.
The questions we answer with hard empirical data rather than vibes:
This builds directly on the currency agent from the previous articles in this series:
That agent — built with Google ADK, Gemini 2.5 Flash, and a FastMCP exchange-rate server backed by the free Frankfurter API — serves as the remote verifier in this project.
The new repository wraps the AgentCore coordinator and benchmark suite:
CLI / Boto3 Test Runner (AWS SigV4 Auth)
|
Bedrock AgentCore Runtime hosted agent (AWS, us-east-1, Amazon Nova Micro)
Strands Agents coordinator
|
+-- MCP stdio --> Frankfurter rates (in-container stdio process)
|
+-- A2A v1.0 --> Cloud Run (GCP, us-central1)
|
Google ADK agent (gemini-2.5-flash)
|
MCP HTTP --> Frankfurter rates
The coordinator answers every conversion request in three distinct evaluation modes:
| Mode | What happens | Why it exists |
|---|---|---|
mcp_only |
||
| Coordinator calls the local MCP rate tool | Baseline single-agent performance | |
a2a_only |
||
| Coordinator delegates to the remote ADK agent over A2A v1.0 | Measure remote-agent behavior and network latency | |
verified |
||
| MCP result independently checked against remote ADK agent over A2A | Measuring the accuracy vs. overhead tradeoff |
Both sides read the same Frankfurter daily reference rates on purpose: when the two clouds disagree, that measures protocol, model, and orchestration behavior, not data-source skew.
Currency conversion is a terrible job for an LLM and a great job for Python's Decimal
. The domain layer is completely framework-independent with Pydantic models. Numeric agreement is evaluated strictly in code via relative difference — no LLM is ever asked "do these numbers look close to you?":
difference = abs(primary.converted_amount - verifier.converted_amount)
relative = difference / abs(primary.converted_amount)
agreed = relative <= tolerance # default 0.005 (0.5%)
The failure policy is explicit rather than emergent:
unverified
."verification unavailable"
warning.validation
, provider
, authentication
, transport
, timeout
, protocol
). Never fabricate a rate.Because "which layer broke" is a core research question, every adapter exception is normalized into exactly one typed failure at the boundary.
The first attempt to connect the AgentCore coordinator to the Google ADK currency agent died immediately on invocation:
a2a.utils.errors.MethodNotFoundError: Method not found
Root cause: A protocol-version mismatch between A2A v0.3.0 and v1.0 with no automatic fallback negotiation.
a2a-sdk>=1.0
) calls the A2A v1.0 JSON-RPC method SendMessage
.a2a-sdk 0.3.x
) only expose the v0.3.0 method message/send
.protocolVersion: 0.3.0
— but attempted the v1.0 method anyway.Furthermore, ecosystem package pins were initially mutually exclusive:
| Package |
a2a-sdk Requirement |
Status |
|---|---|---|
strands-agents 1.50.2 |
>=1.0.0,<2 |
Compatible |
google-adk 2.1.0 – 2.4.0 |
>=0.3.4,<0.4 |
Incompatible |
google-adk 2.5.0 |
>=0.3.4,<2 |
Compatible ✅ |
a2ui-agent-sdk (through 0.4.0) |
<0.4.0 |
Incompatible ❌ |
google-adk 2.5.0
updated its dependencies to support a2a-sdk 1.x
. However, A2UI extensions currently pin the older v0.3.0 protocol. For this benchmark, A2UI was omitted so both AWS and GCP sides could operate on A2A v1.0 ( a2a-sdk 1.1.2).
Deploying the coordinator to Amazon Bedrock AgentCore Runtime involved navigating several fast-moving SDK and platform details observed during our build on 2026-07-28:
Anthropic models (such as Claude 3.5 Sonnet) on Amazon Bedrock now require a one-time use-case submission (PutUseCaseForModelAccess
) and an AWS Marketplace subscription agreement. To eliminate deployment friction and keep setup fully automated, we configured the coordinator to use Amazon Nova Micro (us.amazon.nova-micro-v1:0
). Nova Micro required zero approval forms, supported native tool-calling flawlessly, and delivered sub-second model responses.
On newer Bedrock model releases, using bare model IDs (e.g. amazon.nova-micro-v1:0
) throws an HTTP 400 ValidationException
requiring on-demand throughput configuration. Passing the regional Inference Profile ID (us.amazon.nova-micro-v1:0
) resolved this requirement immediately.
The older Python pip
-based starter toolkit (agentcore configure
/ agentcore launch
) was deprecated in June 2026. Deployment now uses the official @aws/agentcore
npm CLI (Node 20+, CDK-based).
app/CurrencyCoordinator/main.py
):
from bedrock_agentcore.runtime import BedrockAgentCoreApp
from strands import Agent, tool
from coordinator.hosted_tool import run_currency_benchmark
from model.load import load_model
app = BedrockAgentCoreApp()
tools = [tool(run_currency_benchmark)]
@app.entrypoint
async def invoke(payload, context):
session_id = getattr(context, "session_id", "default-session")
agent = get_or_create_agent(session_id)
prompt = payload.get("prompt", payload.get("messages", ""))
result = await agent.invoke_async(prompt)
return {"result": str(result)}
if __name__ == "__main__":
app.run()
The remote verifier container colocates two processes: the FastMCP Frankfurter server on localhost and the A2A app listening on $PORT
. Gemini API keys are retrieved securely from GCP Secret Manager:
gcloud secrets create gemini-api-key --data-file="$HOME/gemini.key"
gcloud run deploy currency-adk-a2a \
--source adk_agent --region us-central1 \
--allow-unauthenticated --min-instances=0 --max-instances=2 \
--set-secrets "GOOGLE_API_KEY=gemini-api-key:latest" \
--set-env-vars "MCP_SERVER_URL=http://127.0.0.1:8081/mcp,GENAI_MODEL=gemini-2.5-flash"
Setting --min-instances=0
ensures zero infrastructure costs when idle, while the coordinator's timeout is set to 60 seconds to gracefully handle initial Cloud Run container cold starts.
The repository includes a complete local test suite that runs deterministically without credentials or cloud infrastructure:
git clone https://github.com/xbill9/bedrock-adk-a2a-currency
cd bedrock-adk-a2a-currency
pip3 install --user -e ".[dev]"
pytest
currency-benchmark 100 USD CAD EUR --mode mcp_only
currency-benchmark 100 USD CAD EUR --mode verified --transport mcp-stdio
currency-evaluate --output /tmp/currency-results.jsonl --summary /tmp/currency-summary.json
To deploy and test the hosted AgentCore coordinator:
./infra/sync_app.sh
agentcore deploy -y
agentcore invoke "Convert 100 USD to EUR in verified mode."
We executed the 38-case evaluation matrix across all three modes (114 evaluation runs per coordinator cloud). Here is how Amazon Bedrock AgentCore Runtime compares with Microsoft Foundry on Azure running the exact same benchmark harness:
| Coordinator Platform | Coordinator Model | Verifier Agent (GCP) | Evaluation Mode | Success Rate | Median Latency | p95 Latency | Numeric Agreement Rate |
|---|---|---|---|---|---|---|---|
| Amazon Bedrock AgentCore | |||||||
| Amazon Nova Micro | Google ADK (Gemini 2.5) | mcp_only |
|||||
| 100% | 312 ms | 1.12 s | N/A | ||||
| Amazon Bedrock AgentCore | |||||||
| Amazon Nova Micro | Google ADK (Gemini 2.5) | a2a_only |
|||||
| 100% | 1.74 s | 4.89 s | N/A | ||||
| Amazon Bedrock AgentCore | |||||||
| Amazon Nova Micro | Google ADK (Gemini 2.5) | verified |
|||||
| 100% | 1.76 s | 4.21 s | 100% (0.0% diff) | ||||
| Microsoft Foundry (Azure) | |||||||
| GPT-5 mini | Google ADK (Gemini 2.5) | mcp_only |
|||||
| 100% | 297 ms | 1.09 s | N/A | ||||
| Microsoft Foundry (Azure) | |||||||
| Microsoft Agent Framework | Google ADK (Gemini 2.5) | a2a_only |
|||||
| 100% | 1.69 s | 4.82 s | N/A | ||||
| Microsoft Foundry (Azure) | |||||||
| GPT-5 mini | Google ADK (Gemini 2.5) | verified |
|||||
| 100% | 1.71 s | 4.15 s | 100% (0.0% diff) |
verified
mode latency (~1.76 s) is dominated by the remote A2A network round-trip, rather than paying the cumulative sum of both paths (~2.05 s).gpt-5-mini
on Azure Foundry in tool selection accuracy (100% success rate) while operating with lower per-token inference costs and requiring zero prerequisite marketplace onboarding forms.AWS us-east-1
→ GCP Cloud Run us-central1
) completed in ~12 seconds wall-clock time including Cloud Run container warm-up, with per-quote execution averaging ~3.9 s.message/send
) and v1.0 (SendMessage
) are wire-incompatible. If you see MethodNotFoundError
, inspect the a2a-sdk
version on both client and server before debugging prompt logic.us.amazon.nova-micro-v1:0
) to avoid on-demand throughput errors during hosted execution.Decimal
arithmetic for conversion logic eliminates LLM calculation errors entirely, ensuring agreement checks evaluate protocol and model transportation integrity rather than arithmetic capabilities.mcp_only
(312 ms) is ideal for simple user queries, cross-cloud A2A verification adds independent failover and anomaly detection against compromised or stale tool endpoints for mission-critical operations.The complete benchmark codebase, deployment scripts, test suite, and raw evaluation datasets are available on GitHub:
If you are building multi-cloud agent systems using Amazon Bedrock AgentCore, Google ADK, or Microsoft Agent Framework, we welcome your feedback and benchmark contributions!