Your API's newest users are agents. They don't read docs, they don't browse dashboards, and they don't file support tickets. They parse OpenAPI specs, call endpoints in loops, and expect deterministic, machine-readable responses. If your API was designed for humans clicking buttons, it's already failing this new class of client.
In this post, I'll walk through a concrete example: building a small "tool API" that an LLM-based agent can call. We'll cover the problem, a solution, and a runnable Python implementation. No hand-waving about "agentic workflows" — just code you can run and a loop with explicit termination conditions.
Human-facing APIs optimize for discoverability and forgiveness. Agents optimize for determinism and low token cost. Three specific failure modes show up when agents hit a human-designed API:
{"error": "bad request"} forces the agent to guess. The agent will retry, hallucinate a fix, or give up.
Here's a minimal example of the kind of handler that causes these problems:
from flask import Flask, request, jsonify
app = Flask(__name__)
@app.route("/create_ticket", methods=["POST"])
def create_ticket_bad():
data = request.get_json(silent=True) or {}
if "title" not in data:
return jsonify({"error": "bad request"}), 400
return jsonify({"ticket": {"id": 1, "title": data["title"], "history": [...]}})
An agent calling this has no way to distinguish "missing field" from "malformed JSON" from "server error," and no safe retry path.
Design the API for a client that is literal, stateless between calls, and token-budgeted. Concretely:
code string the agent can branch on.Idempotency-Key header; store the result keyed by it.
We'll build a small in-memory ticket API with the properties above, then write an agent loop that uses it. The agent loop is deliberately simple: it's a plan-act-observe loop with a hard step limit and a success predicate. No framework required.
import json
import uuid
from dataclasses import dataclass, field, asdict
from typing import Optional
from flask import Flask, request, jsonify
app = Flask(__name__)
@dataclass
class Ticket:
id: str
title: str
status: str = "open"
tickets: dict[str, Ticket] = {}
idempotency_store: dict[str, dict] = {}
TOOL_SCHEMA = {
"name": "create_ticket",
"description": "Create a support ticket. Idempotent on Idempotency-Key header.",
"parameters": {
"type": "object",
"properties": {
"title": {"type": "string", "minLength": 1, "maxLength": 200},
},
"required": ["title"],
"additionalProperties": False,
},
}
@app.route("/tools", methods=["GET"])
def list_tools():
return jsonify({"tools": [TOOL_SCHEMA]})
@app.route("/create_ticket", methods=["POST"])
def create_ticket():
key = request.headers.get("Idempotency-Key")
if not key:
return jsonify({"code": "missing_idempotency_key",
"message": "Provide Idempotency-Key header."}), 400
if key in idempotency_store:
return jsonify(idempotency_store[key]), 200
data = request.get_json(silent=True)
if not isinstance(data, dict):
return jsonify({"code": "invalid_json",
"message": "Body must be a JSON object."}), 400
title = data.get("title")
if not isinstance(title, str) or not title.strip():
return jsonify({"code": "invalid_title",
"message": "'title' must be a non-empty string."}), 422
ticket = Ticket(id=str(uuid.uuid4()), title=title.strip())
tickets[ticket.id] = ticket
payload = {"ticket": asdict(ticket)}
idempotency_store[key] = payload
return jsonify(payload), 201
@app.route("/tickets/<ticket_id>", methods=["GET"])
def get_ticket(ticket_id: str):
ticket = tickets.get(ticket_id)
if ticket is None:
return jsonify({"code": "not_found",
"message": f"No ticket {ticket_id}."}), 404
return jsonify({"ticket": asdict(ticket)})
Key details: every error has a stable code; the response body is small; idempotency is enforced via header, not body, so retries are safe.
The agent loop below is deliberately framework-free. It defines explicit termination conditions: it stops when the goal is met, when the model returns no tool call, or when max_steps is reached. I'm using an OpenAI-chat-completions-style interface here as a stand-in; swap in whichever client you use.
import json
import uuid
from typing import Any
import requests
API = "http://localhost:5000"
MAX_STEPS = 6
def call_model(messages: list[dict]) -> dict:
"""Return a message dict. Replace with your model client."""
raise NotImplementedError("Wire up your model client.")
def call_tool(name: str, args: dict) -> dict:
if name == "create_ticket":
r = requests.post(
f"{API}/create_ticket",
json=args,
headers={"Idempotency-Key": str(uuid.uuid4())},
timeout=10,
)
return {"status": r.status_code, "body": r.json()}
return {"status": 400, "body": {"code": "unknown_tool", "message": name}}
def run_agent(goal: str) -> dict:
tools = requests.get(f"{API}/tools", timeout=10).json()["tools"]
messages = [
{"role": "system", "content": "You are an agent. Use tools to satisfy the goal."},
{"role": "user", "content": goal},
]
for step in range(MAX_STEPS):
msg = call_model(messages)
messages.append(msg)
if not msg.get("tool_calls"):
return {"status": "done", "steps": step, "answer": msg.get("content")}
for call in msg["tool_calls"]:
name = call["function"]["name"]
args = json.loads(call["function"]["arguments"])
result = call_tool(name, args)
messages.append({
"role": "tool",
"tool_call_id": call["id"],
"content": json.dumps(result),
})
return {"status": "max_steps_reached", "steps": MAX_STEPS}
The loop has exactly two exit paths, both explicit. There's no "keep trying until it works" branch, which is where most agent bugs live.
If you extend this to execute code or shell commands, treat the tool boundary as a trust boundary. Never pass model output to eval, exec, or subprocess with shell=True. If you must run generated code, isolate it in a sandbox (a container with no network, a read-only filesystem, and a hard timeout) and validate arguments against the JSON Schema before execution. The example above avoids this entirely by only allowing a single typed tool with a bounded string parameter.
eval on model output is a remote code execution vulnerability with extra steps.
The code above is intentionally minimal so you can fork it and add your own tools. Start by auditing your existing API for the three failure modes in the Problem section — that's where the real work is.