ai agent python code example for FastAPI and OpenAI SDK A developer with over a year of production AI agent experience published a FastAPI and OpenAI SDK implementation pattern covering the core request loop, function-calling tool use, and Redis-backed state for multi-turn memory. The writeup emphasizes letting the OpenAI SDK manage the tool call/response cycle rather than hand-rolling state machines, and includes explicit handling for rate limits and tool timeouts. I’ve been running AI agents in production for over a year now. Most tutorials skip the hard parts - state, rate limits, and what happens when your agent calls a tool that times out. Here’s how I build them with FastAPI and the OpenAI SDK, the way I’d want it documented when I’m paged at 2 a.m. Start with a minimal FastAPI app that accepts a prompt, calls OpenAI, and returns a response. No tools, no state - just the core loop. This is your foundation. python main.py from fastapi import FastAPI, HTTPException from pydantic import BaseModel import openai import os app = FastAPI client = openai.OpenAI api key=os.getenv "OPENAI API KEY" class AgentRequest BaseModel : prompt: str model: str = "gpt-4o" class AgentResponse BaseModel : response: str @app.post "/agent", response model=AgentResponse async def run agent request: AgentRequest : try: completion = client.chat.completions.create model=request.model, messages= {"role": "user", "content": request.prompt} return AgentResponse response=completion.choices 0 .message.content except openai.RateLimitError: raise HTTPException status code=429, detail="Rate limit exceeded" except Exception as e: raise HTTPException status code=500, detail=str e This works for simple Q&A. But real agents need to do more than chat - they need to act. Tool use turns your agent from a talker into a doer. I define functions the agent can call - like querying a database or hitting an internal API - and let the OpenAI SDK handle the function calling loop. python tools.py from typing import Optional import httpx import os async def get user info user id: str - dict: async with httpx.AsyncClient as client: resp = await client.get f"{os.getenv 'INTERNAL API URL' }/users/{user id}" resp.raise for status return resp.json agent tools.py from openai import OpenAI from typing import List, Dict, Any import json client = OpenAI api key=os.getenv "OPENAI API KEY" TOOLS = { "type": "function", "function": { "name": "get user info", "description": "Fetch user details by ID", "parameters": { "type": "object", "properties": { "user id": {"type": "string", "description": "The user ID"} }, "required": "user id" } } } async def run agent with tools prompt: str, model: str = "gpt-4o" - str: messages = {"role": "user", "content": prompt} while True: completion = client.chat.completions.create model=model, messages=messages, tools=TOOLS, tool choice="auto" message = completion.choices 0 .message if message.tool calls: Add the assistant's message with tool calls messages.append message for tool call in message.tool calls: if tool call.function.name == "get user info": args = json.loads tool call.function.arguments result = await get user info args "user id" messages.append { "tool call id": tool call.id, "role": "tool", "content": json.dumps result } Continue the loop to let the agent process the tool result continue else: Final response return message.content I’ve seen teams try to manage this loop manually with state machines. Don’t. Let the SDK handle the tool call/response cycle - it’s less buggy and easier to debug. Production agents aren’t stateless. They need memory across turns - chat history, user preferences, cached tool results. I use Redis for fast state and Pydantic models to keep it typed. python state.py import redis.asyncio as redis from pydantic import BaseModel from typing import Optional, Dict, Any import json redis client = redis.from url os.getenv "REDIS URL" , decode responses=True class AgentState BaseModel : session id: str chat history: List Dict str, str = user preferences: Dict str, Any = {} last tool result: Optional Dict = None async def get state session id: str - AgentState: data = await redis client.get f"agent state:{session id}" if data: return AgentState.model validate json data return AgentState session id=session id async def save state state: AgentState : await redis client.set f"agent state:{state.session id}", state.model dump json , ex=3600 1 hour TTL Each request loads state, runs the agent loop which may mutate state via tools , then saves it back. I’ve had agents corrupt state by writing concurrently - use Redis transactions or a queue if you need strong consistency. Containerize it. Cloud Run scales to zero, which saves money when traffic is spiky. But cold starts hurt - keep your image small and avoid heavy imports at module level. Dockerfile FROM python:3.11-slim WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY . . ENV PORT=8080 EXPOSE 8080 CMD "uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080" I deploy with: gcloud builds submit --tag gcr.io/my-project/ai-agent gcloud run deploy ai-agent --image gcr.io/my-project/ai-agent --platform managed Watch for: --secret in Cloud Build or Cloud Run env vars. Rate limits are the most frequent pager. I’ve seen agents get stuck in retry loops because they didn’t back off. Token limits sneak up when chat history grows. Context overflow makes agents forget instructions. Here’s how I defend against them: python middleware.py from fastapi import Request, Response from starlette.middleware.base import BaseHTTPMiddleware import time from collections import defaultdict request counts = defaultdict list class RateLimitMiddleware BaseHTTPMiddleware : async def dispatch self, request: Request, call next : client ip = request.client.host now = time.time Clean old requests request counts client ip = t for t in request counts client ip if now - t < 60 if len request counts client ip = 10: 10 req/min return Response "Rate limit exceeded", status code=429 request counts client ip .append now return await call next request For token limits, I truncate history: python def truncate history messages, max tokens=3000 : Rough estimate: 4 chars per token total chars = sum len m "content" for m in messages if total chars max tokens 4: Keep system message and recent turns return messages 0 + messages - max tokens//4 : Simplified return messages I log token usage per request: In your agent endpoint usage = completion.usage logger.info f"Token usage: {usage.total tokens} prompt: {usage.prompt tokens}, completion: {usage.completion tokens} " If you see completion tokens near the limit, your agent is looping or over-explaining. I track three things: success rate, latency, and tool accuracy. Success rate means the agent completed the user’s goal without human intervention. I log every turn and use a simple evaluator LLM to judge outcomes. python metrics.py from prometheus client import Counter, Histogram import time AGENT REQUESTS = Counter 'agent requests total', 'Total agent requests', 'status' AGENT LATENCY = Histogram 'agent latency seconds', 'Time spent processing agent requests' TOOL USAGE = Counter 'agent tool usage total', 'Tool usage count', 'tool name' async def run agent evaluated prompt: str, session id: str : start = time.time state = await get state session id try: response = await run agent with tools prompt Simple success heuristic: no error and tool used if needed success = "error" not in response.lower AGENT REQUESTS.labels status="success" if success else "failure" .inc Log for human review logger.info { "session id": session id, "prompt": prompt, "response": response, "state": state.model dump , "success": success } return response finally: AGENT LATENCY.observe time.time - start I’ve used this to catch agents that call tools unnecessarily - wasting money and latency. If your tool usage counter is high but success rate low, your agent is confused about when to act. A FastAPI endpoint that calls OpenAI chat.completions.create with a user prompt and returns the message content. Add error handling for rate limits and you’ve got a baseline. Limit the number of tool call iterations per request e.g., max 5 turns . The OpenAI SDK doesn’t do this by default - wrap your agent loop in a counter and break if exceeded. For production agents needing fine-grained control over tool calls, state, and logging, use the OpenAI SDK directly. LangChain adds abstraction that obscures failure modes and makes debugging harder. At $0.005 per 1K tokens for GPT-4o, a typical agent turn 500 prompt + 150 completion tokens costs ~$0.003. Tool calls add latency but no extra token cost - watch for iteration bloat.