{"slug": "ai-agent-python-code-example-for-fastapi-and-openai-sdk", "title": "ai agent python code example for FastAPI and OpenAI SDK", "summary": "A developer with over a year of production AI agent experience published a FastAPI and OpenAI SDK implementation pattern covering the core request loop, function-calling tool use, and Redis-backed state for multi-turn memory. The writeup emphasizes letting the OpenAI SDK manage the tool call/response cycle rather than hand-rolling state machines, and includes explicit handling for rate limits and tool timeouts.", "body_md": "I’ve been running AI agents in production for over a year now. Most tutorials skip the hard parts - state, rate limits, and what happens when your agent calls a tool that times out. Here’s how I build them with FastAPI and the OpenAI SDK, the way I’d want it documented when I’m paged at 2 a.m.\n\nStart with a minimal FastAPI app that accepts a prompt, calls OpenAI, and returns a response. No tools, no state - just the core loop. This is your foundation.\n\n``` python\n# main.py\nfrom fastapi import FastAPI, HTTPException\nfrom pydantic import BaseModel\nimport openai\nimport os\n\napp = FastAPI()\nclient = openai.OpenAI(api_key=os.getenv(\"OPENAI_API_KEY\"))\n\nclass AgentRequest(BaseModel):\n    prompt: str\n    model: str = \"gpt-4o\"\n\nclass AgentResponse(BaseModel):\n    response: str\n\n@app.post(\"/agent\", response_model=AgentResponse)\nasync def run_agent(request: AgentRequest):\n    try:\n        completion = client.chat.completions.create(\n            model=request.model,\n            messages=[{\"role\": \"user\", \"content\": request.prompt}]\n        )\n        return AgentResponse(response=completion.choices[0].message.content)\n    except openai.RateLimitError:\n        raise HTTPException(status_code=429, detail=\"Rate limit exceeded\")\n    except Exception as e:\n        raise HTTPException(status_code=500, detail=str(e))\n```\n\nThis works for simple Q&A. But real agents need to do more than chat - they need to act.\n\nTool use turns your agent from a talker into a doer. I define functions the agent can call - like querying a database or hitting an internal API - and let the OpenAI SDK handle the function calling loop.\n\n``` python\n# tools.py\nfrom typing import Optional\nimport httpx\nimport os\n\nasync def get_user_info(user_id: str) -> dict:\n    async with httpx.AsyncClient() as client:\n        resp = await client.get(f\"{os.getenv('INTERNAL_API_URL')}/users/{user_id}\")\n        resp.raise_for_status()\n        return resp.json()\n\n# agent_tools.py\nfrom openai import OpenAI\nfrom typing import List, Dict, Any\nimport json\n\nclient = OpenAI(api_key=os.getenv(\"OPENAI_API_KEY\"))\n\nTOOLS = [\n    {\n        \"type\": \"function\",\n        \"function\": {\n            \"name\": \"get_user_info\",\n            \"description\": \"Fetch user details by ID\",\n            \"parameters\": {\n                \"type\": \"object\",\n                \"properties\": {\n                    \"user_id\": {\"type\": \"string\", \"description\": \"The user ID\"}\n                },\n                \"required\": [\"user_id\"]\n            }\n        }\n    }\n]\n\nasync def run_agent_with_tools(prompt: str, model: str = \"gpt-4o\") -> str:\n    messages = [{\"role\": \"user\", \"content\": prompt}]\n\n    while True:\n        completion = client.chat.completions.create(\n            model=model,\n            messages=messages,\n            tools=TOOLS,\n            tool_choice=\"auto\"\n        )\n\n        message = completion.choices[0].message\n\n        if message.tool_calls:\n            # Add the assistant's message with tool calls\n            messages.append(message)\n\n            for tool_call in message.tool_calls:\n                if tool_call.function.name == \"get_user_info\":\n                    args = json.loads(tool_call.function.arguments)\n                    result = await get_user_info(args[\"user_id\"])\n                    messages.append({\n                        \"tool_call_id\": tool_call.id,\n                        \"role\": \"tool\",\n                        \"content\": json.dumps(result)\n                    })\n            # Continue the loop to let the agent process the tool result\n            continue\n        else:\n            # Final response\n            return message.content\n```\n\nI’ve seen teams try to manage this loop manually with state machines. Don’t. Let the SDK handle the tool call/response cycle - it’s less buggy and easier to debug.\n\nProduction agents aren’t stateless. They need memory across turns - chat history, user preferences, cached tool results. I use Redis for fast state and Pydantic models to keep it typed.\n\n``` python\n# state.py\nimport redis.asyncio as redis\nfrom pydantic import BaseModel\nfrom typing import Optional, Dict, Any\nimport json\n\nredis_client = redis.from_url(os.getenv(\"REDIS_URL\"), decode_responses=True)\n\nclass AgentState(BaseModel):\n    session_id: str\n    chat_history: List[Dict[str, str]] = []\n    user_preferences: Dict[str, Any] = {}\n    last_tool_result: Optional[Dict] = None\n\nasync def get_state(session_id: str) -> AgentState:\n    data = await redis_client.get(f\"agent_state:{session_id}\")\n    if data:\n        return AgentState.model_validate_json(data)\n    return AgentState(session_id=session_id)\n\nasync def save_state(state: AgentState):\n    await redis_client.set(\n        f\"agent_state:{state.session_id}\",\n        state.model_dump_json(),\n        ex=3600  # 1 hour TTL\n    )\n```\n\nEach request loads state, runs the agent loop (which may mutate state via tools), then saves it back. I’ve had agents corrupt state by writing concurrently - use Redis transactions or a queue if you need strong consistency.\n\nContainerize it. Cloud Run scales to zero, which saves money when traffic is spiky. But cold starts hurt - keep your image small and avoid heavy imports at module level.\n\n```\n# Dockerfile\nFROM python:3.11-slim\n\nWORKDIR /app\nCOPY requirements.txt .\nRUN pip install --no-cache-dir -r requirements.txt\n\nCOPY . .\n\nENV PORT=8080\nEXPOSE 8080\n\nCMD [\"uvicorn\", \"main:app\", \"--host\", \"0.0.0.0\", \"--port\", \"8080\"]\n```\n\nI deploy with:\n\n```\ngcloud builds submit --tag gcr.io/my-project/ai-agent\ngcloud run deploy ai-agent --image gcr.io/my-project/ai-agent --platform managed\n```\n\nWatch for:\n\n`--secret` in Cloud Build or Cloud Run env vars.\nRate limits are the most frequent pager. I’ve seen agents get stuck in retry loops because they didn’t back off. Token limits sneak up when chat history grows. Context overflow makes agents forget instructions.\n\nHere’s how I defend against them:\n\n``` python\n# middleware.py\nfrom fastapi import Request, Response\nfrom starlette.middleware.base import BaseHTTPMiddleware\nimport time\nfrom collections import defaultdict\n\nrequest_counts = defaultdict(list)\n\nclass RateLimitMiddleware(BaseHTTPMiddleware):\n    async def dispatch(self, request: Request, call_next):\n        client_ip = request.client.host\n        now = time.time()\n\n        # Clean old requests\n        request_counts[client_ip] = [t for t in request_counts[client_ip] if now - t < 60]\n\n        if len(request_counts[client_ip]) >= 10:  # 10 req/min\n            return Response(\"Rate limit exceeded\", status_code=429)\n\n        request_counts[client_ip].append(now)\n        return await call_next(request)\n```\n\nFor token limits, I truncate history:\n\n``` python\ndef truncate_history(messages, max_tokens=3000):\n    # Rough estimate: 4 chars per token\n    total_chars = sum(len(m[\"content\"]) for m in messages)\n    if total_chars > max_tokens * 4:\n        # Keep system message and recent turns\n        return [messages[0]] + messages[-(max_tokens//4):]  # Simplified\n    return messages\n```\n\nI log token usage per request:\n\n```\n# In your agent endpoint\nusage = completion.usage\nlogger.info(f\"Token usage: {usage.total_tokens} (prompt: {usage.prompt_tokens}, completion: {usage.completion_tokens})\")\n```\n\nIf you see completion tokens near the limit, your agent is looping or over-explaining.\n\nI track three things: success rate, latency, and tool accuracy. Success rate means the agent completed the user’s goal without human intervention. I log every turn and use a simple evaluator LLM to judge outcomes.\n\n``` python\n# metrics.py\nfrom prometheus_client import Counter, Histogram\nimport time\n\nAGENT_REQUESTS = Counter('agent_requests_total', 'Total agent requests', ['status'])\nAGENT_LATENCY = Histogram('agent_latency_seconds', 'Time spent processing agent requests')\nTOOL_USAGE = Counter('agent_tool_usage_total', 'Tool usage count', ['tool_name'])\n\nasync def run_agent_evaluated(prompt: str, session_id: str):\n    start = time.time()\n    state = await get_state(session_id)\n\n    try:\n        response = await run_agent_with_tools(prompt)\n        # Simple success heuristic: no error and tool used if needed\n        success = \"error\" not in response.lower()\n        AGENT_REQUESTS.labels(status=\"success\" if success else \"failure\").inc()\n\n        # Log for human review\n        logger.info({\n            \"session_id\": session_id,\n            \"prompt\": prompt,\n            \"response\": response,\n            \"state\": state.model_dump(),\n            \"success\": success\n        })\n\n        return response\n    finally:\n        AGENT_LATENCY.observe(time.time() - start)\n```\n\nI’ve used this to catch agents that call tools unnecessarily - wasting money and latency. If your tool usage counter is high but success rate low, your agent is confused about when to act.\n\nA FastAPI endpoint that calls OpenAI chat.completions.create with a user prompt and returns the message content. Add error handling for rate limits and you’ve got a baseline.\n\nLimit the number of tool call iterations per request (e.g., max 5 turns). The OpenAI SDK doesn’t do this by default - wrap your agent loop in a counter and break if exceeded.\n\nFor production agents needing fine-grained control over tool calls, state, and logging, use the OpenAI SDK directly. LangChain adds abstraction that obscures failure modes and makes debugging harder.\n\nAt $0.005 per 1K tokens for GPT-4o, a typical agent turn (500 prompt + 150 completion tokens) costs ~$0.003. Tool calls add latency but no extra token cost - watch for iteration bloat.", "url": "https://wpnews.pro/news/ai-agent-python-code-example-for-fastapi-and-openai-sdk", "canonical_source": "https://dev.to/ayush_kumar_085a0f2c54e3f/ai-agent-python-code-example-for-fastapi-and-openai-sdk-22jg", "published_at": "2026-09-30 10:13:30+00:00", "updated_at": "2026-09-30 10:17:26.009951+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models", "ai-infrastructure"], "entities": ["FastAPI", "OpenAI", "Redis", "Pydantic", "httpx", "GPT-4o"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agent-python-code-example-for-fastapi-and-openai-sdk", "markdown": "https://wpnews.pro/news/ai-agent-python-code-example-for-fastapi-and-openai-sdk.md", "text": "https://wpnews.pro/news/ai-agent-python-code-example-for-fastapi-and-openai-sdk.txt", "jsonld": "https://wpnews.pro/news/ai-agent-python-code-example-for-fastapi-and-openai-sdk.jsonld"}}