How (and Why) to Build an AI Agent from Scratch in Python A tutorial published with companion code in the GitHub repository balapriyac/ai-agent-from-scratch shows developers how to build an AI agent in plain Python using the Anthropic API and the Anthropic Python SDK, without an orchestration framework. The guide requires Python 3.10 or later and an ANTHROPIC_API_KEY environment variable, and walks through tool calling, the request-execute-respond loop, and persistent memory across turns. It warns that Anthropic will retire Claude Sonnet 4.5 on November 30, 2026, so the model string "claude-sonnet-4-5" used in the examples will stop working after that date. In this article, you will learn what an AI agent is and how to build one from scratch in plain Python using the Anthropic API, without relying on any orchestration framework. Topics we will cover include: - Why tool calling is necessary and how to define a tool the model can use. - How to implement the request-execute-respond loop that drives an agent’s behavior. - How to add persistent memory so the agent maintains context across multiple turns. Introduction A simple call to a large language model https://www.ibm.com/think/topics/large-language-models is enough when the answer can be produced from what the model already knows: explaining a policy, drafting a reply, summarizing text. It is not enough once the answer depends on data outside that training set, such as a live order status or a row in your database; this gap is what an agent is built to close. An AI agent https://cloud.google.com/discover/what-are-ai-agents is what results once the model can call out to something else — a function, a database query, an API — and use the result before it answers. Reaching for an agent orchestration framework https://www.langchain.com/resources/ai-agent-frameworks to get this working is a reasonable move, eventually. Writing one from scratch in plain Python, using a raw API call, makes the underlying architecture clear. At its core, there is simply a model, a loop that manages the interaction, and a small set of well-defined functions that give the model the capabilities it needs. Stripping away the frameworks and abstractions makes it easier to see how these pieces fit together and what is actually happening under the hood. This article covers: - Why tool calling is necessary and what you need installed to try it - How to describe a tool so the model knows when and how to call it - What the request-execute-respond loop looks like in code - How to add memory so the agent keeps context across turns You can find the companion code for this article in this GitHub repository https://github.com/balapriyac/ai-agent-from-scratch . Prerequisites You need Python 3.10 https://www.python.org/downloads/release/python-3100/ or later, an Anthropic API key, and the Anthropic Python SDK: pip install anthropic 1 pip install anthropic Set your key as an environment variable so the client picks it up without hardcoding it anywhere: export ANTHROPIC API KEY="your-key-here" 1 export ANTHROPIC API KEY="your-key-here" This is all you need to get started. You can find all the code in the agent.py https://github.com/balapriyac/ai-agent-from-scratch/blob/main/agent.py script. ⚠️ A note to readers : Anthropic will officially retire Claude Sonnet 4.5 on November 30, 2026 https://platform.claude.com/docs/en/about-claude/model-deprecations 2026-09-30-claude-sonnet-4-5-model . If you are following this tutorial after that date, the specific model code used in these examples will no longer work. To ensure your project runs smoothly, replace the model string in the API calls with a newer version or the latest available model. Setting Up the Model Call A minimal wrapper around the API is a function that sends a prompt and returns text: python import anthropic client = anthropic.Anthropic def ask prompt : response = client.messages.create model="claude-sonnet-4-5", max tokens=512, system="You are a helpful support assistant. Be direct and factual.", messages= {"role": "user", "content": prompt} , return response.content 0 .text 123456789101112 import anthropic client = anthropic.Anthropic def ask prompt : response = client.messages.create model="claude-sonnet-4-5", max tokens=512, system="You are a helpful support assistant. Be direct and factual.", messages= {"role": "user", "content": prompt} , return response.content 0 .text This handles a large share of questions well — explaining a policy, drafting a reply, summarizing a paragraph. It fails the moment the answer depends on data the model was never shown. Ask ask "What is the status of order 4471?" and the model has no mechanism to check: that information lives in your database, not in its training data. It either states that it does not know, or produces a plausible answer anyway, and no amount of prompting changes that, since no prompt grants the model access to data it was never given. Tool calling https://www.ibm.com/think/topics/tool-calling gives the model a defined way to request that data instead of inferring it. Read The Roadmap to Mastering Tool Calling in AI Agents https://machinelearningmastery.com/the-roadmap-to-mastering-tool-calling-in-ai-agents/ to learn more. Defining a Tool the Model Can Ask For Giving the model access to a function means writing the function, and writing a description of it that the model can read. The description matters because it tells the model what the function does and when to reach for it. orders db = { "4471": {"status": "shipped", "carrier": "UPS", "eta": "2 days"}, "4472": {"status": "processing", "carrier": None, "eta": None}, } def get order status order id : return orders db.get order id, {"error": "No order found with that ID"} get order status schema = { "name": "get order status", "description": "Looks up the current status of a customer order by its ID. " "Use this any time a question depends on current order data " "rather than general policy information." , "input schema": { "type": "object", "properties": { "order id": {"type": "string", "description": "The order ID to look up, e.g. '4471'"}, }, "required": "order id" , }, } 1234567891011121314151617181920212223 orders db = { "4471": {"status": "shipped", "carrier": "UPS", "eta": "2 days"}, "4472": {"status": "processing", "carrier": None, "eta": None},} def get order status order id : return orders db.get order id, {"error": "No order found with that ID"} get order status schema = { "name": "get order status", "description": "Looks up the current status of a customer order by its ID. " "Use this any time a question depends on current order data " "rather than general policy information." , "input schema": { "type": "object", "properties": { "order id": {"type": "string", "description": "The order ID to look up, e.g. '4471'"}, }, "required": "order id" , },} The schema is structured metadata: a name, a description, and a definition of the parameters the function expects. Tool use https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview works the same way across most providers: you pass the model a list of these schemas alongside your message, and the model decides on its own whether answering the question requires calling one of them. Executing the Tool and Feeding the Result Back Pass the schema in, and something different happens. Instead of a plain text answer, the response comes back with a stop reason of tool use and a content block describing which function the model wants to call and with what arguments. The model has not run anything yet — it is paused, and has returned a request for you to act on. messages = {"role": "user", "content": "What is the status of order 4471?"} response = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools= get order status schema , messages=messages, if response.stop reason == "tool use": tool call = next block for block in response.content if block.type == "tool use" result = get order status tool call.input messages.append {"role": "assistant", "content": response.content} messages.append { "role": "user", "content": { "type": "tool result", "tool use id": tool call.id, "content": str result , } , } final = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools= get order status schema , messages=messages, print final.content 0 .text 123456789101112131415161718192021222324252627282930 messages = {"role": "user", "content": "What is the status of order 4471?"} response = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools= get order status schema , messages=messages, if response.stop reason == "tool use": tool call = next block for block in response.content if block.type == "tool use" result = get order status tool call.input messages.append {"role": "assistant", "content": response.content} messages.append { "role": "user", "content": { "type": "tool result", "tool use id": tool call.id, "content": str result , } , } final = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools= get order status schema , messages=messages, print final.content 0 .text Two round trips happen here. The first asks the model what it wants to do. You run the function yourself — the lookup against orders db , or in a production system, a call to your order service — append the result back into the conversation as a tool result , and send everything again. The second round trip is where the model reads that result and gives you an answer grounded in it, instead of a guess. Looping Until the Model Has What It Needs This pattern generalizes once you stop hardcoding a single tool and a single round trip. A more complex question might need two or three tool calls in sequence — look up the order, then check the shipping carrier’s tracking API, then format a reply — and you will not know the number in advance. So instead of writing out each step, you loop until the model stops asking for tools. python def run agent user input, tools, tool map, max iterations=6 : messages = {"role": "user", "content": user input} for in range max iterations : response = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools=tools, messages=messages, messages.append {"role": "assistant", "content": response.content} if response.stop reason = "tool use": return response.content 0 .text tool results = for block in response.content: if block.type == "tool use": function = tool map block.name output = function block.input tool results.append { "type": "tool result", "tool use id": block.id, "content": str output , } messages.append {"role": "user", "content": tool results} return "Stopped after max iterations without a final answer." 12345678910111213141516171819202122232425262728 def run agent user input, tools, tool map, max iterations=6 : messages = {"role": "user", "content": user input} for in range max iterations : response = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools=tools, messages=messages, messages.append {"role": "assistant", "content": response.content} if response.stop reason = "tool use": return response.content 0 .text tool results = for block in response.content: if block.type == "tool use": function = tool map block.name output = function block.input tool results.append { "type": "tool result", "tool use id": block.id, "content": str output , } messages.append {"role": "user", "content": tool results} return "Stopped after max iterations without a final answer." The max iterations cap prevents a specific failure: without it, a model that keeps deciding it needs “one more” tool call will loop until you run out of budget or patience. Capping the loop and returning a clear fallback message is a small line of code that saves you from a confusing incident in production. Strip away the surrounding code and this is the whole idea: a model, a loop, and a set of functions it can ask you to run. Nothing here is specific to order lookups. You can swap in a search call, an internal API, or a query of your own, and the same loop handles it, as long as the schema describes it clearly. Adding Memory run agent as written forgets everything the moment it returns. Call it twice in a row — first asking about order 4471, then asking “and when will it arrive?” — and the second call has no idea what “it” refers to, because each call builds a brand new messages list from scratch. Memory here means keeping that list around between calls instead of discarding it. The simplest way to do that is to stop passing messages in fresh each time and start keeping it on an object instead: python class Agent: def init self, tools, tool map, max iterations=6 : self.tools = tools self.tool map = tool map self.max iterations = max iterations self.messages = def run self, user input : self.messages.append {"role": "user", "content": user input} for in range self.max iterations : response = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools=self.tools, messages=self.messages, self.messages.append {"role": "assistant", "content": response.content} if response.stop reason = "tool use": return response.content 0 .text tool results = for block in response.content: if block.type == "tool use": function = self.tool map block.name output = function block.input tool results.append { "type": "tool result", "tool use id": block.id, "content": str output , } self.messages.append {"role": "user", "content": tool results} return "Stopped after max iterations without a final answer." 1234567891011121314151617181920212223242526272829303132333435 class Agent: def init self, tools, tool map, max iterations=6 : self.tools = tools self.tool map = tool map self.max iterations = max iterations self.messages = def run self, user input : self.messages.append {"role": "user", "content": user input} for in range self.max iterations : response = client.messages.create model="claude-sonnet-4-5", max tokens=512, tools=self.tools, messages=self.messages, self.messages.append {"role": "assistant", "content": response.content} if response.stop reason = "tool use": return response.content 0 .text tool results = for block in response.content: if block.type == "tool use": function = self.tool map block.name output = function block.input tool results.append { "type": "tool result", "tool use id": block.id, "content": str output , } self.messages.append {"role": "user", "content": tool results} return "Stopped after max iterations without a final answer." self.messages persists across calls to run , so the model sees the full conversation each time, including the earlier tool call and its result. Ask about order 4471, then ask “and when will it arrive?”, and the model can resolve “it” from the history already sitting in self.messages . This covers memory within a single running process. It does not cover what happens once that list grows past the model’s context window, or what happens if the process restarts and self.messages resets to empty. Those are real constraints of this approach, and they are the reason production agents usually add a step that trims or summarizes older turns, along with a place to persist the conversation outside the process itself — a database row, a file, a cache key tied to a session ID. Summary and Next Steps What exists at this point is a model call, a tool schema describing a function, a loop that executes tool calls until the model has enough information to answer, and a class that keeps the conversation across turns. That is a working agent, and every piece of it is code you wrote and can read back line by line. Here are a few directions worth exploring from here: - Testing which tool the model picks once there are several competing for the same question - Persisting and trimming conversation history so it does not overflow the context window - Handling a failing tool call, such as a timeout or a bad response from an upstream API, so it does not crash the loop - Adding logging around each tool call for debugging and later review Happy building