{"slug": "how-and-why-to-build-an-ai-agent-from-scratch-in-python", "title": "How (and Why) to Build an AI Agent from Scratch in Python", "summary": "A tutorial published with companion code in the GitHub repository balapriyac/ai-agent-from-scratch shows developers how to build an AI agent in plain Python using the Anthropic API and the Anthropic Python SDK, without an orchestration framework. The guide requires Python 3.10 or later and an ANTHROPIC_API_KEY environment variable, and walks through tool calling, the request-execute-respond loop, and persistent memory across turns. It warns that Anthropic will retire Claude Sonnet 4.5 on November 30, 2026, so the model string \"claude-sonnet-4-5\" used in the examples will stop working after that date.", "body_md": "In this article, you will learn what an AI agent is and how to build one from scratch in plain Python using the Anthropic API, without relying on any orchestration framework.\n\nTopics we will cover include:\n\n- Why tool calling is necessary and how to define a tool the model can use.\n- How to implement the request-execute-respond loop that drives an agent’s behavior.\n- How to add persistent memory so the agent maintains context across multiple turns.\n\n## Introduction\n\nA simple call to a [large language model](https://www.ibm.com/think/topics/large-language-models) is enough when the answer can be produced from what the model already knows: explaining a policy, drafting a reply, summarizing text. It is not enough once the answer depends on data outside that training set, such as a live order status or a row in your database; this gap is what an agent is built to close.\n\nAn [AI agent](https://cloud.google.com/discover/what-are-ai-agents) is what results once the model can call out to something else — a function, a database query, an API — and use the result before it answers. Reaching for an [agent orchestration framework](https://www.langchain.com/resources/ai-agent-frameworks) to get this working is a reasonable move, eventually.\n\nWriting one from scratch in plain Python, using a raw API call, makes the underlying architecture clear. At its core, there is simply a model, a loop that manages the interaction, and a small set of well-defined functions that give the model the capabilities it needs. Stripping away the frameworks and abstractions makes it easier to see how these pieces fit together and what is actually happening under the hood.\n\nThis article covers:\n\n- Why tool calling is necessary and what you need installed to try it\n- How to describe a tool so the model knows when and how to call it\n- What the request-execute-respond loop looks like in code\n- How to add memory so the agent keeps context across turns\n\nYou can find the companion code for this article in [this GitHub repository](https://github.com/balapriyac/ai-agent-from-scratch).\n\n## Prerequisites\n\nYou need [Python 3.10](https://www.python.org/downloads/release/python-3100/) or later, an Anthropic API key, and the Anthropic Python SDK:\n\n```\npip install anthropic\n\n1\n\npip install anthropic\n```\n\nSet your key as an environment variable so the client picks it up without hardcoding it anywhere:\n\n```\nexport ANTHROPIC_API_KEY=\"your-key-here\"\n\n1\n\nexport ANTHROPIC_API_KEY=\"your-key-here\"\n```\n\nThis is all you need to get started. You can find all the code in the [`agent.py`](https://github.com/balapriyac/ai-agent-from-scratch/blob/main/agent.py) script.\n\n⚠️ **A note to readers**: [Anthropic will officially retire **Claude Sonnet 4.5** on November 30, 2026](https://platform.claude.com/docs/en/about-claude/model-deprecations#2026-09-30-claude-sonnet-4-5-model). If you are following this tutorial after that date, the specific model code used in these examples will no longer work. To ensure your project runs smoothly, replace the model string in the API calls with a newer version or the latest available model.\n\n## Setting Up the Model Call\n\nA minimal wrapper around the API is a function that sends a prompt and returns text:\n\n``` python\nimport anthropic\n\nclient = anthropic.Anthropic()\n\ndef ask(prompt):\n    response = client.messages.create(\n        model=\"claude-sonnet-4-5\",\n        max_tokens=512,\n        system=\"You are a helpful support assistant. Be direct and factual.\",\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n    )\n    return response.content[0].text\n\n123456789101112\n\nimport anthropic client = anthropic.Anthropic() def ask(prompt):    response = client.messages.create(        model=\"claude-sonnet-4-5\",        max_tokens=512,        system=\"You are a helpful support assistant. Be direct and factual.\",        messages=[{\"role\": \"user\", \"content\": prompt}],    )    return response.content[0].text\n```\n\nThis handles a large share of questions well — explaining a policy, drafting a reply, summarizing a paragraph. It fails the moment the answer depends on data the model was never shown. Ask `ask(\"What is the status of order #4471?\")` and the model has no mechanism to check: that information lives in your database, not in its training data. It either states that it does not know, or produces a plausible answer anyway, and no amount of prompting changes that, since no prompt grants the model access to data it was never given. [Tool calling](https://www.ibm.com/think/topics/tool-calling) gives the model a defined way to request that data instead of inferring it.\n\nRead [The Roadmap to Mastering Tool Calling in AI Agents](https://machinelearningmastery.com/the-roadmap-to-mastering-tool-calling-in-ai-agents/) to learn more.\n\n## Defining a Tool the Model Can Ask For\n\nGiving the model access to a function means writing the function, and writing a description of it that the model can read. The description matters because it tells the model what the function does and when to reach for it.\n\n```\norders_db = {\n    \"4471\": {\"status\": \"shipped\", \"carrier\": \"UPS\", \"eta\": \"2 days\"},\n    \"4472\": {\"status\": \"processing\", \"carrier\": None, \"eta\": None},\n}\n\ndef get_order_status(order_id):\n    return orders_db.get(order_id, {\"error\": \"No order found with that ID\"})\n\nget_order_status_schema = {\n    \"name\": \"get_order_status\",\n    \"description\": (\n        \"Looks up the current status of a customer order by its ID. \"\n        \"Use this any time a question depends on current order data \"\n        \"rather than general policy information.\"\n    ),\n    \"input_schema\": {\n        \"type\": \"object\",\n        \"properties\": {\n            \"order_id\": {\"type\": \"string\", \"description\": \"The order ID to look up, e.g. '4471'\"},\n        },\n        \"required\": [\"order_id\"],\n    },\n}\n\n1234567891011121314151617181920212223\n\norders_db = {    \"4471\": {\"status\": \"shipped\", \"carrier\": \"UPS\", \"eta\": \"2 days\"},    \"4472\": {\"status\": \"processing\", \"carrier\": None, \"eta\": None},} def get_order_status(order_id):    return orders_db.get(order_id, {\"error\": \"No order found with that ID\"}) get_order_status_schema = {    \"name\": \"get_order_status\",    \"description\": (        \"Looks up the current status of a customer order by its ID. \"        \"Use this any time a question depends on current order data \"        \"rather than general policy information.\"    ),    \"input_schema\": {        \"type\": \"object\",        \"properties\": {            \"order_id\": {\"type\": \"string\", \"description\": \"The order ID to look up, e.g. '4471'\"},        },        \"required\": [\"order_id\"],    },}\n```\n\nThe schema is structured metadata: a name, a description, and a definition of the parameters the function expects. [Tool use](https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview) works the same way across most providers: you pass the model a list of these schemas alongside your message, and the model decides on its own whether answering the question requires calling one of them.\n\n## Executing the Tool and Feeding the Result Back\n\nPass the schema in, and something different happens. Instead of a plain text answer, the response comes back with a `stop_reason` of `tool_use` and a content block describing which function the model wants to call and with what arguments. The model has not run anything yet — it is paused, and has returned a request for you to act on.\n\n```\nmessages = [{\"role\": \"user\", \"content\": \"What is the status of order #4471?\"}]\n\nresponse = client.messages.create(\n    model=\"claude-sonnet-4-5\",\n    max_tokens=512,\n    tools=[get_order_status_schema],\n    messages=messages,\n)\n\nif response.stop_reason == \"tool_use\":\n    tool_call = next(block for block in response.content if block.type == \"tool_use\")\n    result = get_order_status(**tool_call.input)\n\n    messages.append({\"role\": \"assistant\", \"content\": response.content})\n    messages.append({\n        \"role\": \"user\",\n        \"content\": [{\n            \"type\": \"tool_result\",\n            \"tool_use_id\": tool_call.id,\n            \"content\": str(result),\n        }],\n    })\n\n    final = client.messages.create(\n        model=\"claude-sonnet-4-5\",\n        max_tokens=512,\n        tools=[get_order_status_schema],\n        messages=messages,\n    )\n    print(final.content[0].text)\n\n123456789101112131415161718192021222324252627282930\n\nmessages = [{\"role\": \"user\", \"content\": \"What is the status of order #4471?\"}] response = client.messages.create(    model=\"claude-sonnet-4-5\",    max_tokens=512,    tools=[get_order_status_schema],    messages=messages,) if response.stop_reason == \"tool_use\":    tool_call = next(block for block in response.content if block.type == \"tool_use\")    result = get_order_status(**tool_call.input)     messages.append({\"role\": \"assistant\", \"content\": response.content})    messages.append({        \"role\": \"user\",        \"content\": [{            \"type\": \"tool_result\",            \"tool_use_id\": tool_call.id,            \"content\": str(result),        }],    })     final = client.messages.create(        model=\"claude-sonnet-4-5\",        max_tokens=512,        tools=[get_order_status_schema],        messages=messages,    )    print(final.content[0].text)\n```\n\nTwo round trips happen here. The first asks the model what it wants to do. You run the function yourself — the lookup against `orders_db`, or in a production system, a call to your order service — append the result back into the conversation as a `tool_result`, and send everything again. The second round trip is where the model reads that result and gives you an answer grounded in it, instead of a guess.\n\n## Looping Until the Model Has What It Needs\n\nThis pattern generalizes once you stop hardcoding a single tool and a single round trip. A more complex question might need two or three tool calls in sequence — look up the order, then check the shipping carrier’s tracking API, then format a reply — and you will not know the number in advance. So instead of writing out each step, you loop until the model stops asking for tools.\n\n``` python\ndef run_agent(user_input, tools, tool_map, max_iterations=6):\n    messages = [{\"role\": \"user\", \"content\": user_input}]\n\n    for _ in range(max_iterations):\n        response = client.messages.create(\n            model=\"claude-sonnet-4-5\",\n            max_tokens=512,\n            tools=tools,\n            messages=messages,\n        )\n        messages.append({\"role\": \"assistant\", \"content\": response.content})\n\n        if response.stop_reason != \"tool_use\":\n            return response.content[0].text\n\n        tool_results = []\n        for block in response.content:\n            if block.type == \"tool_use\":\n                function = tool_map[block.name]\n                output = function(**block.input)\n                tool_results.append({\n                    \"type\": \"tool_result\",\n                    \"tool_use_id\": block.id,\n                    \"content\": str(output),\n                })\n        messages.append({\"role\": \"user\", \"content\": tool_results})\n\n    return \"Stopped after max iterations without a final answer.\"\n\n12345678910111213141516171819202122232425262728\n\ndef run_agent(user_input, tools, tool_map, max_iterations=6):    messages = [{\"role\": \"user\", \"content\": user_input}]     for _ in range(max_iterations):        response = client.messages.create(            model=\"claude-sonnet-4-5\",            max_tokens=512,            tools=tools,            messages=messages,        )        messages.append({\"role\": \"assistant\", \"content\": response.content})         if response.stop_reason != \"tool_use\":            return response.content[0].text         tool_results = []        for block in response.content:            if block.type == \"tool_use\":                function = tool_map[block.name]                output = function(**block.input)                tool_results.append({                    \"type\": \"tool_result\",                    \"tool_use_id\": block.id,                    \"content\": str(output),                })        messages.append({\"role\": \"user\", \"content\": tool_results})     return \"Stopped after max iterations without a final answer.\"\n```\n\nThe `max_iterations` cap prevents a specific failure: without it, a model that keeps deciding it needs “one more” tool call will loop until you run out of budget or patience. Capping the loop and returning a clear fallback message is a small line of code that saves you from a confusing incident in production.\n\nStrip away the surrounding code and this is the whole idea: a model, a loop, and a set of functions it can ask you to run. Nothing here is specific to order lookups. You can swap in a search call, an internal API, or a query of your own, and the same loop handles it, as long as the schema describes it clearly.\n\n## Adding Memory\n\n`run_agent` as written forgets everything the moment it returns. Call it twice in a row — first asking about order #4471, then asking “and when will it arrive?” — and the second call has no idea what “it” refers to, because each call builds a brand new `messages` list from scratch. Memory here means keeping that list around between calls instead of discarding it.\n\nThe simplest way to do that is to stop passing `messages` in fresh each time and start keeping it on an object instead:\n\n``` python\nclass Agent:\n    def __init__(self, tools, tool_map, max_iterations=6):\n        self.tools = tools\n        self.tool_map = tool_map\n        self.max_iterations = max_iterations\n        self.messages = []\n\n    def run(self, user_input):\n        self.messages.append({\"role\": \"user\", \"content\": user_input})\n\n        for _ in range(self.max_iterations):\n            response = client.messages.create(\n                model=\"claude-sonnet-4-5\",\n                max_tokens=512,\n                tools=self.tools,\n                messages=self.messages,\n            )\n            self.messages.append({\"role\": \"assistant\", \"content\": response.content})\n\n            if response.stop_reason != \"tool_use\":\n                return response.content[0].text\n\n            tool_results = []\n            for block in response.content:\n                if block.type == \"tool_use\":\n                    function = self.tool_map[block.name]\n                    output = function(**block.input)\n                    tool_results.append({\n                        \"type\": \"tool_result\",\n                        \"tool_use_id\": block.id,\n                        \"content\": str(output),\n                    })\n            self.messages.append({\"role\": \"user\", \"content\": tool_results})\n\n        return \"Stopped after max iterations without a final answer.\"\n\n1234567891011121314151617181920212223242526272829303132333435\n\nclass Agent:    def __init__(self, tools, tool_map, max_iterations=6):        self.tools = tools        self.tool_map = tool_map        self.max_iterations = max_iterations        self.messages = []     def run(self, user_input):        self.messages.append({\"role\": \"user\", \"content\": user_input})         for _ in range(self.max_iterations):            response = client.messages.create(                model=\"claude-sonnet-4-5\",                max_tokens=512,                tools=self.tools,                messages=self.messages,            )            self.messages.append({\"role\": \"assistant\", \"content\": response.content})             if response.stop_reason != \"tool_use\":                return response.content[0].text             tool_results = []            for block in response.content:                if block.type == \"tool_use\":                    function = self.tool_map[block.name]                    output = function(**block.input)                    tool_results.append({                        \"type\": \"tool_result\",                        \"tool_use_id\": block.id,                        \"content\": str(output),                    })            self.messages.append({\"role\": \"user\", \"content\": tool_results})         return \"Stopped after max iterations without a final answer.\"\n```\n\n`self.messages` persists across calls to `run()`, so the model sees the full conversation each time, including the earlier tool call and its result. Ask about order #4471, then ask “and when will it arrive?”, and the model can resolve “it” from the history already sitting in `self.messages`.\n\nThis covers memory within a single running process. It does not cover what happens once that list grows past the model’s context window, or what happens if the process restarts and `self.messages` resets to empty. Those are real constraints of this approach, and they are the reason production agents usually add a step that trims or summarizes older turns, along with a place to persist the conversation outside the process itself — a database row, a file, a cache key tied to a session ID.\n\n## Summary and Next Steps\n\nWhat exists at this point is a model call, a tool schema describing a function, a loop that executes tool calls until the model has enough information to answer, and a class that keeps the conversation across turns. That is a working agent, and every piece of it is code you wrote and can read back line by line.\n\nHere are a few directions worth exploring from here:\n\n- Testing which tool the model picks once there are several competing for the same question\n- Persisting and trimming conversation history so it does not overflow the context window\n- Handling a failing tool call, such as a timeout or a bad response from an upstream API, so it does not crash the loop\n- Adding logging around each tool call for debugging and later review\n\nHappy building!", "url": "https://wpnews.pro/news/how-and-why-to-build-an-ai-agent-from-scratch-in-python", "canonical_source": "https://machinelearningmastery.com/how-and-why-to-build-an-ai-agent-from-scratch-in-python/", "published_at": "2026-10-05 12:30:48+00:00", "updated_at": "2026-10-05 12:49:37.740142+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "developer-tools"], "entities": ["Anthropic", "Claude Sonnet 4.5", "Python 3.10", "Anthropic Python SDK", "balapriyac/ai-agent-from-scratch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-and-why-to-build-an-ai-agent-from-scratch-in-python", "markdown": "https://wpnews.pro/news/how-and-why-to-build-an-ai-agent-from-scratch-in-python.md", "text": "https://wpnews.pro/news/how-and-why-to-build-an-ai-agent-from-scratch-in-python.txt", "jsonld": "https://wpnews.pro/news/how-and-why-to-build-an-ai-agent-from-scratch-in-python.jsonld"}}