如何用纯 Python 从零构建 AI 智能体
How (and Why) to Build an AI Agent from Scratch in Python
文章讲解如何不依赖编排框架,用纯 Python 和 Anthropic API 从零构建一个 AI 智能体,并给出完整代码。核心步骤包括定义工具 schema 让模型调用、实现请求-执行-响应循环(含 max_iterations 上限防止无限调用),以及用类保存 messages 列表实现跨轮次记忆。
In this article, you will learn what an AI agent is and how to build one from scratch in plain Python using the Anthropic API, without relying on any orchestration framework.
Topics we will cover include:
- Why tool calling is necessary and how to define a tool the model can use.
- How to implement the request-execute-respond loop that drives an agent’s behavior.
- How to add persistent memory so the agent maintains context across multiple turns.

Introduction
A simple call to a large language model is enough when the answer can be produced from what the model already knows: explaining a policy, drafting a reply, summarizing text. It is not enough once the answer depends on data outside that training set, such as a live order status or a row in your database; this gap is what an agent is built to close.
An AI agent is what results once the model can call out to something else — a function, a database query, an API — and use the result before it answers. Reaching for an agent orchestration framework to get this working is a reasonable move, eventually.
Writing one from scratch in plain Python, using a raw API call, makes the underlying architecture clear. At its core, there is simply a model, a loop that manages the interaction, and a small set of well-defined functions that give the model the capabilities it needs. Stripping away the frameworks and abstractions makes it easier to see how these pieces fit together and what is actually happening under the hood.
This article covers:
- Why tool calling is necessary and what you need installed to try it
- How to describe a tool so the model knows when and how to call it
- What the request-execute-respond loop looks like in code
- How to add memory so the agent keeps context across turns
You can find the companion code for this article in this GitHub repository.
Prerequisites
You need Python 3.10 or later, an Anthropic API key, and the Anthropic Python SDK:
|
1 |
pip install anthropic |
Set your key as an environment variable so the client picks it up without hardcoding it anywhere:
|
1 |
export ANTHROPIC_API_KEY="your-key-here" |
This is all you need to get started. You can find all the code in the agent.py script.
⚠️ A note to readers: Anthropic will officially retire Claude Sonnet 4.5 on November 30, 2026. If you are following this tutorial after that date, the specific model code used in these examples will no longer work. To ensure your project runs smoothly, replace the model string in the API calls with a newer version or the latest available model.
Setting Up the Model Call
A minimal wrapper around the API is a function that sends a prompt and returns text:
|
1 2 3 4 5 6 7 8 9 10 11 12 |
import anthropic client = anthropic.Anthropic() def ask(prompt): response = client.messages.create( model="claude-sonnet-4-5", max_tokens=512, system="You are a helpful support assistant. Be direct and factual.", messages=[{"role": "user", "content": prompt}], ) return response.content[0].text |
This handles a large share of questions well — explaining a policy, drafting a reply, summarizing a paragraph. It fails the moment the answer depends on data the model was never shown. Ask ask("What is the status of order #4471?") and the model has no mechanism to check: that information lives in your database, not in its training data. It either states that it does not know, or produces a plausible answer anyway, and no amount of prompting changes that, since no prompt grants the model access to data it was never given. Tool calling gives the model a defined way to request that data instead of inferring it.
Read The Roadmap to Mastering Tool Calling in AI Agents to learn more.
Defining a Tool the Model Can Ask For
Giving the model access to a function means writing the function, and writing a description of it that the model can read. The description matters because it tells the model what the function does and when to reach for it.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 |
orders_db = { "4471": {"status": "shipped", "carrier": "UPS", "eta": "2 days"}, "4472": {"status": "processing", "carrier": None, "eta": None}, } def get_order_status(order_id): return orders_db.get(order_id, {"error": "No order found with that ID"}) get_order_status_schema = { "name": "get_order_status", "description": ( "Looks up the current status of a customer order by its ID. " "Use this any time a question depends on current order data " "rather than general policy information." ), "input_schema": { "type": "object", "properties": { "order_id": {"type": "string", "description": "The order ID to look up, e.g. '4471'"}, }, "required": ["order_id"], }, } |
The schema is structured metadata: a name, a description, and a definition of the parameters the function expects. Tool use works the same way across most providers: you pass the model a list of these schemas alongside your message, and the model decides on its own whether answering the question requires calling one of them.
Executing the Tool and Feeding the Result Back
Pass the schema in, and something different happens. Instead of a plain text answer, the response comes back with a stop_reason of tool_use and a content block describing which function the model wants to call and with what arguments. The model has not run anything yet — it is paused, and has returned a request for you to act on.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 |
messages = [{"role": "user", "content": "What is the status of order #4471?"}] response = client.messages.create( model="claude-sonnet-4-5", max_tokens=512, tools=[get_order_status_schema], messages=messages, ) if response.stop_reason == "tool_use": tool_call = next(block for block in response.content if block.type == "tool_use") result = get_order_status(**tool_call.input) messages.append({"role": "assistant", "content": response.content}) messages.append({ "role": "user", "content": [{ "type": "tool_result", "tool_use_id": tool_call.id, "content": str(result), }], }) final = client.messages.create( model="claude-sonnet-4-5", max_tokens=512, tools=[get_order_status_schema], messages=messages, ) print(final.content[0].text) |
Two round trips happen here. The first asks the model what it wants to do. You run the function yourself — the lookup against orders_db, or in a production system, a call to your order service — append the result back into the conversation as a tool_result, and send everything again. The second round trip is where the model reads that result and gives you an answer grounded in it, instead of a guess.

Looping Until the Model Has What It Needs
This pattern generalizes once you stop hardcoding a single tool and a single round trip. A more complex question might need two or three tool calls in sequence — look up the order, then check the shipping carrier’s tracking API, then format a reply — and you will not know the number in advance. So instead of writing out each step, you loop until the model stops asking for tools.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 |
def run_agent(user_input, tools, tool_map, max_iterations=6): messages = [{"role": "user", "content": user_input}] for _ in range(max_iterations): response = client.messages.create( model="claude-sonnet-4-5", max_tokens=512, tools=tools, messages=messages, ) messages.append({"role": "assistant", "content": response.content}) if response.stop_reason != "tool_use": return response.content[0].text tool_results = [] for block in response.content: if block.type == "tool_use": function = tool_map[block.name] output = function(**block.input) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": str(output), }) messages.append({"role": "user", "content": tool_results}) return "Stopped after max iterations without a final answer." |
The max_iterations cap prevents a specific failure: without it, a model that keeps deciding it needs “one more” tool call will loop until you run out of budget or patience. Capping the loop and returning a clear fallback message is a small line of code that saves you from a confusing incident in production.
Strip away the surrounding code and this is the whole idea: a model, a loop, and a set of functions it can ask you to run. Nothing here is specific to order lookups. You can swap in a search call, an internal API, or a query of your own, and the same loop handles it, as long as the schema describes it clearly.
Adding Memory
run_agent as written forgets everything the moment it returns. Call it twice in a row — first asking about order #4471, then asking “and when will it arrive?” — and the second call has no idea what “it” refers to, because each call builds a brand new messages list from scratch. Memory here means keeping that list around between calls instead of discarding it.

The simplest way to do that is to stop passing messages in fresh each time and start keeping it on an object instead:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 |
class Agent: def __init__(self, tools, tool_map, max_iterations=6): self.tools = tools self.tool_map = tool_map self.max_iterations = max_iterations self.messages = [] def run(self, user_input): self.messages.append({"role": "user", "content": user_input}) for _ in range(self.max_iterations): response = client.messages.create( model="claude-sonnet-4-5", max_tokens=512, tools=self.tools, messages=self.messages, ) self.messages.append({"role": "assistant", "content": response.content}) if response.stop_reason != "tool_use": return response.content[0].text tool_results = [] for block in response.content: if block.type == "tool_use": function = self.tool_map[block.name] output = function(**block.input) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": str(output), }) self.messages.append({"role": "user", "content": tool_results}) return "Stopped after max iterations without a final answer." |
self.messages persists across calls to run(), so the model sees the full conversation each time, including the earlier tool call and its result. Ask about order #4471, then ask “and when will it arrive?”, and the model can resolve “it” from the history already sitting in self.messages.
This covers memory within a single running process. It does not cover what happens once that list grows past the model’s context window, or what happens if the process restarts and self.messages resets to empty. Those are real constraints of this approach, and they are the reason production agents usually add a step that trims or summarizes older turns, along with a place to persist the conversation outside the process itself — a database row, a file, a cache key tied to a session ID.
Summary and Next Steps
What exists at this point is a model call, a tool schema describing a function, a loop that executes tool calls until the model has enough information to answer, and a class that keeps the conversation across turns. That is a working agent, and every piece of it is code you wrote and can read back line by line.
Here are a few directions worth exploring from here:
- Testing which tool the model picks once there are several competing for the same question
- Persisting and trimming conversation history so it does not overflow the context window
- Handling a failing tool call, such as a timeout or a bad response from an upstream API, so it does not crash the loop
- Adding logging around each tool call for debugging and later review
Happy building!
No comments yet.
来源:Hacker News · AI · machinelearningmastery.com