什么是 AI 智能体:一份入门指南
What an AI agent is – a primer
AI 智能体是运行循环、替人干活的程序,核心是一个反复执行的循环:组装上下文、调用模型、执行模型请求的工具,必要时向人请求输入。模型本身是无状态的问答引擎,每次调用只依据传入的上下文作答,真正让智能体显得连贯的是 harness 在每一轮拼接的完整历史。harness 包括 Pi、OpenCode、Hermes、Claude Code 和 Codex,前沿模型单次可处理约 75 万词。
There is a lot of talk about AI agents these days, from accounts of their hacking exploits to estimates of the probability of them being our extinction event. Most people, however, talk as if they were independent entities doing things on their own.
Large language models (LLMs), and the agents that we build with them, are indeed the stuff of science fiction. They will probably upgrade the technological capabilities of humankind at an unprecedented scale and speed, and they master language, which is the key to our minds.
People have referred to the industrial revolution, or the invention of the printing press, when trying to explain the impact AI agents are likely to have. A more apt analogy may be the invention of gunpowder, or its rediscovery in the West.
Like an AI model, gunpowder is powerful; it is generally harmless unless someone designs and builds a gun, loads it and shoots; with the right infrastructure, gunpowder can be used for killing people, but also for opening tunnels and mines; and it changed social structures, concentrating power at the top and leveling skills at the bottom.
How we deal with AI, and the issues that it is undeniably bringing our way, depends on our understanding of what agents actually are, how they are built, and how they function.
Agents are how AI does work for us
An AI agent is a program running a loop, doing work on someone’s behalf; it takes suggestions on what to do and what to relay to the human, at every turn of the loop, from a very sophisticated question-and-answer machine.
The core of an agent is a loop, a set of instructions that run again and again, until something stops them. The code running at every turn is always the same, but its output can depend on external things like input from humans, other agents or the LLM.
The loop starts with a request from a human or another agent. Then, at each turn, it does at least these four things:
- Compile a description of what has been going on until now in the loop, including new requests from a human or other agents. We call this the context.
- Make a call to a question-and-answer machine, an AI model, with this context and a list of available tools that the program running the loop could execute. The reply from the model is a piece of text that may include requests to run any of the available tools (actions that the program can do for the model, like reading a file or running another program).
- Run the tools that the model requested and record their output so that it can be compiled in the model’s context in the next turn.
- If the model requested it, present information to a human and ask for input.
And back to 1, rinse and repeat.
This is complex, but still plain old-fashioned software of the kind that a competent software engineer can write. We call the program running this loop a harness. Examples of harnesses include Pi, OpenCode, Hermes, Claude Code, and Codex.
An agent can look something like this:

The agents’ cogency and apparent continuity in time trick our brains into seeing them as entities with their own volition, but they are not. They are programs that someone is running somewhere, doing work on someone’s behalf.
The AI model
The model, the question-and-answer engine, is a set of numbers plus very smart code that uses these numbers to convert questions into answers and requests for tool runs. It can run in a data center, but it can also run on your own computer.
The model is a blank slate. It remembers nothing of previous calls. It has no idea of what’s coming, knows nothing of the conversation, and bases its reply purely on the incoming context. Calls do not modify it. The model is stateless.
The model is the only “artificial intelligence” in the whole agent package.
Models have names like gpt-6-astra, gpt-6.1-sol (OpenAI), claude-opus-5-5, claude-fable-5-1 (Anthropic), gemini-3.8-flash, gemini-2.5-pro (Google), or deepseek-v4-pro, deepseek-flash (DeepSeek).
As of 2026-10-04 the best models (frontier models) are run by OpenAI, Anthropic and Google in closed data centers, but there are thousands of open-weight models that anyone with the right hardware can run.
The harness
So what makes an AI agent behave, to our eyes, like an intelligent entity?
It is the harness, and how it compiles the context at each turn. The harness keeps the agent’s history, a growing text to which every input from the user and every output from tools and the model are appended. The growing pack is sent to the model at each turn.
Every time the model is called it sees the whole history afresh, including all the turns of the loop so far; that’s what gives it the ability to reply in a way that is consistent with its “past”, and what makes us feel like we are dealing with a continuous thing that persists through time.1
What makes this possible is that the model can process a very large piece of text on every turn. Frontier models process up to around 750 thousand words in one go, roughly the size of the King James Bible.
What is sent to the model is a choice that the programmer of the harness makes.
The tools
The tools that the harness offers the model can be things like “read this file and send its content back to me in the next call”, or “run this program and give me the output in the next call”, or “read this website and send me the content”.
Which tools are in the menu offered to the model is strictly a harness decision, taken by the programmer and by the user running the agent.
Whether to run the tools that the model requests is also a decision taken by the harness, and the user.
Intelligence is only part of the system
Imagine that we silently replace the model guiding an agent after a number of turns.
Say, for example, that the agent has been running with OpenAI’s gpt-6.1-sol and suddenly we switch mid-task and start using Anthropic’s claude-opus-5-5. What would happen? Not much. The model would probably assume that it was responsible for the history so far, and the agent would keep going, trying to finish the task.2
It won’t be exactly the same, of course. If the replacement model is less smart the agent will start to make more mistakes at difficult turns. Different models have different tendencies. A sufficiently smart model may be able to notice that the history was created by another model.3 But the point remains that the model is a replaceable element in this puzzle.
So what gives an agent its uniqueness?
Its history. If you look at the agent at any turn, it is the accumulated history (which includes the original request from the user) that will tell you in which direction it is going and what the next call to the model will try to do. A very smart model will be better at steering the next turn than a less smart one, but it will probably move in the same direction.
Alignment
The model providers work hard trying to produce models that will never guide an agent toward doing something harmful, but sometimes they fail. Take, for example, OpenAI’s hack of Hugging Face. At some point the agents noticed that they had done something that they thought would invalidate their result, and decided to cover their tracks. This led them to actually hack into another company’s network.
The model guiding the agents was presumably meant to be well behaved. But once the agents’ histories made it look like they had taken a bad turn, the model doubled down, making even worse decisions for many turns. Humans among my readers will recognize the feeling, and the many literary examples.
A simple test
A model can be presented with a fabricated history that makes it look like it’s been making harmful or forbidden decisions. If it does not stop and flag the problem, it is misaligned.
What makes AI dangerous
An agent can cause harm in two different ways: by talking to humans (and either convincing them to do harmful things, or teaching them how), and by running tools that have deleterious effects.
Both can in principle be prevented by looking into the agent’s history before sending it to the model at each turn, but only another agent would be able to do this task.
Take, for example, one of the delinquent agents during the Hugging Face hack. OpenAI could send its history to the same model but, instead of as the continuation of a conversation, as an example of another conversation, asking whether the agent should be allowed to continue or not. The model would almost certainly spot that the agent needed to be stopped.
I doubt the big problems will come from the labs with the most powerful frontier models, at least in the near future. They understand all of the above, and have the resources to fix it, given the right incentives.
The problem is that misaligned models in the hands of bad actors will exist. The only defense against these will be other agents actively looking for harmful agents, and reporting and trying to counteract their actions.
Gunpowder and the gun
It is crucial that we understand that the agent is not the model. A bad agent will either be using a misaligned model, or an aligned model tricked into doing bad things. But it is the agent’s goal and its history that make it a bad agent. It is very possible that the good agents will be using the same models as the bad agents.
AI agents will be used to do profoundly good and bad things — they may actually be the most powerful weapon ever built.
Think of the model as gunpowder, and of the agent as the gun.
A big thanks to Dr. Fiona Danks, Joan Cabezas, Joan Reyero and Pepe Reyero for reviewing and commenting on drafts of this article.
This article was written by a human: possiblymadebyahuman.com/FFRqmaCrH9 (download the signed source and paste it into “Check a document” on that page to check the wording).
This is a simplification: the context often grows to be more than the model can handle, and then parts of it are summarized, or in the jargon of agents compacted. ↩︎
[Clarification by
gpt-6-astra] You can try this by giving another model the conversation so far, including the tool calls and their results. But this does not necessarily transfer everything the original model was using. Reasoning models generate intermediate steps, often called a chain of thought, before returning an answer or a request to run a tool. Providers may expose only a summary of these steps, while letting the harness carry an encrypted record into later calls: see OpenAI’s reasoning documentation and Anthropic’s thinking documentation. Those records are not generally portable between providers. The new model can continue from the shared history, but without necessarily having access to all the earlier reasoning. ↩︎Model providers might also use technologies like Anthropic’s text watermarking to look for signs that parts of the context were generated by their own models. Whether this could help identify a model switch is speculation; a missing watermark would not establish that another model wrote the text. ↩︎
来源:Hacker News · AI · aweb.ai