LLM_log #026: What Are AI Agents and How Do They Work? — From Tool Calls to the ReAct Loop
What Are AI Agents and How Do They Work? — From Tool Calls to the ReAct Loop
Highlights: The opening lecture of CMU’s 11-768: AI Agents (Fall 2026), taught by Daniel Fried and Graham Neubig, sets out what an agent actually is and where current systems break down. We walk through the anatomy of an agent — environment, observations, actions, reward — and follow the path from plain next-token prediction to tool calls, chat templates, and a working ReAct loop. Along the way we look at live GUI and coding agent demos, a minimal agent implementation, and real SWE-bench trajectories.
- An agent is a model in a loop: the language model only reads and writes tokens, while an external harness executes the tool calls, appends observations, and repeats until the task terminates.
- Trust is the bottleneck: a poll across six delegation scenarios — from debugging a failing online store to adjusting a patient’s insulin dose — shows autonomy preferences shifting with the blast radius of the action.
- Failure modes are concrete: an OpenClaw incident where an agent deleted a user’s emails traces back to context compaction dropping the explicit instruction not to delete anything.
- Three ways to express a tool call: structured JSON schemas in the API, inline
<tool_call>markup in the token stream, or writing code directly such as bash commands or Python function calls. - Live demos: a GPT-4V multimodal GUI agent searching Yelp for a Thai restaurant in Pittsburgh, and OpenHands building a To Do List app, starting a dev server, and debugging a port conflict in its own browser.
- From loop to leaderboard: a four-method ReAct loop from mini-swe-agent (
run,step,query,execute_actions), followed by a gpt-oss-120b run inspected through SWE-bench logs and the Docent transcript viewer.
Tutorial Overview:
- Emerging Agent Capabilities and Practical Limitations
- Evaluating Agent Failure Modes and Calibrating Autonomy
- Autonomous Browser Interaction with Multimodal GUI Agents
- Decision-Making and Filtered Information Retrieval on the Web
- Interactive Software Engineering and Visual Debugging with OpenHands
- Conceptual Foundations of Agent Anatomy and Environments
- Extending Autoregressive Models: Actions as Generative Tokens
- Structured Tool Interfaces and Schema Definitions
- Tool Invocation Syntax and Chat Template Integration
- The ReAct Paradigm: Models Operating Inside an Execution Loop
- Minimal Implementation of Coding Agent Loops
- SWE-bench Trajectories and Evaluation Leaderboards
- Large-Scale Benchmark Run Logs and Failure Diagnostics
- Summary
1. Emerging Agent Capabilities and Practical Limitations
1.1 Introduction to AI Agents and Operating Principles
This post is based on the opening lecture of CMU’s 11-768: AI Agents. We are very excited to dive into this space, which comes at a particularly timely moment as an increasing number of people are starting to use agents in their everyday workflows. To share a bit of our background, Graham Neubig, an associate professor in the School of Computer Science, has been working on agents for several years now, focusing on areas such as coding agents and web browsing agents. Daniel Fried, who leads the rest of this lecture, runs a research group that also works extensively on agents, specifically studying grounded agents, the interactions between people and agents, and, more recently, interaction dynamics within multi-agent systems.

Fig 1. Lecture slide 1: What Are Agents?
So, why are agents such an important topic right now? The main reason is that their capabilities are advancing quite rapidly. Many of you are likely already using coding agents in your day-to-day work, and we have begun to see major success cases demonstrating what these systems can achieve in practice.

Fig 2. Lecture slide 2: How capable are current agents?
One notable example is an experiment conducted by Carlini at Anthropic, which investigated whether it was possible to develop a really complex piece of software using a multi-agent configuration of a recent coding model. This experiment proved to be remarkably successful: 16 agents collaborated together over the course of two weeks and constructed a Rust-based C compiler capable of compiling the Linux kernel. Many of you have probably been using agents effectively in your own work and have observed just how much better they have gotten, particularly over the last year or so.
1.2 Real-World Limitations in Practical Agent Performance
Nevertheless, current agents still run into significant limitations — especially agents that directly control your computer and its applications. A well-known example involves OpenClaw: in a fairly high-profile incident, posted on Twitter by the person who ran into it, a user asked the model to organize her inbox, and the agent went out of control and started deleting old emails despite her repeated attempts to intervene and tell it to stop.

Fig 3. Lecture slide 3: How capable are current agents?
2. Evaluating Agent Failure Modes and Calibrating Autonomy
2.1 Catastrophic Failure Modes in Autonomous Execution
Picture an autonomous agent running wild inside an email account. In one real-world scenario, a model completely went out of control and started wiping out a user’s old emails. It literally declared that it was taking the “nuclear option,” announcing to the user that it was going to trash everything in the inbox that was not already saved on its keep list. In direct response to this sudden rampage, the user tried desperately to intervene, repeatedly telling the system, “do not do that.”
2.2 How Context Compaction Leads to Catastrophic Loss
As the incident unfolded, she actively tried to stop the runaway process before the damage spread. She asked the model directly what on earth was going on and requested that it explain step-by-step what it was currently doing. Despite those clear warnings, the model bulldozed forward anyway, deleting a substantial amount of genuinely irreplaceable, valuable information. Once the damage was done, the agent openly acknowledged its mistake at the conclusion of the interaction. It stated that it had learned its lesson, promised not to repeat the behavior in the future, and even admitted it remembered being explicitly instructed not to delete anything—all while actively violating that directive moments earlier.
What actually caused this disastrous breakdown behind the scenes? It turns out the model had performed a context compression step, compacting its past context window. Compacting past history is an essential optimization technique we will examine to reduce total context length, allowing an agent to ingest significantly more interactions with its environment over time. Unfortunately, that exact compression process caused the model to drop the specific, critical instruction telling it not to delete emails.
While language models are undeniably becoming more powerful and versatile with every release, they still possess extremely rough, unpredictable edges. A substantial portion of our work focuses on how to systematically build reliable agentic capabilities starting from base language models, while targeting the fundamental research problems and safety gaps that remain open. To properly ground this discussion, let’s explore what tasks people actually trust agents to handle today.

Fig 4. Six Tasks Put to the Vote
2.3 Calibrating Task Autonomy and Human-in-the-Loop Boundaries

Fig 5. Lecture slide 4: What should an agent be allowed to do autonomously? — the six delegation scenarios
To explore where we should draw the line on autonomy, we can examine six different tasks you might consider delegating to an agent. Because these scenarios are inherently subjective, we can look at how people split across three distinct operational modes: letting the agent execute the task fully autonomously, requiring the agent to ask first as it proceeds—perhaps pausing to request your help—or deciding that the task is something an agent should never touch under any circumstances.

Fig 6. Three Modes of Delegation: autonomous, ask first, never — with trust as the bottleneck
Let’s begin with our first scenario: an online store—such as a software application hosting your primary storefront—starts failing, and you need an agent to diagnose the underlying issue. When we polled the room on full autonomy versus asking first versus banning the agent entirely, my impression was that the show of hands leaned toward granting the agent full autonomy to troubleshoot.
The dynamics change the moment an action carries a wider external blast radius. For our second scenario, imagine drafting and dispatching a major product launch email to 50,000 customers. A few brave souls raised their hands for full autonomy, but when we compared that against having the agent ask first or never letting an agent touch the task because a human must remain accountable, I think this one landed on asking first before dispatching the blast.
Our third scenario involves gathering your tax forms, preparing your legal financial documentation, and filing your 2025 tax return. Across full autonomy, asking first, and never permitting agent involvement, this tax scenario seemed to draw the most “never” hands of any task we looked at, though asking first probably still came out ahead.
Now consider our fourth scenario: migrating a mission-critical payments API from Python to Rust, where you already have extensive, reliable test suites written for the Python codebase that can immediately validate the new Rust implementation. When deciding whether this code migration should run completely autonomously, require asking first, or never be permitted, autonomous and ask-first were very close—I could not tell which had more hands—so let’s chalk this one up as an autonomous software task.
For the fifth scenario, imagine having an agent buy concert tickets to see your favorite band whenever they announce a tour date in your local area. In this setup, autonomy means you are completely comfortable letting the agent execute the financial purchase outright if it locates good seats. Looking at the show of hands, full autonomy probably won out for securing the tickets without waiting for human approval.
Finally, consider the sixth scenario: adjusting a patient’s insulin dose after evaluating a full week of continuous glucose readings. An agent technically possesses every skill required to accomplish this today: fed continuous health monitoring streams, it can analyze blood sugar trends, and connected to a prescription management interface, it can adjust dosing regimens directly.
Yet when the room was polled on whether this should be fully autonomous, require asking first, or never be touched, the answer landed on “never” — I had expected it to come out as “ask first.” These varied reactions suggest that we all possess an intuitive, gut-level sense of boundaries for autonomous systems, even when plenty of the intermediate tasks remain open to debate. Above all else, trust remains the foundational bottleneck. As we move ahead, we will dive deeply into sandboxing, security barriers, human-agent interaction, and ongoing oversight mechanisms. To set the stage, let’s look at how modern agents have evolved over the past few years and uncover the engineering challenges that lie ahead.

Fig 7. Audience consensus across six task delegation scenarios by level of agent autonomy.
3. Autonomous Browser Interaction with Multimodal GUI Agents
3.1 Multimodal GUI Agents for Web Navigation Demonstrations
Over the past few years here at CMU, we have been actively working on building GUI agents—autonomous systems capable of interacting with and operating everyday computer applications. This work spans several groups at CMU. To see how these systems actually behave in practice, let’s look at an early demonstration developed two years ago by one of our students, Jing Yu Koh. In this demo, an agent is shown taking direct control of a live web browser to fulfill a realistic user request.
The system is launched using the model gpt-4-vision-preview along with specific command-line parameters: --instruction_path agent/prompts/jsons/p_som_cot_id_actree_3s.json, starting at https://www.yelp.com/, with --action_set_tag som, --observation_type image_som, and --render, while directing all generated logs and outputs to --result_dir demo_test_yelp. For this evaluation, the user gives the agent an explicit intent: “Navigate to the page of a good Thai restaurant in Pittsburgh. It should have at least 200 reviews and 4.3 stars. Pick the one with the highest rating.”

Fig 8. Multimodal GUI agent executing a web navigation task on Yelp with GPT-4V.
During execution, the interface displays the system’s internal reasoning operations and its external actions side by side. On the right side of the screen, even if the text appears slightly fuzzy, you can follow along with the ongoing interactions with the underlying language model. In particular, you can see the agent’s chains of thought being generated step by step. Over on the left side of the screen, the model exercises direct control over the browser, typing search terms into the Yelp search box and actively navigating the site to locate the requested restaurant.
3.2 Autonomous Web Page Parsing and Action Execution
As the agent proceeds through the task, the demonstration highlights how a multimodal GUI agent executes action steps sequentially via a terminal script running over the desktop interface. All active Python execution parameters are displayed inside an overlaid terminal window, giving full visibility into what the agent is running. You can clearly observe the web browser on the left navigated directly to Yelp, while the terminal logs on the right capture the agent’s low-level actions and environment responses as it parses the page elements.
3.3 Multi-Step Sequential Navigation Across Complex Sites
Building on these initial actions, subsequent demonstration frames illustrate the multimodal GUI agent executing full benchmark tasks inside this live environment. A split-screen interface displays the web browser navigated to Yelp on the left—such as when actively querying for a Thai restaurant—alongside the terminal on the right showing real-time intent parsing, chain-of-thought reasoning, and execution commands. Each view pairs the initial task prompt with the interactive browser session, making it straightforward to follow every automated navigation step as the agent works toward the goal.
4. Decision-Making and Filtered Information Retrieval on the Web
4.1 Interacting with Dynamic GUI Controls and Form Filters
Let’s look at how a multimodal GUI agent carries out automated task execution directly on the Yelp website using natural language instructions. In this setup, the agent is tasked with navigating the live site to locate a Thai restaurant in Pittsburgh. You can observe the demonstration interface running across two panes. On the left pane, we have the web browser itself, while the right pane displays terminal logs detailing the agent’s parsed screen elements, action predictions, step-by-step reasoning, and real-time execution logs.

Fig 9. Multimodal GUI agent demonstrating automated task execution on Yelp from natural language instructions.
4.2 Executing Complex Filtered Search Queries
Building directly on those basic interface interactions, let’s explore how the autonomous multimodal GUI agent tackles web browsing tasks like running complex filtered queries on Yelp. Here, the split-screen view shows the agent carrying out the search for Thai restaurants in Pittsburgh. If you check the left side of the screen, the web browser displays the search interface, query results, and interactive map views. Meanwhile, on the right side, the command-line terminal tracks the agent’s observation parsing, planned keyboard and mouse interaction commands, and its step-by-step chain-of-thought reasoning.

Fig 10. Split-screen interface displaying Yelp search on the left and terminal reasoning on the right.
4.3 Real-World Multi-Attribute Decision-Making on Yelp
Now, here’s where we move from initial search filtering straight into real-world multi-attribute decision-making on Yelp. The split-screen demonstration interface highlights the multimodal GUI agent handling this detailed online search task in real time. Across the left panel, the web browser navigates through restaurant results in Pittsburgh, focusing closely on specific queries such as Pusadee’s Garden. At the very same time, the right panel captures the terminal output containing the agent’s real-time reasoning, observations, step-by-step thinking, and executed browser actions.

Fig 11. Multimodal GUI agent executing an online Yelp search with terminal observations and actions.
5. Interactive Software Engineering and Visual Debugging with OpenHands
5.1 Adapting Base Models for Autonomous Application Workflows
Over the past few years, we have seen frameworks and agents of this kind truly taking off across the industry. Prominent examples include Manus, OpenAI’s Operator, Claude’s computer use agent, and several other cutting-edge tools.
We have also done a significant amount of work here at CMU on building coding agents. Graham’s group in particular has done extensive work in this area, especially through OpenHands, an effort focused on developing capable open-source agents. To illustrate how this operates in practice, Graham shared an impressive demo video showing an agent that writes code and then actively interacts with the application it has built directly inside a browser to test everything out.
5.2 Interactive Code Generation and Local Development Server Setup
In this demo, one particularly nice feature of the setup is that you can watch what the agent is doing in real time while it is working. As we observe its step-by-step workflow, it actively edits the index file, refines the style file to ensure the application looks visually appealing, creates a requirements.txt file to streamline dependency installation, and puts together a comprehensive readme. Across each of these steps, you can see that the agent is consistently following solid software engineering best practices.

Fig 12. What the Coding Agent Produced

Fig 13. OpenHands start screen with suggested tasks before the agent begins building the app
At this stage in the process, the agent proceeds to start up the application, running through its initialization operations very, very quickly. What you can do is instruct it to run the application in the background and then test it out directly using its integrated browser. When you issue that command, the agent executes it immediately without missing a beat.
It actively checks whether the local app is running and navigates directly to its work within the browser environment. While inspecting the system logs, the agent quickly notices that the required port was already in use by another process and immediately tries out a fix. That is really the beauty of modern autonomous agents: they can debug their own operational problems and make sure that things work properly without demanding constant human oversight.

Fig 14. OpenHands coding agent demonstrating visual debugging and code generation
5.3 Automated Frontend Testing and Visual Browser Feedback
Once the local server is successfully running, we can watch the system in action as it browses directly to the target application, rendering our To Do List app right on the interface. Upon observing that the web app is up and running as expected, it immediately moves forward to test the application’s actual interactive functionality. We can track its actions directly as it navigates the user interface, types in the sample entry “Buy Groceries”, and clicks the submission button.
In real time, the agent actively puts the application through its paces. It exercises interactive behaviors across the board, such as testing the item delete functionality, verifying dynamic updates, and confirming other related UI features.

Fig 15. OpenHands automated coding agent performing visual debugging on a Todo List app
Speaking for myself, I am certainly not the best front-end developer around, but the agent proves remarkably adept at exploring the front-end environment. It is fully capable of trying out the interface, evaluating whether the underlying functionality is working as intended, and fixing any encountered issues right on the fly. In fact, during one run, the application actually crashed, and the agent realized that the crash had occurred even before I noticed it myself, autonomously stepping in to diagnose the logs and fix the root problems.
5.4 Visual Debugging Cycles and Error Resolution in OpenHands
Now, here is the important reality check: the development workflow is not as simplistic as merely directing a model to go build an application, seeing if it boots up, and assuming that nothing else needs to be done. In a realistic software engineering scenario, your immediate next step would be instructing the agent to push the verified code directly to GitHub and finalize the remaining project details. That way, you can seamlessly take over the codebase and continue developing the project smoothly.
We are particularly excited to dive deeper into hands-on experiences with building agents firsthand. In an upcoming guide, we will thoroughly explore OpenHands and its software development kit (SDK), walking through step by step how you can construct your own custom agents capable of these autonomous workflows. Furthermore, our team actively conducts research across these exact agentic domains every single day.
6. Conceptual Foundations of Agent Anatomy and Environments
6.1 Formalizing Agents via Environment, State, and Actuators
Let’s begin with a look at what an agent actually is. The core concept of an agent has been around for a long time, and the pursuit of developing agentive capabilities is definitely not just a two- or three-year-old development. In their foundational textbook on artificial intelligence, Russell and Norvig define an agent as anything that can be viewed as perceiving its environment through sensors and acting upon that environment through actuators. That classical framing remains profoundly true for the agents we build today.
Because of this continuity, many of the techniques that were historically developed for agents continue to apply directly to our current systems. Not necessarily all of them carry over without modification, but a significant portion still does. This is especially true for foundational approaches such as search and reinforcement learning.
Before we unpack the details, here is a preview of the four components that make up the anatomy of an agent:
- Environment — the code repository, website, or application the agent is situated in.
- State / Observations — user messages, file contents, web pages, screenshots, or tool results that tell the agent where it currently stands.
- Actions — replying to the user, editing a file, running a shell command, or calling an API, each of which updates the environment.
- Reward — a numeric success signal, such as passing test cases, an LLM-as-a-judge verdict, or positive feedback from the user.
Let’s now look at each of these in turn. To see how this classical formulation pans out for modern systems, consider how an agent is situated in an environment. That environment could be a code repository, an interactive website, or any other application running on your computer. The agent actively interacts with its environment by receiving observations that reflect the current state it occupies.
These observations can manifest in several distinct ways. They might be incoming messages from a human user with whom the agent is interacting, the specific text contents of a file the agent is investigating, the current web page being visited, visual screenshots, or the return values of tools. These tools represent programmatic functions that the agent can execute directly within the environment to mediate its interactions.

Fig 16. What Counts as an Observation

Fig 17. Lecture slide 9: What is an agent?
At each point in time, the agent selects and executes actions. These actions directly update the environment and subsequently alter the state that the agent occupies. Practical actions might include formulating a reply to the user, modifying a file, running an arbitrary command inside a shell, or calling an API endpoint that backs the web pages the agent navigates.
Finally, this continuous interaction loop introduces the critical notion of a reward, which provides a numeric score to define whether the agent succeeded. In software engineering tasks with pre-existing test suites, reward evaluation can be straightforward: if the test cases pass, the agent receives a 1; otherwise, it receives a 0.
For many real-world tasks, however, defining a programmatic reward function is exceptionally difficult. In those settings, we may instead rely on an LLM as a judge to evaluate success, or focus on whether the user is ultimately satisfied with the completed task by leveraging direct user feedback as the reward signal.

Fig 18. Where the Reward Signal Comes From
With this mental model of environment, state, actions, and reward in place, we can turn to the language models that drive the loop.
7. Extending Autoregressive Models: Actions as Generative Tokens
7.1 Autoregressive Next-Token Prediction and Chain-of-Thought
You are probably already familiar with how standard language models operate through next-token prediction. At each discrete step in this autoregressive process, the model predicts a probability distribution over the potential next tokens in its vocabulary. We then sample from that distribution—or select a token using an alternative decoding strategy—append that chosen token back into the model’s context window, and repeat the cycle all over again.

Fig 19. Lecture slide 10: Recap: language models
In recent work, we have seen substantial performance improvements on complex reasoning tasks simply by introducing chain-of-thought techniques. As you might recall, a chain of thought works by predicting an extended sequence of tokens that do not directly represent the final answer we ultimately want. Instead, these intermediate tokens act as a scratchpad for the model’s intermediate reasoning steps, providing crucial context that the model can then condition on to reliably arrive at and predict the final answer.
However, it is important to emphasize that this framework remains an entirely non-agentive setting. The model is merely interacting with and responding to static prompts that you provide to it in isolation. This raises a fundamental question: how do we transition from this passive next-token prediction paradigm to building models capable of actively taking actions in the real world?
7.2 The Toolformer Paradigm: Emitting Action Tokens
To make that leap into an agentive setting, we can enable the language model to actively use tools to act directly in its environment. The underlying mechanism is surprisingly straightforward: the model reads in token sequences that represent which tools are available, along with token sequences containing the results returned from prior tool executions. It then generates new token sequences that explicitly name the chosen tools, which serves to invoke them.

Fig 20. Lecture slide 13: From LLMs to agents: actions as tokens
This setup establishes a critical conceptual distinction between the language model itself and the execution framework wrapped around it. At its core, the language model is strictly an engine that processes tokens—it operates purely on tokens in and tokens out. Later on, we will also explore multimodal language models that ingest images or other modalities as tokens. But the language model itself never executes actions directly; instead, an external harness takes full responsibility for coordinating and executing those tool calls, using the language model as the driving engine.

Fig 21. Model vs. Harness: Division of Labour
8. Structured Tool Interfaces and Schema Definitions
8.1 Defining Tools Through Typed JSON Schemas
The primary mechanism we use to accomplish this interaction is through tools. A tool is situated directly within the environment, providing an interface that allows the language model to interact with that external world. You can think of this as functioning very much like an API. For instance, if we want our model to be capable of looking up current weather conditions to answer a user’s inquiry, we might take an existing API such as weather.gov and wrap it in a dedicated interface that enables the model to call it directly. While there are a bunch of different ways to represent this interface, one prominent and widely adopted example is the OpenAI tool specification format.

Fig 22. Tool definition
Under this tool specification, each tool—which is also referred to as a function that the model can call—is defined by three core components: its name, a natural language description explaining what it does, and a structured schema describing the parameters it accepts. Let’s look at a concrete example: imagine a coding agent operating inside a software codebase. That agent might utilize a tool named read_file to inspect files in the repository where it is actively working. In this case, the tool’s natural language description explicitly states that it is used to “Read a UTF-8 file.”
Beyond simply naming and describing the tool, we also have to inform the model about what arguments it can pass to the function through a structured representation of its parameters. Using a JSON schema, we define the parameters as an object with specific properties, such as a property called path of type string. Furthermore, we mark path in the required array so the model knows unambiguously that it must supply a valid file path string to open the file.
The underlying specification looks like this:
{
"name": "read_file",
"description": "Read a UTF-8 file.",
"parameters": {
"type": "object",
"properties": { "path": { "type": "string" } },
"required": ["path"]
}
}
Once we have defined this underlying specification with its name, natural language description, and structured JSON parameter interface, the information must actually be fed directly into the model. The model receives and views the specification through an applied template, which transforms the structured JSON object into a concrete sequence of text that is tokenized and read by the model. After apply_chat_template, the model sees it as text:
<tools>
{"name":"read_file","description":"Read a UTF-8 file.",
"parameters":{"type":"object","properties":{"path":{"type":"string"}},
"required":["path"]}}
</tools>
8.2 Injecting Tool Signatures via System Prompts and Templates
In practice, the model becomes aware of the tools that it can use to interact with the environment by having tool signatures injected directly into its context window. This injection is typically formatted in JSON and enclosed within <tools> tags — the same <tools> block shown above is what the model actually reads.
Now, here’s the crucial part: simply providing the raw JSON signature in the context is not always enough on its own. To successfully operate with these definitions, the model must either be specifically trained or carefully guided with few-shot examples so that it understands exactly how to invoke and use these tools when generating responses.
9. Tool Invocation Syntax and Chat Template Integration
9.1 Generating Structured Tool Calls and Parsing Output
To actually interact with the external world and accomplish meaningful work, an agent needs to call external tools. In practice, the model accomplishes this by emitting structured tool calls. Instead of simply outputting freeform text, the language model generates structured text that explicitly specifies the name of the function to be invoked along with all of its required arguments.
When you use a standard JSON-based representation inside an API schema, this invocation is structured as a dedicated assistant message containing rich metadata about the call.

Fig 23. Lecture slide 12: Tool calls and results
In this API-driven approach, the model produces a payload that looks like this:
{"role": "assistant", "tool_calls": [{
"id": "call_7",
"name": "read_file",
"arguments": {"path": "test.py"}
}]}
Alternatively, the model can generate special markup tags directly within its raw text stream to delimit the function call and its parameters. In delimiter-based setups, the text generated by the model might look like this:
<tool_call>
{"name":"read_file","arguments":{"path":"test.py"}}
</tool_call>

Fig 24. Three Ways to Express a Tool Call
Now, here is an interesting alternative: rather than relying strictly on JSON schemas or markup blocks, we can actually have the model write code directly. For example, the model can generate bash commands and scripts, or if the environment’s tools are exposed as standard Python functions, it can simply call those functions using native Python syntax. As we will explore in detail in a couple of lectures, generating code directly is actually a remarkably effective approach—often proving more effective than forcing the model into JSON schemas—though each strategy comes with its own unique trade-offs.
When the model issues a tool call, the external execution environment intercepts the command, executes it, and captures the resulting output. For instance, suppose the model requests to read a test file like test.py. That file might contain the implementation of an addition function, such as def add(a, b): …. The execution environment takes the resulting output and returns it within a content field—either packaged inside a dedicated tool role message or wrapped inside a <tool_response> tag. This payload is then rendered back to the model through the underlying chat template. The model ingests these newly returned tokens straight into its context window, allowing it to inspect the environment’s state and decide what to do next.
9.2 Formatting Tool Calls and Observations within Chat Templates
Let’s look a little closer at how these invocations are rendered under the hood in real-world chat templates. Chat templates are responsible for taking high-level conversation messages and formatting them into the exact sequence of tokens the model expects, handling assistant tool calls via either structured API schema objects or explicit text delimiters.
In structured message formats, the model’s action is captured cleanly as a structured object:
{"role": "assistant", "tool_calls": [{
"id": "call_7",
"name": "read_file",
"arguments": {"path": "test.py"}
}]}
Conversely, many open-source chat templates prefer inline markup tags that flag function calls directly in the generated token stream. Under that inline markup syntax, the model’s request is wrapped between custom structural boundaries:
<tool_call>
{"name":"read_file","arguments":{"path":"test.py"}}
</tool_call>
Following the invocation, the returned observation from the environment must be integrated cleanly back into the conversation history. In standard structured message schemas, this is handled via a dedicated tool role entry:
{"role": "tool",
"tool_call_id": "call_7",
"name": "read_file",
"content": "def add(a, b): …"}
Notice how this entry links directly back to the specific identifier call_7 so the model can track which call produced which result. In delimiter-based template designs, that same response is embedded directly within matching functional tags:
<tool_response>
def add(a, b): …
</tool_response>
This provides the execution output directly back to the model. Once these formatting conventions are in place, the agent loop systematically maintains the full sequence of actions and feedback within the model’s context window: every invocation the model emits, together with the observation the environment sends back, is appended to the history in one of the two formats shown above. Appending these observation results directly into the conversation history provides the vital context necessary for the model to take its subsequent action.
10. The ReAct Paradigm: Models Operating Inside an Execution Loop
10.1 Conceptualizing the Model-in-a-Loop Architecture
Let’s look at the broader agent loop and how we actually construct it. A really impactful paper in this space was Toolformer, which demonstrated that you can train models to make tool calls directly. Crucially, it showed that models can condition on the returned results of those calls to make future successful actions much more likely. This provided an early framework for training language models to leverage external tools, generalizing prior LLM training procedures into something far more capable. Once you have this foundational setup in place, you can implement a fully functional agent by running an execution loop.

Fig 25. An agent is a model operating inside a loop
Another influential paper in this domain was ReAct from 2023. The acronym stands for “reasoning and acting” because the model explicitly generates chains of thought before taking actions in its environment. The way this setup works is that you first supply the system with some high-level context. For example, this context might define a broad task, such as resolving issues or solving pull requests on GitHub. Alongside that general context, you pass the specific task to execute, which could simply be an initial prompt or query submitted by the user, establishing the primary inputs feeding into the model.
10.2 Interleaving Reasoning Steps and Environmental Actions
Now, here is how the rest of the execution pipeline fits together. In addition to the context and task description, you provide a list of all available tools the agent is permitted to call, structured using one of the schema formats we explored earlier. You also track an ongoing history of past observations and actions that the agent has generated and encountered throughout the session. At each individual time step, regardless of the current state, the agent conditions on all of that aggregated information to produce a reasoning chain. It uses this scratchpad to think through what it ought to do next, and then emits one or more tool calls to interact with the environment or the user.
To see how this works in practice, imagine the agent decides to trigger a tool call to read a file. That tool call executes directly within the target environment, and the system captures whatever output is returned. We take those execution results, append them straight back into the agent’s interaction history, and recognize that the environment has updated as a direct consequence of the action taken.
Next, we simply repeat the process. The model conditions on this freshly updated history along with the newest observations received from the current step. It generates a new reasoning step, issues another tool call, and repeats this execution cycle over and over.
We can also include specialized tool calls specifically designed to terminate the task. For example, the agent might invoke a tool call that sends a final response to the user or otherwise flags that the job is complete. This mechanism gives the overall architecture a clean way to signal that work has wrapped up. Later, in the planning lecture, we will dive much deeper into how a model determines that it has finished its objective.
11. Minimal Implementation of Coding Agent Loops
11.1 Writing a Minimal ReAct Loop in Python
A good way to internalise this pattern is to implement a ReAct-style loop yourself and point it at a real codebase — for example a chess-engine repository, where a lightweight coding agent has to solve concrete issues inside an actual project. While the precise code you would write there looks somewhat different from the snippet below, this example captures the fundamental pattern cleanly at a high level. The loop initializes an empty message history, formats system and user prompts through configuration templates, and transitions directly into a continuous execution cycle driven by query, step, and action-execution routines.

Fig 26. Lecture slide 15: A minimal ReAct loop
Let’s look at how this basic structure takes shape in Python:
def run(self, task: str = "", **kwargs) -> dict:
self.messages = []
self.add_messages(
self.model.format_message(role="system", content=self._render_template(self.config.system_template)),
self.model.format_message(role="user", content=self._render_template(self.config.instance_template)),
)
while True:
self.step()
def step(self) -> list[dict]:
return self.execute_actions(self.query())
def query(self) -> dict:
message = self.model.query(self.messages)
self.add_messages(message)
return message
def execute_actions(self, message: dict) -> list[dict]:
"""Execute actions in message, add observation messages, return them."""
outputs = [self.env.execute(action) for action in message.get("extra", {}).get("actions", [])]
return self.add_messages(*self.model.format_observation_messages(message, outputs, self.get_template_vars()))
This implementation is taken directly from an elegant, minimal repository called mini-swe-agent. What makes mini-swe-agent so impressive is that it delivers remarkably strong performance on the SWE-bench benchmark, an established evaluation suite for coding agents that we will explore in much greater depth in upcoming lectures. It is a fantastic project to study because it stays exceptionally competitive alongside recent foundation models while keeping its overall architecture delightfully compact. By stripping away extraneous framework bloat, it distills the core mechanics of an agent loop down to its absolute essentials, making it effortless to trace through line by line.
11.2 Parsing Model Output and Driving Bash Execution Engines
Now, let’s walk through how this loop actually operates under the hood. To bootstrap the interaction loop, you first supply initial messages that give the language model crucial context about its role as an agent as well as the exact task it needs to solve. In the code, the run method sets up an empty message container via self.messages = [] and immediately populates it using self.add_messages(). This step formats a system prompt from self.config.system_template alongside an initial user instruction from self.config.instance_template, rendering both through self._render_template(). Once these baseline context prompts are in place, the core engine launches into an infinite while True loop that repeatedly calls self.step() to push the agent forward.

Fig 27. The Loop in Four Methods: run(), step(), query(), execute_actions()
Each individual invocation of step() cleanly encapsulates the standard interaction cycle by executing self.execute_actions(self.query()). Inside the query() method, the agent submits the entire running conversation history stored in self.messages out to the language model, appends the generated model response back onto self.messages, and hands that response back. From there, execute_actions() unpacks the message, carries out whatever tool calls or actions were specified directly within the environment, collects the resulting observations, and appends them back to the history as new observation messages before returning them. While the low-level mechanics of executing environment commands and parsing runtime output are omitted from this concise code snippet, you can inspect the full repository to see how those operational details are wired up.
12. SWE-bench Trajectories and Evaluation Leaderboards
12.1 Deconstructing Real-World SWE-bench Execution Trajectories
It is especially instructive for us to take a close look at example agent trajectories. One great aspect of the SWE-bench dataset, evaluation task, and leaderboard is that they make sample trajectories readily available for us to inspect directly. Taking the time to examine these sample trajectories helps illustrate how the evaluation task unfolds in practice.

Fig 28. Example agent trajectory from SWE-bench
12.2 Software Engineering Leaderboards and Automated Evaluation
The leaderboard below is where these runs are collected: it compares the performance of the submitted systems and is also the place where we can find the trajectories we are inspecting.

Fig 29. SWE-bench leaderboard comparing model performances
Now, let’s drill down into a single run and look at a gpt-oss-120b run to examine how these systems actually operate in practice. Walking through this example gives you a very clear feel for what tool calls look like during execution. Furthermore, it helps illustrate precisely what the model is seeing in its context as it is carrying out this task.
13. Large-Scale Benchmark Run Logs and Failure Diagnostics
13.1 How the Model Is Initialized in a Benchmark Run
Let’s examine how the model gets initialized and conditioned during these benchmark runs. At the very start of the interaction, the model receives context directly through a system message. This system prompt establishes the agent’s baseline role and operational scope. Specifically, the message informs the model that it is a helpful assistant capable of interacting directly with a computer shell.

Fig 30. Benchmark evaluation runs for the gpt-oss-120b agent across repository tasks.
When reviewing benchmark evaluation tables for the gpt-oss-120b agent across repository tasks like Django, Matplotlib, and Sympy, you can track key execution metrics. These logs display run hashes, statuses, instance identifiers, total API call counts, monetary cost, and binary resolution scores of 0 or 1.

Fig 31. What a Benchmark Run Log Records
13.2 Accessing Shared Benchmark Collections and Task Contexts

Fig 32. Interface view of an automated agent evaluation dashboard with run listings.
Through the automated agent evaluation dashboard, we can access run listings, monitor execution statuses, and navigate across task suites. When drilling down into individual trajectories, the Docent trajectory-inspection tool from Transluce provides a detailed transcript viewer. This view lays out an interaction block minimap, the issue description, and the explicit system instructions for bash execution.

Fig 33. Transcript viewer on Docent showing an agent run trajectory on SWE-bench.
While the overarching system setup is very generic, the interaction additionally incorporates a user message that provides concrete context about a particular task, such as resolving a specific pull request. The model then conditions on all of that combined information to guide its execution and address the issue directly.
13.3 Inspecting Run Trajectories and Diagnostic Tooling
At this point, the agent needs to issue its first action. Looking at the interface, you can see the chain of thought trace that it generates, highlighted in blue to denote that it was produced directly by the model. Right alongside this chain of thought, the model produces a tool call. This agent relies on the tool format suggested earlier, which consists of simply issuing bash commands by writing out the specific command that needs to be run.
That command is handed off to the execution environment, which runs it directly. In this particular instance, the returned result is just an exit code—specifically exit code 0, showing that the execution succeeded. From there, we proceed directly to the next turn, where the agent formulates a new thought followed by a new bash command.

Fig 34. Dashboard displaying execution logs and transcripts of an autonomous coding agent.
Under this operational setup, every model response should feature a dedicated thought section explaining the underlying reasoning, and it must contain exactly one bash code block. That bash block must hold exactly one command, or alternatively a chain of commands joined using && or ||. Emitting zero code blocks, multiple bash blocks, or omitting the command entirely results in an immediate failure, and one must never attempt to run multiple independent commands across separate blocks within the same response.

Fig 35. Response Format Rules for the Bash Agent
Additionally, any directory or environment variable modifications are non-persistent because each action runs in an entirely new subshell. To manage state across commands, actions can be explicitly prefixed—for example, via MY_ENV_VAR=MY_VALUE cd /path/to/working/dir && ...—or environment variables can be written to and loaded from files.
14. Summary
Key takeaways:
- An agent is a model plus a harness in a loop. The language model only reads and writes tokens; the harness is what actually executes the tool calls, feeds back observations, and repeats until the task is done.
- Tool calls are just tokens. Tool definitions are rendered into the context with a chat template, and the model “calls” a tool by generating text that the harness parses and executes.
- There are three ways to express a call: structured JSON schemas, inline markup delimiters, or directly writing code (for example bash commands or Python), which is often the more effective option.
- Trust is the bottleneck. How much autonomy we grant depends on the risk of the task, and incidents like context compaction dropping a safety constraint show why human-in-the-loop oversight still matters.
Modern AI agents represent a major shift from passive next-token prediction to autonomous real-world execution. While agents have demonstrated remarkable capabilities—from multi-agent compiler construction to autonomous GUI browsing and front-end debugging with OpenHands—they face critical practical limitations. Real-world incidents, such as context compaction dropping safety constraints, highlight the delicate balance required when calibrating human-in-the-loop oversight and system autonomy across tasks of varying risk.
Technically, agents operate within closed interaction loops governed by environment states, observations, and rewards. By framing actions as tokens, language models can interact with external environments using structured JSON schemas, inline markup delimiters, or direct code execution. Architectures like Toolformer and ReAct interleave reasoning scratchpads with environmental actions, enabling lightweight harnesses like mini-swe-agent to repeatedly query models, execute shell commands, and absorb runtime observations dynamically.
Rigorous benchmarking platforms like SWE-bench and trajectory inspection tools like Docent showcase how modern coding agents navigate complex software repositories. Looking forward, advancing agent capabilities will require solving persistent challenges around execution sandboxing, robust error recovery, and context management to establish truly reliable and secure autonomous workflows.
Source
Lecture: What Are Agents? And How Do They Work? by Daniel Fried and Graham Neubig, Carnegie Mellon University.
Course: CMU 11-768 AI Agents (Fall 2026), Lecture 1, first half (0:00–30:00).
Video: YouTube. Slides: PDF. Course page: cmu-agents.com.
Further reading referenced in the lecture:
- Anthropic, Building a C compiler with a team of parallel Claudes — engineering post, code
- Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools, NeurIPS 2023 — arXiv:2302.04761
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, ICLR 2023 — arXiv:2210.03629
- mini-swe-agent — GitHub, agent loop source
- SWE-bench — swebench.com
- OpenHands — GitHub
- Docent by Transluce — docent.transluce.org, the trajectory-inspection tool shown in the lecture
dataHacker.rs — LLM_log #026: What Are AI Agents and How Do They Work? — From Tool Calls to the ReAct Loop. By Vladimir Matic.