CodeAct

CodeAct

CodeAct is a principle for AI agents in which the model does not describe its actions as text or a list, but formulates them directly as executable program code. This makes the agent more flexible and reliable, because code is more precise than natural language and its steps can easily be combined and reused.

When an AI is supposed to complete a task — for example renaming a file or analyzing a spreadsheet — it somehow has to describe what it wants to do. In many systems, this happens via a fixed list of commands: “Step 1: open file”, “Step 2: calculate value”. CodeAct takes a different approach. Instead, the model writes actual program code, usually in Python, and has it executed directly. The result comes back, and the model writes the next piece of code — and so on until the task is complete.

Why code is better than a command list

A fixed command list is rigid. If the developer hasn’t anticipated certain actions, the agent cannot perform them. Code doesn’t have this limitation: anyone who can write Python can use it to describe almost any process — loops, conditions, function calls.

There’s also a practical advantage: multiple steps can be combined into a single function. Instead of sending ten individual commands one after another, the agent writes a compact routine. This reduces communication between the model and the environment and makes the process faster and more robust.

Code is also unambiguous. Natural language leaves room for interpretation — “delete the old files” could mean many things. A line of Python, on the other hand, does exactly what it says. Errors occur earlier and are easier to find.

How a CodeAct agent works

The agent receives a task in natural language. It translates this into a short snippet of code and sends it to a runtime environment — a safe space in which code is allowed to execute. The environment returns the result or an error message.

The model reads the result and decides: is the task complete? If not, it writes the next code snippet. This cycle continues until the goal is reached or the agent determines that it cannot proceed any further. This gives the agent genuine memory across multiple steps, since variables from earlier code remain available in the runtime environment.

An important difference from purely language-based agents: CodeAct agents can actually compute things. If a user asks by what percentage a number has increased, the agent doesn’t produce an estimate — it carries out the calculation and returns the exact result.

CodeAct in products and research

The principle appears under various names in real-world products. OpenAI’s ChatGPT uses very similar logic with its “Code Interpreter” (formerly Advanced Data Analysis): the model writes Python code, executes it, and displays the result — such as a chart or an analysis of an uploaded spreadsheet.

The term CodeAct itself was coined by a research team at the University of Illinois in a 2024 study. The researchers compared agents that work with text commands to those that write code — and found that CodeAct agents performed significantly better on complex, multi-step tasks. The study was quickly picked up by industry.

Agent systems such as OpenDevin or various open-source frameworks for autonomous software development also build on CodeAct. Anyone who hears in the future about an AI agent that independently writes, tests, and improves code is likely dealing with some variant of this principle.

Related Products

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.