Skip to main content

Bounded Agent Loops

A tool loop lets a model observe a result, choose another action, and continue. That is useful for debugging and multi-step work, but it also turns one bad assumption into a sequence of writes, retries, and growing context.

The important split is simple: the model proposes; the harness decides what may run and whether the result counts as success.

Common loop shapes​

PatternUseful forTypical failure
in-session tool loopinteractive investigationcontext fills with old logs and repeated guesses
fresh-context task loopwork that spans many model windowsthe task ledger drifts away from the repository
implement–verify looprepairs with meaningful teststhe agent satisfies a weak test instead of the real contract
evaluator–optimizer looprepeated drafting or measurable tuningworker and evaluator share the same blind spot
benchmark loopoptimizing a scalar measurementnoise is mistaken for improvement

The pattern matters less than the stop conditions around it.

What the harness must own​

Budget and stopping​

Set limits before the run: steps, wall time, model use, tool use, and allowed files or services. Stop when a limit is reached, when the user cancels, when required evidence is unavailable, or when repeated failures show that the current strategy is not improving.

A repeated-call breaker can help, but its threshold should match the operation. Retrying a read is not the same as retrying a payment or deployment.

Tool validation​

Validate tool names, argument schemas, paths, permissions, and side effects outside the model. Prefer direct argument arrays to shell interpolation, while still checking dangerous flags and path semantics. Treat tool output as untrusted input rather than as new instructions.

Verification​

A model's final sentence is not a completion signal. The verifier may be a test suite, schema check, exact artifact comparison, browser assertion, or human review. It should test the intended behavior and be independent enough that the agent cannot make it green by deleting or weakening the check.

Durable state​

Keep the complete transcript and raw tool output outside the active prompt. For long work, save a small task record with:

  • objective and allowed scope;
  • files or artifacts changed;
  • checks already run and their results;
  • failed approaches worth avoiding;
  • open question and next action.

This record helps a fresh context resume the job, but it should be updated from observed results rather than from the model's confidence.

Cancellation and cleanup​

Long-running tools need timeouts, process tracking, and cleanup. Cancelling the UI should cancel the underlying subprocess or background job as well. Otherwise an apparently stopped agent can continue writing to disk.

A minimal control loop​

The following is pseudocode; tool and model interfaces vary by harness.

for step in range(max_steps):
if cancelled() or budget.exhausted():
return STOPPED

context = select_working_context(state)
proposal = model.propose(context)

if proposal.is_final:
return verify_final(proposal, state)

checked = policy.validate(proposal.tool_call)
if not checked.allowed:
state.record_denial(checked.reason)
continue

result = tools.execute(checked.call, timeout=checked.timeout)
state.record(result)

if repeated_failure(result, state):
return NEEDS_REVIEW

Real implementations also need exception handling, idempotency rules, concurrency control, and redacted logs. The pseudocode is useful because it shows where those concerns belong: around the model call, not inside a prompt asking the model to be careful.

Signs of real progress​

A shrinking failure set, a new discriminating test, a smaller justified diff, or a resolved ambiguity is progress. Rephrasing the same diagnosis, adding abstractions without evidence, or running longer is not.

The best loop is usually the least autonomous one that can finish the task reliably. Add parallel workers, planning layers, and self-critique only when a measured failure calls for them.

The general control pattern above can be compared with the versioned Pi implementation below. Configuration and extension ownership are covered in the Pi harness case; recovery across processes is covered in long-running agents.

How a Pi Agent Loop Actually Runs​

The Pi architecture video depicts an Agent Loop as a model calling a tool and receiving the result. That direction is correct, but Pi's current agent-loop.ts and AgentSession show three layers: the session prepares a request, the low-level loop coordinates the model and tools, and the session layer handles persistence, compaction, and retries.

Much happens before the model call​

User input does not go straight to the model. AgentSession.prompt() first lets Extensions handle or transform it, then expands an explicitly invoked Skill or Prompt Template. If the Agent is already running, a new message must enter the Steering or Follow-up queue instead of appearing in the middle of a tool call.

The session then checks the model and authentication and decides whether old context needs compaction. A before_agent_start event may also append custom messages or alter the system prompt for this run. Only after these steps does the user message enter the low-level Agent Loop.

The low-level implementation has two loops​

The inner loop performs successive model calls. Before another call, Pi can refresh the system prompt, tools, model, and thinking level, and inject Steering messages at a safe turn boundary. It then transforms the context into the message format expected by the active provider and streams an Assistant Message.

If that response contains tool calls, Pi executes them, turns every result into a toolResult message, and calls the model again. The inner loop stops only when there are no tool calls and no waiting Steering messages.

The outer loop handles Follow-up messages. It retrieves them only after the current tool chain and Steering queue are exhausted, then enters the inner loop again. Steering therefore corrects direction before the next model call; Follow-up continues after the current task was about to end.

In Pi 0.84.3, the keybinding reference assigns Enter to Steering and Alt+Enter to Follow-up; Windows and WSL use Ctrl+Q for Follow-up by default. Alt+Up retrieves queued messages into the editor (Alt+Q by default on Windows and WSL), while Escape aborts the run and restores messages that were not processed.

A requested tool does not run immediately​

Pi first checks that the tool exists, prepares its arguments, and validates them. beforeToolCall can block execution. After the tool returns, afterToolCall can rewrite the result or turn a success into an error. Tool calls may run in parallel, while tools marked Sequential, or a global sequential setting, force ordered execution.

There is another useful guard. If the model response is truncated by its output limit, Pi does not execute tool calls whose arguments merely appear valid. It marks every call in that response as failed and asks the model to submit them again. Otherwise, incomplete JSON could happen to validate and trigger the wrong action.

A stopped Loop is not a successful task​

The low-level Loop stops when no tool calls, Steering, or Follow-up messages remain, or when an error, cancellation, or external stop condition ends it. Pi then emits agent_end.

That means only that the Agent will not make another model call. It does not establish that the code is correct, the files match the request, or the user's goal is complete. Tests, schema checks, browser assertions, and human acceptance remain outside the Loop. A model saying "done" and a task being done are different events.

Compaction and retries belong to the session layer​

Pi's compaction mechanism is more than checking a Token count once per turn. It checks before a new Prompt, before the next model call after tool results, and after a low-level Agent Run. When the threshold is crossed, older messages become a summary while recent messages remain. A recoverable context overflow can trigger compaction followed by one retry.

Messages are also written to a tree-shaped JSONL Session. That tree stores conversation and tool results, not filesystem snapshots. Switching a session branch does not undo files already written to disk.

The Loop worth remembering​

An Agent Loop is not "let the model think until it finishes." It is an explicit handoff: the session prepares context, the model proposes an action, the harness validates and executes a tool, and the result returns to context. The session layer persists, compacts, retries, and queues. External acceptance finally decides whether the run actually completed the task.

Explore connectionsOpen network