Multi-Agent Coordination: Subagents, Worktrees, Reviewers, and Event Logs
Handing one task to several agents is not hard because of how many agents to start. It is hard because of three questions: what each agent can see, what it can change, and who combines the results and verifies them. How a single agent proposes actions inside a checked loop is covered in Harness and Bounded Agent Loops; this page covers what changes when several of those loops run at once.
The main case is Meta's Muse Code, launched in beta on August 5, 2026 and out of beta since August 31, which was designed around coordinating agents from the start. OpenAI's Agents API, Claude Code, and Antigravity serve as comparisons. Product details were checked on 2026-10-01.
Four Roles
Multi-agent systems usually combine four roles with different responsibilities:
- a coordinator that splits the task, assigns the pieces, combines the results, and answers for the final outcome;
- bounded workers that complete only their own piece and return something checkable;
- background observers that take no task of their own and watch one quality dimension, offering suggestions;
- explicit reviewers that examine a result once it is finished.
Muse Code has all four. The lead session spawns subagents, each with one bounded task. A set of background observers handles memory recall, skill recall, goal tracking, and verification; in Meta's words, these background agents "remain active throughout each session, rather than being spawned for individual tasks" (launch post). Observers can only advise: "it proposes, a reconciler decides," and only an accepted proposal reaches the main agent's next turn. Review therefore neither interrupts the main agent nor bypasses it to edit code directly.
Three Boundaries to Keep Apart
"Give each subagent its own environment" sounds like one decision. It actually involves at least three separate boundaries:
The first two are the easiest to confuse. Separate contexts do not mean separate files: OpenAI's documentation states that the coordinator and its subagents share one environment's filesystem, and "creating a subagent does not create another environment." Separate files do not mean separate everything, either. According to the Git manual, a worktree keeps only per-worktree files such as HEAD and index to itself and shares the rest of the repository, including the object store and, by default, the repository configuration. Two workers that each start a development server in their own worktree can still fight over one port or write to one test database.
Permissions have a further trap. Claude Code's documentation warns that disabling only Write and Edit does not stop file writes, because Bash remains available. Judge a tool restriction by the effects that remain possible, not by tool names. The wider relationship between permissions, approvals, and sandboxes is covered in Agent Security.
One Bounded Parallel Change
A fictional example: a service needs a new API field, an updated frontend display, and matching documentation, all at once.
- The coordinator fixes the interface first. The field name, type, and default go into every worker's task description. This is the only decision the workers must share, so it has to be settled before assigning work.
- Assign directories by whether a worker writes code. The backend and frontend workers each get a worktree; a worker that only gathers information stays in the shared checkout, which is what Muse Code's documentation recommends.
- Workers return evidence, not conclusions. Each returns commits and its own test results, not a statement that it is done.
- The coordinator integrates once and runs one shared acceptance check. Only tests on the merged result catch two halves that each pass but disagree with each other.
- Conflicts go back to the coordinator. If the interface must change, the coordinator decides and reassigns; workers do not edit each other's code.
In this site's view, step 1 deserves the most attention. In June 2025, Cognition argued from examples that parallel agents make implicit decisions that may not fit together, so every action should be "informed by the context of all relevant decisions made by other parts of the system" (post). That is an engineering argument built on cases, not a controlled comparison, and it does not measure how often the failure occurs, but it describes one way split work can go wrong.
Messages, Cancellation, and Capacity
When agents pass messages, establish where a message comes from and what it is allowed to do. Muse Code's session messaging connects independent interactive sessions of the same user on the same macOS or Linux machine, carrying plain text of up to 8,192 bytes; a received message is treated as "unverified agent-provided data" and does not carry the sender's authority. Antigravity CLI 1.2.9 added @<subagent> <message> syntax for writing directly to a subagent's conversation, and the Antigravity 2.0 desktop app added live subagent cards and direct messages in late September.
Cancellation is usually not instantaneous. Muse Code documents it as cooperative: a cancelled subagent that has not reached a checkpoint keeps running, and one in the middle of a write finishes that write. After stopping every agent, check the actual state of the working directories.
Capacity is bounded too. By default one Muse Code agent tree runs at most eight agents at once, including the lead, while an unconfigured ultra root uses 64; the limit can be set from 1 to 64, and grandchildren share the same capacity. When the tree is full, a new spawn request is rejected.
What an Event Log Can Recover
Muse Code records every run as an append-only event log covering model calls, tool calls and their results, and approval decisions. After an interruption, resuming means rebuilding the conversation from that log. The log also separates two kinds of side effects: those confirmed as complete count as done, while an effect that was announced but never confirmed makes the agent check the actual state before deciding whether to retry (audit and resume).
"Replayable" needs a precise meaning. Muse Code's deterministic replay rebuilds the recorded context from events and "does not call a model, run a tool, or use the network." That makes it useful for auditing and regression tests, but it does not guarantee that an external side effect happens only once. Preventing a retry from charging a card twice or sending a message twice still depends on the idempotency keys and state checks described in Long-Running Agents.
When Several Agents Are Worth It
The public evidence points one way: independent tasks split well, tightly coupled tasks do not. In June 2025 Anthropic reported that its multi-agent research system outperformed a single Opus 4 agent by 90.2% on an internal research evaluation, while using about 15 times as many tokens as chats, and that it suits tasks with heavy interdependence less well (post). Those are vendor-internal results that apply to the research tasks it measured.
Review has a cost as well. On the Sonnet 5.5 release page, Anthropic explains that at maximum effort the model more often launched a code review split across many subagents; in two cases Cognition examined, that led to a timeout or to edits beyond the task's scope, and the FrontierCode score fell below the next-lower setting. That establishes a failure mechanism, not how often it occurs.
This site's reading of that evidence: get the task working with one agent first, then split off the parts that are genuinely independent and can be accepted separately, such as reading different documents, investigating problems in different modules, or running unrelated experiments in parallel. Tightly coupled changes, designs that depend on many shared implicit decisions, and repeated reviews tend to get slower and more expensive when split. Compare on the same tasks by completion time, tokens, and rework, not by how many agents ran at once.
How Current Tools Handle It
Checked on 2026-10-01:
Terminal persistence is yet another layer. Herdr restores several agents' terminals after a restart, but it does not split tasks, merge results, or isolate working directories. When choosing tools, first decide which of these layers you actually need.