Choosing an AI Coding Agent
The first version of this comparison tried to describe every product with a list of strengths, traps, caveats, and biases. The result was thorough and not very useful. The practical question is smaller: which tool should I try first for the way I work, and what would make me switch?
What the public numbers say
Terminal-Bench 2.0 tests multi-step work in isolated terminal environments. These were representative entries visible in the 2026-08-10 snapshot:
Codex had the strongest result among these named configurations. That makes it a sensible first candidate for terminal-heavy work. It does not prove that the Codex harness alone caused the difference: the model, version, prompt, permissions, and date all changed across rows.
Cursor and Pi did not have entries in that snapshot. Cursor is centered on an IDE workflow that a terminal benchmark barely measures. Pi can be extended into many different harnesses, so “Pi” without its model and extensions is not one stable test subject.
That is the only benchmark warning worth repeating. The table narrows a shortlist; it does not finish the choice.
What each tool buys
The eight tools do not need the same role. Cursor can be the best place to review a diff while Codex or Claude Code does the terminal work. Pi is useful when the experiment itself is the harness. OpenCode emphasizes provider freedom, Copilot CLI minimizes switching for GitHub-centered teams, and Aider keeps the workflow close to Git.
How I would choose today
- For terminal-first repository work, start with Codex CLI or Claude Code.
- For an IDE-centered review loop, start with Cursor.
- For a Google stack or very large source survey, try Gemini CLI.
- For provider switching, try OpenCode.
- For a GitHub-centered workflow, try GitHub Copilot CLI.
- For Git-first terminal pairing, try Aider.
- For harness experiments and controlled model comparisons, use Pi.
Then keep the one that fails least expensively on the actual repository. Familiarity matters: a slightly weaker tool with predictable permissions and recovery can be more useful than a benchmark leader that does not fit the workflow.
A personal replay that would change the answer
Give each serious candidate the same clean repository copy, budget, permissions, and tasks:
Keep task completion, wall time, model/token cost, human interventions, dangerous actions, and cleanup effort. Pass rate alone hides the part that usually determines whether a tool is pleasant to use.
My current shortlist can change quickly because all of these products move quickly. The replay protocol should change much less. That is the durable part of this note.