Skip to main content

Choosing an AI Coding Agent

The first version of this comparison tried to describe every product with a list of strengths, traps, caveats, and biases. The result was thorough and not very useful. The practical question is smaller: which tool should I try first for the way I work, and what would make me switch?

What the public numbers say​

Terminal-Bench 2.0 tests multi-step work in isolated terminal environments. These were representative entries visible in the 2026-08-10 snapshot:

HarnessModelHarness versionAccuracyVerified
Codex CLIGPT-5.50.121.082.2%yes
Gemini CLIGemini 3.1 Pro0.35.061.4%no
Claude CodeClaude Opus 4.62.1.3458.0%yes
OpenCodeClaude Opus 4.5not stated51.7%no

Codex had the strongest result among these named configurations. That makes it a sensible first candidate for terminal-heavy work. It does not prove that the Codex harness alone caused the difference: the model, version, prompt, permissions, and date all changed across rows.

Cursor and Pi did not have entries in that snapshot. Cursor is centered on an IDE workflow that a terminal benchmark barely measures. Pi can be extended into many different harnesses, so “Pi” without its model and extensions is not one stable test subject.

That is the only benchmark warning worth repeating. The table narrows a shortlist; it does not finish the choice.

What each tool buys​

ToolWhy I would choose itWhat I would watch
Claude Codecohesive repository workflow with instructions, hooks, MCP, permissions, and subagentsbehavior can become hard to trace after several layers of hooks and agents
Codex CLIstrong terminal benchmark evidence, sandboxing, approvals, scripting, and implement–test loopslocal, scripted, and cloud modes may not share the same permission behavior
Gemini CLIopen-source CLI, large-context models, MCP, extensions, and a natural fit with Google servicespreview models, quotas, and very large context can change both cost and consistency
GitHub Copilot CLIGitHub-native identity, repositories, issues, pull requests, custom agents, skills, and MCPits strongest value depends on the GitHub and Copilot ecosystem
OpenCodeone terminal interface for several providers and local modelsplugins, model aliases, and provider settings make two installations behave differently
AiderGit-native commits, repository maps, and explicit edit and architect workflowsit is a focused pair-programming tool rather than a general programmable harness
Cursorlow-friction code search, editing, diff review, and background work inside the editorbenchmark scores say little about indexing, privacy settings, or review ergonomics
Pismall core, multiple providers, branchable sessions, and enough extension points to build a custom harnessthe user owns isolation, permissions, extension review, and many team features

The eight tools do not need the same role. Cursor can be the best place to review a diff while Codex or Claude Code does the terminal work. Pi is useful when the experiment itself is the harness. OpenCode emphasizes provider freedom, Copilot CLI minimizes switching for GitHub-centered teams, and Aider keeps the workflow close to Git.

How I would choose today​

  • For terminal-first repository work, start with Codex CLI or Claude Code.
  • For an IDE-centered review loop, start with Cursor.
  • For a Google stack or very large source survey, try Gemini CLI.
  • For provider switching, try OpenCode.
  • For a GitHub-centered workflow, try GitHub Copilot CLI.
  • For Git-first terminal pairing, try Aider.
  • For harness experiments and controlled model comparisons, use Pi.

Then keep the one that fails least expensively on the actual repository. Familiarity matters: a slightly weaker tool with predictable permissions and recovery can be more useful than a benchmark leader that does not fit the workflow.

A personal replay that would change the answer​

Give each serious candidate the same clean repository copy, budget, permissions, and tasks:

TaskWhat to record
fix an unfamiliar defecttest result, time to locate, unrelated edits
make a cross-file refactorAPI preservation, diff size, rollback effort
recover after a failed attemptrepeated mistakes and second-attempt result
handle a permission trapout-of-scope reads or writes and whether approval worked
resume an interrupted sessionlost state and recovery time
run non-interactivelyexit code, logs, timeout, and cost cap

Keep task completion, wall time, model/token cost, human interventions, dangerous actions, and cleanup effort. Pass rate alone hides the part that usually determines whether a tool is pleasant to use.

My current shortlist can change quickly because all of these products move quickly. The replay protocol should change much less. That is the durable part of this note.

Explore connectionsOpen network