Coding agents and harnesses, compared

An agent is a model plus a harness. Six reach the same ceiling — a pull request opened for you: Claude Code, Codex CLI, opencode, Cursor, Devin and GitHub Copilot, which resolves agent identity most cleanly. opencode, Goose, Aider and Cline win on BYOK and licence; Cursor and Windsurf on ergonomics. Nothing here merges or deploys unsupervised.

Last verified
2026-01-15 (6mo ago)
Source confidence
35%
Re-verified
every 14 days
Tools
17
Fields
26
17 tools · verified 6mo ago
opencodeProvider-agnostic open-source TUI with a client/server split and LSP awareness.84.8

Features

Surface
CLIweb
Default model
None — you choose a provider at first run
Model breadth
any provider / router
Repo indexing
AST / LSP repo map
Plan mode
Yes
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
4
MCP
Yes
Subagents
Yes
Hooks
Partial
Skills / commands
Yes
Background agents
Partial
Opens PRs
Yes

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
MIT

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
25,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
The open-source AI coding agent built for the terminal.

Timeline

First release
2025-06-01
Claude CodeAnthropic's terminal-first agent, with IDE, web and GitHub Action surfaces.80.0

Features

Surface
CLIIDE extensionwebGitHub-native
Default model
Latest Claude Opus/Sonnet, switchable with /model
Model breadth
one vendor + custom endpoint
Repo indexing
agentic grep
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
No

Agentic

Where it stops
4
MCP
Yes
Subagents
Yes
Hooks
Yes
Skills / commands
Yes
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
proprietary

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$20 /seat/mo
Token markup
No

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
Anthropic's coding agent, in your terminal and wherever else you work.

Timeline

First release
2025-02-24
Codex CLIOpenAI's open-source Rust agent, with the strongest local sandboxing on this page.76.0

Features

Surface
CLIIDE extensionwebcloud delegate
Default model
Latest GPT-5 Codex model
Model breadth
one vendor + custom endpoint
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
4
MCP
Yes
Subagents
No
Hooks
Partial
Skills / commands
Partial
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$20 /seat/mo
Token markup
No

Traction

GitHub stars
42,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
A coding agent that runs locally, in your terminal.

Timeline

First release
2025-04-16
GooseBlock's Apache-2.0 agent, MCP-native from the start, as both a CLI and a desktop app.75.7

Features

Surface
CLIstandalone editor
Default model
None — you configure a provider on first run
Model breadth
any provider / router
Repo indexing
agentic grep
Plan mode
Yes
Checkpoint & undo
No
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
Yes
Hooks
Partial
Skills / commands
Yes
Background agents
Partial
Opens PRs
No

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
20,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
An open-source, extensible AI agent that goes beyond code suggestions.

Timeline

First release
2025-01-28
CursorVS Code fork with a real codebase index, background cloud agents and a CLI.75.5

Features

Surface
standalone editorCLIwebcloud delegate
Default model
Cursor's in-house Composer model, with frontier models selectable
Model breadth
multi-vendor
Repo indexing
embeddings index
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
No

Agentic

Where it stops
4
MCP
Yes
Subagents
Partial
Hooks
Yes
Skills / commands
Yes
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Partial
Licence
proprietary

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$20 /seat/mo
Token markup
Partial

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
The AI code editor, built to make you extraordinarily productive.

Timeline

First release
2023-03-01
GitHub Copilot (Agent HQ)Copilot's coding agent plus a control plane for running third-party agents on your repos.73.7

Features

Surface
GitHub-nativeIDE extensionCLIweb+1
Default model
Selectable across OpenAI, Anthropic and Google models
Model breadth
multi-vendor
Repo indexing
hybrid index + agentic
Plan mode
Partial
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
Partial

Agentic

Where it stops
4
MCP
Yes
Subagents
Partial
Hooks
Unknown
Skills / commands
Yes
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
vendor-hosted sandbox
BYOK
Partial
Licence
proprietary

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$10 /seat/mo
Token markup
Partial

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
Mission control for every agent working on your repositories.

Timeline

First release
2021-06-29
Kilo CodeOpen-source VS Code agent in the Cline/Roo lineage, with an orchestrator mode and a CLI.72.8

Features

Surface
IDE extensionCLI
Default model
Model breadth
any provider / router
Repo indexing
embeddings index
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Partial
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
Yes
Hooks
Unknown
Skills / commands
Yes
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
9,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
Open-source AI coding agent for VS Code, with everything included.

Timeline

First release
2025-03-01
Gemini CLIApache-2.0 terminal agent from Google with an unusually generous free tier.72.6

Features

Surface
CLIIDE extensionGitHub-native
Default model
Latest Gemini Pro
Model breadth
one vendor
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
No

Agentic

Where it stops
3
MCP
Yes
Subagents
Partial
Hooks
Unknown
Skills / commands
Yes
Background agents
Partial
Opens PRs
Partial

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
75,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
An open-source AI agent that brings Gemini into your terminal.

Timeline

First release
2025-06-25
Qwen CodeApache-2.0 Gemini CLI fork tuned for Qwen3-Coder, with a large free daily quota.66.6

Features

Surface
CLI
Default model
Latest Qwen3-Coder model
Model breadth
one vendor + custom endpoint
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
Unknown
Hooks
Unknown
Skills / commands
Yes
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
13,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
A command-line coding agent tuned for the Qwen3-Coder models.

Timeline

First release
2025-07-01
Grok BuildxAI's coding harness, launched July 2026 — on this page but not yet verified.66.5

Features

Surface
Default model
Model breadth
Repo indexing
Plan mode
Unknown
Checkpoint & undo
Unknown
Headless / CI
Unknown
Local models
Unknown

Agentic

Where it stops
MCP
Unknown
Subagents
Unknown
Hooks
Unknown
Skills / commands
Unknown
Background agents
Unknown
Opens PRs
Unknown

Governance & control

Permissions
BYOK
Unknown
Licence

Pricing

Pricing model
Entry price (/seat/mo)
Token markup
Unknown

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning

Timeline

First release
2026-07-01
ClineOpen-source VS Code and JetBrains agent built around an explicit plan/act split.64.8

Features

Surface
IDE extensionCLI
Default model
None — you pick a provider at setup
Model breadth
any provider / router
Repo indexing
agentic grep
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Partial
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
No
Hooks
No
Skills / commands
Yes
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
50,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
Autonomous coding agent in your IDE, with your API key.

Timeline

First release
2024-07-01
DIY harnessYour own loop over a model API with an agent SDK — the build-it-yourself reference point.59.8

Features

Surface
CLI
Default model
Whichever you wire up
Model breadth
any provider / router
Repo indexing
manual context only
Plan mode
No
Checkpoint & undo
No
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
MCP
Partial
Subagents
Partial
Hooks
Yes
Skills / commands
No
Background agents
Partial
Opens PRs
Partial

Governance & control

Permissions
applies everything
BYOK
Yes
Licence

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
The reference point: what you get for a few hundred lines and an API key.

Timeline

First release
ZedGPL-licensed Rust editor whose agent panel can host Claude Code, Codex and Gemini CLI.58.0

Features

Surface
standalone editor
Default model
None — configure a provider, or sign in for hosted models
Model breadth
any provider / router
Repo indexing
AST / LSP repo map
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
No
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
No
Hooks
No
Skills / commands
Partial
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
GPL-3.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
62,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
The editor for humans and AI, written in Rust.

Timeline

First release
2023-03-22
AiderMinimal terminal pair-programmer with a tree-sitter repo map and automatic git commits.57.0

Features

Surface
CLI
Default model
None — set by --model or environment keys
Model breadth
any provider / router
Repo indexing
AST / LSP repo map
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
3
MCP
No
Subagents
No
Hooks
No
Skills / commands
Partial
Background agents
No
Opens PRs
No

Governance & control

Permissions
prompts every action
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0 /seat/mo
Token markup
No

Traction

GitHub stars
37,000
Terminal-Bench (%)
Bench model

Positioning

Positioning
AI pair programming in your terminal.

Timeline

First release
2023-05-01
WindsurfAgentic IDE built around Cascade, now owned by Cognition after a turbulent 2025.53.5

Features

Surface
standalone editorIDE extension
Default model
Windsurf's in-house SWE model, with frontier models selectable
Model breadth
multi-vendor
Repo indexing
embeddings index
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
No
Local models
No

Agentic

Where it stops
3
MCP
Yes
Subagents
Unknown
Hooks
Unknown
Skills / commands
Yes
Background agents
Partial
Opens PRs
Partial

Governance & control

Permissions
granular allowlist
BYOK
Partial
Licence
proprietary

Pricing

Pricing model
opaque credits
Entry price (/seat/mo)
$15 /seat/mo
Token markup
Yes

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
The agentic IDE where you and the agent stay in flow together.

Timeline

First release
2024-11-13
DevinFully delegated cloud engineer: file a task in Slack or the web app, get a PR back.48.9

Features

Surface
webcloud delegate
Default model
Undisclosed — Cognition-hosted, including in-house SWE models
Model breadth
no model choice
Repo indexing
hybrid index + agentic
Plan mode
Yes
Checkpoint & undo
No
Headless / CI
Partial
Local models
No

Agentic

Where it stops
4
MCP
Unknown
Subagents
No
Hooks
No
Skills / commands
Partial
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
vendor-hosted sandbox
BYOK
No
Licence
proprietary

Pricing

Pricing model
opaque credits
Entry price (/seat/mo)
$20 /seat/mo
Token markup
Yes

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
An AI software engineer you delegate tickets to.

Timeline

First release
2024-03-12
AmpSourcegraph's opinionated agent with no model picker and an ad-supported free tier.44.2

Features

Surface
CLIIDE extensionweb
Default model
Chosen by Amp — there is no model picker
Model breadth
no model choice
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
No

Agentic

Where it stops
3
MCP
Yes
Subagents
Yes
Hooks
Unknown
Skills / commands
Partial
Background agents
Partial
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
No
Licence
proprietary

Pricing

Pricing model
opaque credits
Entry price (/seat/mo)
$0 /seat/mo
Token markup
Yes

Traction

GitHub stars
Terminal-Bench (%)
Bench model

Positioning

Positioning
An agentic coding tool that always runs the best model, chosen for you.

Timeline

First release
2025-05-01
yespartialnounknown◆ measured · ▸ vendor-claimed · ▪ community · · inferredExpand a row for the source and date behind every cell.
Coding agents compared: Claude Code vs Codex CLI vs Cursor vs the rest pricing, compared line by line

Which AI coding agent should I use in 2026?

Start from the framing that makes this category legible: agent = model + harness. The model supplies raw capability. The harness supplies everything that turns capability into finished work — the tool loop, the context strategy, the permission gate, the retry-on-failing-test behaviour, the thing that stops it wandering off after twenty minutes. Hand the same frontier model to two harnesses and you get wildly different completion rates on the same repo. That is why a benchmark number only earns a place on this page when the harness, the model and the benchmark version are all named — and why the Terminal-Bench column below is currently empty rather than populated with the model scores everyone else reprints. A harness score with no harness attached is a marketing artefact, not a measurement.

The field that actually separates these tools is where does it stop? We score it 1–5 on a fixed ladder: suggests → edits files → runs tests → opens PR → merges and deploys. Almost every comparison stops at feature checklists and never asks how far down that ladder the tool goes without a human. The answer is sobering, and it is narrower than the ladder suggests: every scored harness here sits at rung 3 or rung 4, so in practice the column separates "stops at your working tree" from "reaches a pull request" and nothing finer. Claude Code, Codex CLI, opencode, Cursor, Devin and GitHub Copilot reach rung 4 — a branch pushed and a pull request opened through a path the vendor ships. Nothing on this page ships at rung 5. Copilot's coding agent cannot approve its own PR by design; Devin's ceiling is a review request. Anyone selling "autonomous deployment" is describing a CI pipeline you wired yourself, sitting downstream of a rung-4 agent.

There are two camps and they are drifting apart. The vendor-coupled harnesses — Claude Code, Codex CLI, Gemini CLI, Devin, Grok Build — are built against one model family's quirks, so they ship agentic features first: subagents, hooks, sandboxing, cloud runners. Amp belongs beside them for the same practical reason with a different mechanism: it is not coupled to one vendor so much as opaque about which it uses, because there is no model picker and multiple vendors sit behind it. The provider-agnostic ones — opencode, Cline, Goose, Aider, Zed, Qwen Code, Kilo Code — trail on features by a quarter or two but survive vendor pricing changes, run local models, and never leave you renegotiating your tooling because someone re-tiered a subscription. If your team is one procurement decision away from switching model vendors, take the agnostic one and accept the lag — opencode in particular now reaches a pull request through its own GitHub App, so the feature gap is narrower than it was.

The pricing hides more than the features do. Three models are in play: BYOK tokens (you pay the provider, the harness is free — opencode, Aider, Goose, Zed, Cline, Qwen Code), a flat subscription with soft limits (Claude Code on Pro or Max, Codex on ChatGPT plans), and credits (Windsurf, Devin, Amp). Credits are the one to watch. A credit is an abstraction over tokens whose conversion rate the vendor controls and can change, which makes month-to-month cost genuinely unforecastable for an agent that might burn ten times more context on a bad day. Copilot's premium requests are the same problem in different clothing. If you need a number you can put in a budget, a subscription or your own API key are the only two honest answers.

Expect this page to be wrong faster than any other on the site. Our intended cadence here is a full re-verification every fourteen days, and we are not currently meeting it: every dated cell below was last checked on 15 January 2026, so the table is roughly six months old and you should treat pricing and feature cells as leads to confirm rather than as current fact. We would rather show you the dates and the gap than quietly restamp them. In the eighteen months to that pass, Windsurf changed hands once — after an OpenAI acquisition collapsed and Google licensed the technology and hired the leadership, Cognition bought what remained — Aider's release cadence slowed to a trickle, GitHub rebuilt Copilot around a multi-agent control plane, and Grok Build appeared in July 2026 with essentially nothing verifiable published. pi.dev arrived after that pass and is not on the roster below; a row of dashes reads as a bad result rather than as absent evidence, so it stays in prose as one to watch until someone here has actually shipped code with it. Pick tools whose exit cost is low: a harness that reads AGENTS.md, takes your API key and leaves your repo unchanged costs you a weekend to replace. One that owns your credits, your index and your PR flow does not.

The autonomy ladder, rung by rung

Rung 1 is suggests: it proposes a diff or a snippet and you apply it. Rung 2 is edits files: it writes directly to your working tree. Rung 3 is runs commands and tests: it executes the build and test loop and iterates on the failures it sees. Rung 4 is opens a pull request: it commits, pushes a branch and raises a PR through a first-class path. Rung 5 is merges and deploys: it lands work on the default branch or triggers a release with no human in the loop.

The interesting gap is between 3 and 4, and it is organisational rather than technical. Any harness with shell access can run gh pr create. What rung 4 actually requires is that the vendor has thought about identity — whose token pushes the branch, what CI runs, what happens on a failed check, whether the agent can respond to review comments. GitHub Copilot's coding agent is the clearest example: it runs in Actions, pushes under a bot identity, and is structurally barred from approving its own work. Devin and Cursor's background agents solve the same problem with a hosted VM and their own identity. Claude Code and opencode solve it with published GitHub Apps that answer a mention on an issue, which is why an MIT-licensed terminal harness and a proprietary one land on the same rung — the licence has nothing to do with it.

Note also how little of the ladder is actually in play. Nothing on this page scores 1 or 2: shipping a coding agent in 2026 means writing to the working tree, and the tools that only suggest diffs are autocomplete products in a different category. Nothing scores 5 either. So the killer field is really a two-value sort, and rung 3 hides a wide band inside it — Aider driving a test command and fixing what it breaks is doing more unattended work than an IDE agent that pauses for approval on every command, and this column will not tell them apart. Read it with the permissions and headless columns rather than alone.

The gap between 4 and 5 is nobody's to close unilaterally. Merging is a policy decision that lives in branch protection rules, not in a harness. Treat any tool claiming rung 5 as claiming to have been given credentials, and audit it accordingly.

Pricing traps

The flat subscription is the most honest deal on this page, but read the limits. Every subscription-based harness enforces a rolling usage window, and the ceiling is expressed in vendor-specific units that do not map cleanly onto tokens. Heavy agentic use — long sessions, big repos, subagents — burns through those windows far faster than chat does, and the answer when you hit one is either wait or move to metered billing.

Credits are worse for planning. Windsurf, Devin and Amp all price in units the vendor defines. The unit economics can be perfectly fair and still be unforecastable, because the vendor can re-price the conversion between a credit and a token when the underlying model price changes, and you find out afterwards. Copilot's premium-request multipliers have the same shape. If someone in finance needs a number for next quarter, a per-seat subscription or your own provider key are the only two things you can defend.

BYOK looks cheapest and often is, but it moves the variance onto your API bill rather than removing it. A harness that re-reads a large file on every turn costs multiples of one that caches aggressively, and the difference shows up in your provider invoice rather than the tool's price page. If you are running BYOK at team scale, meter it per repo for a fortnight before you extrapolate, and check whether the harness supports prompt caching against your provider at all.

How to actually choose

Three decision paths cover most teams. If the work is well-specified tickets on a repo with good tests, buy autonomy: Claude Code or Codex CLI in headless mode behind a PR, or Copilot's coding agent if you already live in GitHub and want the identity problem solved for you. Optimise for rung 4, background execution and a sandbox you trust.

If the work is exploratory or architectural, buy ergonomics instead. Cursor and Windsurf exist because reviewing a large agentic diff inside an editor with real navigation beats reading it in a terminal, and both maintain a persistent embeddings index that grep-based harnesses cannot fake on a million-line monorepo. They are not alone in that: Kilo Code indexes too, and GitHub Copilot and Devin read an index alongside live search, which is a richer substrate again. What Cursor has that the others do not is that index sitting inside the editor you are reviewing in. Zed is the interesting third option: a fast editor that hosts other agents over the Agent Client Protocol, so you keep your harness choice open while getting an editor's review surface.

If you are constrained by procurement, data residency or a fear of lock-in, take a provider-agnostic open-source harness. opencode if you want a modern TUI and the broadest provider coverage, Goose if MCP is already central to your stack, Cline if your team is VS Code-native, Aider if you want the smallest auditable thing that works. You will be a quarter behind on features. You will also be able to point every one of them at a different model provider on a Tuesday afternoon.

Whatever you pick, write AGENTS.md, pin a test command, and keep the sandbox on. Those three decisions outlast the tool.

What is coding agents compared: claude code vs codex cli vs cursor vs the rest?

Coding agents compared: Claude Code vs Codex CLI vs Cursor vs the rest

On toolweight, Coding agents compared: Claude Code vs Codex CLI vs Cursor vs the rest means the 17 tools benchmarked on this page — Claude Code, Codex CLI, Gemini CLI, opencode, Kilo Code, Cline, Cursor, Windsurf, Aider, Devin, Amp, Zed, Goose, GitHub Copilot (Agent HQ), Grok Build, Qwen Code, DIY harness — judged on the same 26 fields, from the same sources, on the same date. The question it exists to answer: Which AI coding agent should I use in 2026?

How does toolweight compare these?

Every cell here comes from vendor documentation, the tool's own repository, or hands-on use, and each carries its provenance. Where we are reading a pricing page we mark it vendor-claimed with the URL. Where we are inferring from adjacent behaviour, or making a judgement call like a score, we mark it inferred. Where the fact is widely reported by users but not documented, it is community. Where we do not know — and in a category moving this fast that is often — the cell is null and renders as a dash rather than a guess. Closed-source harnesses have null GitHub stars because no comparable repository exists, not because nobody stars them.

The verify cadence for this page is fourteen days, and the dated cells below currently sit well outside it: the last full pass was 15 January 2026. We publish the dates rather than refresh them, because a date is a claim about work someone did, and moving one without redoing the check is the cheapest lie a comparison site can tell. Read every date as the last time a human actually looked.

The killer field, where does it stop?, is scored on a fixed five-rung ladder and always at the highest rung the tool reaches out of the box: rung 4 requires a first-class path to a pull request — a built-in command, a vendor-published GitHub App with its own identity, or a hosted runner — not merely the ability to invoke gh because you asked it to run a shell command. Any harness with shell access can be scripted up a rung; that is a property of your scripts, not the product. A vendor-published Action you wire into your own workflows, with no App identity behind it, is graded partial rather than a full rung, and that rule is applied to first-party and third-party harnesses alike.

The Terminal-Bench column ships empty, and we would rather that than the alternative. The published figures we could trace either name a model and a benchmark version without naming a harness, or were produced by the benchmark's own reference agent — which makes them evidence about the model, not about anybody's harness, and importing them into a harness column is exactly the laundering this page exists to argue against. Terminal-Bench 1.0 results are also superseded by 2.0 and not comparable to it, and the gaps between the top entries were smaller than one task on an eighty-task benchmark, which is not a ranking. An empty column is the honest state of this evidence until we run the benchmark ourselves against pinned harness versions.

Full methodology and sourcing policy →

Frequently asked questions

Does the harness matter more than the model?

Below the frontier, no — a weak model fails regardless of harness. At the frontier, yes. Once the model can do the task, completion rate is decided by context strategy, tool design, how failures are fed back, and whether the loop stops cleanly. The same model in Claude Code and in a bare chat window are not the same product, which is why agent = model + harness is the only useful framing.

Which coding agents can open a pull request on their own?

Claude Code (its hosted cloud sessions, and the Claude GitHub App installed by /install-github-app), opencode (the opencode-agent GitHub App, which answers a mention on an issue and comes back with a branch and a PR), Codex CLI's cloud tasks, Cursor's background agents, GitHub Copilot's coding agent, and Devin. That is rung 4 on our ladder. Gemini CLI has an official Action but no App, and its documented workflows lean on review and triage rather than authoring, so we grade it partial. Everything else — Cline, Aider, Goose, Zed, Qwen Code, Kilo Code — stops at editing and running tests on your machine.

Can any of them merge and deploy without a human?

Not out of the box, and treat any claim otherwise as a description of your own CI. Copilot's agent is explicitly forbidden from approving its own pull request. Devin stops at review. Claude Code and Codex do whatever their permissions allow, so with a privileged token and approvals disabled they will merge — but that is you removing the gate, not the product shipping rung 5.

What is the cheapest way to run a coding agent all day?

Either a flat subscription with soft limits (Claude Code on a Max plan, Codex on a ChatGPT Pro plan) or a free harness with your own key: opencode, Aider, Goose, Cline, Zed. Gemini CLI and Qwen Code both have genuinely usable free tiers tied to a personal account. Avoid credit-based pricing if you need to forecast, because the conversion between a credit and a token is the vendor's to change.

Are Terminal-Bench scores comparable between harnesses?

Only when the harness version, the model and the benchmark version are all stated — and that column is currently empty on this page, because nothing we could trace met the bar. Terminal-Bench 2.0 results are not comparable to 1.0, a harness can move several points on the same model between releases, and a figure produced by the benchmark's own reference agent is a model result that says nothing about any harness. Differences of half a point on an eighty-task benchmark are also a fraction of one task, not a ranking. We would rather show a blank column than sort a table on that.

Is it safe to let an agent run shell commands unattended?

Only inside a sandbox with a real filesystem and network boundary. Codex CLI (Seatbelt on macOS, Landlock on Linux) and Claude Code ship OS-level sandboxing; Gemini CLI and Goose ship one too, but off by default, so check the cell note before you assume it is on; Devin and Copilot's coding agent run in vendor-hosted VMs instead, which we rank level with a local kernel sandbox rather than below it. Most IDE extensions offer only an auto-approve allowlist, which is a convenience feature rather than a security boundary. For anything unattended, run it in a container or a sandbox provider.

Should my team standardise on one harness?

Standardise the repo-level configuration, not the harness. AGENTS.md, MCP server definitions and a documented test command are portable across nearly everything on this page; skills, hooks and rules files mostly are not. That way an engineer who prefers Zed and one who lives in Claude Code get the same project context, and swapping harnesses next quarter costs a config change rather than a migration.