Get in touch

9io.ai / Blog

Claude Code vs Codex vs Cursor in 2026: benchmarks, limits and cost

Claude Code vs Codex vs Cursor, plus Pi, as of 7 October 2026: the models each runs, independent benchmarks, plan prices, usage limits and which to pick.

Key takeaways

  • Claude Opus 5.5 leads Vals’ Terminal-Bench 4.0 at 65.15%, ahead of GPT-6 Astra at 59.60%.
  • GPT-6.1 Sol, OpenAI’s recommended Codex model, scores 55.05% on the same test for $1.72 a task.
  • OpenAI has proposed ending its model supply to Cursor on 12 November, and Cursor has no GPT-6 model.
  • Claude Code, Codex, Cursor and Pi all read AGENTS.md, so one file can serve a mixed team.
  • Claude Pro, ChatGPT Plus and Cursor Pro cost $20 a month, and top individual plans $200 to $500.

Claude Code vs Codex comes down to model family and budget, as of 7 October 2026. Claude Code runs Anthropic’s Claude models, and Claude Opus 5.5 leads Vals’ independent Terminal-Bench 4.0 at 65.15%. Codex is OpenAI’s, and GPT-6.1 Sol, the model OpenAI recommends for coding in Codex, scores 55.05% on the same test for $1.72 a task, 13% of Opus 5.5’s $13.20.1 Cursor is the choice for teams that want an editor first and models from several labs, though OpenAI has proposed cutting off its models there on 12 November.2

All four tools here are agents that read a repository, edit files and run commands. Claude Code and Codex now run in terminals, IDEs and their makers’ clouds. Cursor is a code editor that SpaceX has owned since August, and Pi, which reached 1.0 on 1 October, is a minimal open-source terminal agent.34

This post compares the models each tool can run, benchmark results and what they measure, plans and usage limits, features and developer surveys, and ends with a decision table. It uses public sources dated on or before 7 October 2026. We haven’t run our own benchmark for it, and every figure links to its source.

Claude Code vs Codex vs Cursor vs Pi at a glance

The table summarises the four tools as of 7 October 2026, and the later sections give the sources.

Claude Code Codex Cursor Pi
Made by Anthropic OpenAI Cursor, part of SpaceX since 14 August 2026 Earendil
Kind of tool Terminal agent, also in VS Code, JetBrains, a desktop app and the web Agent in a CLI, an IDE extension, the ChatGPT apps and OpenAI’s cloud AI code editor with a CLI and cloud agents Minimal terminal agent that you extend yourself
Source code Proprietary CLI is open source (Apache 2.0) Proprietary Open source (MIT)
Models Claude models, Opus 5.5 by default on subscriptions OpenAI models, GPT-6.1 Sol by default in the CLI, plus custom providers Claude, Gemini, GPT-5.6, Muse Spark, Grok and Composer Most hosted providers, compatible endpoints and local models
Price for one developer $20 to $200 a month Limited on the free plan, then $8 to $500 a month Free tier, then $20 to $200 a month Free, plus whatever the model costs

Anthropic publishes the Claude Code repository under an all-rights-reserved licence,5 OpenAI publishes the Codex CLI under Apache 2.0,6 and Pi is MIT-licensed.7

Which models each agent can run

Claude Code runs Claude models through Anthropic’s API, Amazon Bedrock, Google Cloud, Microsoft Foundry or an LLM gateway.8 Claude Opus 5.5 arrived on 22 September as the default Opus model, and the same release made Opus the default on Pro and Team Standard plans, as it already was on Max, Team Premium and Enterprise. Sonnet 5.5 followed on 28 September and Haiku 5.5 on 7 October.9 Fable 5.1 has its own weekly limit on Max plans, listed as 50% of the plan’s weekly limits, and on Pro it runs only on paid usage credits.10 Real-SWE ran Z.ai’s GLM 5.3 inside Claude Code, which shows that other labs’ models can be wired into it.11

Codex runs OpenAI’s models. GPT-6 Astra is the most capable, and GPT-6.1 Sol, launched at DevDay on 29 September, is described by OpenAI as close to Astra at a fifth of Astra’s standard token prices.12 OpenAI’s model guide recommends GPT-6.1 Sol for complex coding where it’s available.13 GPT-6.1 Sol became the default model in Codex CLI 0.161.0 on 7 October, and GPT-5.5 leaves Codex on 14 October.14 The CLI can also use other providers, including local models through Ollama or LM Studio.15

Cursor offers Anthropic’s four current Claude models, Google’s Gemini 3.1 Pro and 3.8 Flash, Meta’s Muse Spark 1.3 and OpenAI’s GPT-5.6 Sol, Terra and Luna at each provider’s API price. Grok 4.7, 4.6 and 4.5 and Cursor’s Composer 2.5 come from a separate pool with more included usage, and Cursor’s docs treat them as first-party models. Some others, such as Z.ai’s GLM 5.3 and Moonshot’s Kimi K3, are hidden by default.16 SpaceX completed its purchase of Cursor on 14 August 2026.3 Two weeks later, OpenAI said it would wind down its contract supplying models to Cursor, proposed 12 November 2026 as the shutoff date and said it would provide no future models in the meantime.2 As of 7 October, Cursor lists no GPT-6 model, and its CursorBench leaderboard has none.17

Pi works with almost any model. Its 1.0.4 documentation covers API keys for Anthropic, OpenAI, Google, DeepSeek, Mistral, xAI, OpenRouter and many other providers, subscription sign-in through /login, local GGUF models through llama.cpp, and any OpenAI-, Anthropic- or Google-compatible endpoint.7

Independent benchmarks score the model in a shared harness

Vals runs every model on Terminal-Bench 4.0 through mini-swe-agent, a minimal harness that gives the model one bash tool and no step or cost limit. It averages three runs over 66 tasks, allows eight hours per task and gives no partial credit.1 Because the harness is the same for every model, the scores rank models. They don’t measure Claude Code or Codex, which add planning, subagents, context management and permission checks on top of the model, and those layers can move a result either way.

Scores and costs are as of 7 October 2026, the date of Vals’ latest update. Cost per solved task is our arithmetic, Vals’ cost per task divided by the pass rate, at the API prices Vals reports.

Model Available in Terminal-Bench 4.0 Cost per task Cost per solved task
Claude Opus 5.5 Claude Code (default), Cursor 65.15% (58.08%)¹ $13.20 $20.26
Claude Sonnet 5.5 Claude Code, Cursor 64.14% (62.63%)¹ $16.51 $25.74
GPT-6 Astra Codex 59.60% $9.58 $16.07
Claude Fable 5.1 Claude Code, Cursor 58.08% (50.00%)¹ $17.18 $29.58
GPT-6.1 Sol Codex (CLI default) 55.05% $1.72 $3.12
GPT-5.6 Sol Codex, Cursor 37.88% $7.98 $21.07
Grok 4.7 Cursor 28.79% $18.09 $62.83
  1. The figure in brackets counts Anthropic’s provider-side fallback as failure. Vals reports that 22 of Opus 5.5’s 198 task attempts were served by Opus 5 or Claude Opus 4.8, and that scoring those as failures drops Opus 5.5 to 58.08%, behind Sonnet 5.5 and GPT-6 Astra.1

Pi can call the Claude and GPT models in this table through each provider’s API.

On the raw scores, Opus 5.5 is 10.1 points ahead of GPT-6.1 Sol and 5.55 points ahead of GPT-6 Astra. With fallback attempts counted as failures, it sits 1.52 points behind Astra. The cost gap is larger than the score gap. GPT-6.1 Sol solves a task for about $3.12 in tokens against $20.26 for Opus 5.5, so a team paying API prices gets about 6.5 times as many solved tasks per dollar from Sol. Cursor’s own models score lower on this test. Grok 4.7, which Cursor calls its most capable model, scores 28.79%,3 and GPT-5.6 Sol, the strongest OpenAI model Cursor offers on this test, scores 37.88%.

Product-level results from Real-SWE and CursorBench

Two newer benchmarks test the products themselves, and both come with caveats.

Real-SWE, from Specific, runs each vendor’s own agent on tasks from private codebases licensed from real companies. Each task gets eight runs, scored as pass@1.11 It is the closest public test of Claude Code against Codex as products. Cost per resolved task is our arithmetic, as above.

Agent and model Resolved Cost per run Cost per resolved task
Codex CLI with GPT-6 Astra 46.25% $4.67 $10.10
Claude Code with Claude Fable 5.1 45.00% $6.96 $15.47
Gemini CLI with Gemini 3.8 Flash 38.75% $2.50 $6.45
Claude Code with GLM 5.3 37.50% $5.12 $13.65
Codex CLI with GPT-5.6 Sol 26.25% $2.65 $10.10

Codex with Astra and Claude Code with Fable are 1.25 points apart, and the page shows its confidence intervals only in a chart, so treat the two as level. The page analyses a sample of tasks without giving the benchmark’s total size, its verifiers are private, and Specific built and runs the benchmark itself. It also leaves out the default models. There are no results for Opus 5.5, Sonnet 5.5 or GPT-6.1 Sol, and Cursor and Pi aren’t tested.

CursorBench 4.0 is Cursor’s own benchmark, built from ambiguous, multi-file tasks taken from real Cursor sessions and last updated on 7 October.17 It shows how models do inside Cursor, and the effort setting matters as much as the model.

Model and effort CursorBench 4.0 Cost per task
Claude Opus 5.5, max 57.8% $13.43
Claude Opus 5.5, high 56.0% $3.97
Claude Sonnet 5.5, extra high 53.1% $2.81
Claude Fable 5.1, max 51.8% $17.28
Grok 4.7, extra high 46.3% $6.01
GPT-5.6 Sol, max 41.7% $8.23
Composer 2.5 27.7% $0.68

Opus 5.5 at high effort scores 1.8 points below max effort for under 30% of the cost per task. Claude models fill the top 13 places, and Grok 4.7 is the best model from another lab. Treat these numbers as vendor-reported. Cursor builds the benchmark and sells its own models on the same leaderboard, and the page doesn’t say how many tasks it has or how they are graded.

Plans, prices and usage limits as of 7 October 2026

Claude Pro, ChatGPT Plus and Cursor Pro all cost $20 a month, but the three vendors meter use in different ways, so seat prices don’t buy equal amounts of work. The table lists monthly US prices as of 7 October 2026, before tax.

Claude Code Codex Cursor
Free Not included on Claude Free Free and Go ($8) get GPT-6 Luna in the desktop app, subject to rollout Hobby, with limited Agent requests
Entry plan Pro, $20 ($17 billed annually) Plus, $20 Pro, $20
Heavier individual plans Max 5x $100, Max 20x $200 Pro at $100, $200 or $500 Pro+ $60, Ultra $200
Team seat Standard $25 ($20 annually), Premium $125 ($100 annually) Business $25 ($20 annually) Teams Standard $40, Premium $120
Enterprise $20 a seat plus usage at API rates Contact sales Custom
How limits work Rolling five-hour window plus weekly limits, shared with Claude chat Five-hour estimates per model, weekly limits may apply Monthly included usage in two pools, then on-demand

Sources are Anthropic’s pricing and Claude Code pages,10 OpenAI’s Codex pricing page18 and Cursor’s pricing documentation.16 Pi itself is free, and you pay the model provider.

Claude. Every plan has limits that reset on a rolling five-hour window, and paid plans add weekly limits. Claude chat and Claude Code draw from the same pool. Max gives 5 or 20 times Pro’s usage per five-hour session, and a Team Premium seat gives 5 times a Standard seat. Anthropic publishes no fixed message counts, and paid plans can buy usage credits at API rates once the limits run out.10

Codex. OpenAI publishes ranges of local messages per five hours on Plus and Standard Business, 5 to 45 with GPT-6 Astra, 15 to 160 with GPT-6.1 Sol and 350 to 3,000 with GPT-6 Luna. Cloud tasks can use more, weekly limits may also apply, and Pro plans had no five-hour limit as of 7 October.18 OpenAI says the $500 Pro plan gives 25 times the Plus allowance.12 Past the limits, Plus and Pro users can buy credits or run Codex with an API key at API rates.

Cursor. Pro, Pro+ and Ultra include two pools that reset monthly. The Cursor Models pool covers Grok and Composer with more included usage, and the Other Models pool is charged at each provider’s API price. Cursor’s own guidance is that daily agent users typically spend $60 to $100 a month in total usage and power users often $200 or more. Teams and Enterprise add a Cursor Token Rate of $0.25 per million tokens on third-party models.16

For a team of ten developers on monthly billing, entry team seats cost $250 a month on Claude Team Standard (10 × $25) or ChatGPT Business (10 × $25). Cursor Teams Standard costs $400 (10 × $40). The heavier seats cost $1,250 on Claude Team Premium (10 × $125) and $1,200 on Cursor Teams Premium (10 × $120). Annual billing brings the Claude and ChatGPT seats down to the prices in brackets in the table. Paying per token is the other route. Anthropic’s documentation puts the average Claude Code cost across enterprise deployments at about $13 per developer per active day and $150 to $250 per developer per month, with 90% of users under $30 per active day.19 On Claude Enterprise, that suggests $1,700 to $2,700 a month for the same ten developers, $200 in seats plus $1,500 to $2,500 in usage. Our post on what coding agents cost per engineer works through budgets for teams of 1, 5 and 25 and the caps that keep them in bounds.

AGENTS.md, subagents, MCP and cloud agents compared

On features, the four tools have largely converged. This table is as of 7 October 2026.

Feature Claude Code Codex Cursor Pi
Reads AGENTS.md Yes, since 18 September, when there is no CLAUDE.md Yes, per directory, with AGENTS.override.md Yes, in the root and nested folders Yes, as well as CLAUDE.md
Subagents Yes, plus agent teams (experimental) and scripted workflows Yes, on by default Yes, in the editor, CLI and cloud agents Not built in; added through extensions
MCP Yes, 2026-07-28 revision by default Yes, stdio and streamable HTTP with OAuth Yes, stdio, SSE and streamable HTTP with OAuth Yes, since 1.0
Cloud or background work Cloud sessions, routines, Projects (beta), Slack Codex Cloud, automatic code review Cloud agents, Projects (beta), Bugbot None built in
IDE integration VS Code and JetBrains extensions, desktop app Extension for VS Code, Cursor and Windsurf; JetBrains and Xcode integrations It is the editor None; RPC and an SDK for building your own
Default permission handling Auto mode, where a classifier model reviews actions Operating-system sandbox plus an approval policy Sandbox setting in the CLI No sandbox and no approval before each tool call

Sources are the vendors’ documentation for each feature.820212223712

AGENTS.md is now common ground. Claude Code 2.1.277, released on 18 September, reads AGENTS.md as the project instructions when a repository has no CLAUDE.md, and version 2.1.281 extended that to Bedrock, Vertex AI, Foundry and gateways on 23 September.9 If a repository has both files, Claude Code reads only CLAUDE.md, unless that file imports AGENTS.md or you change its Project instructions setting.24 Codex, Cursor and Pi read AGENTS.md natively,21237 and the format is now stewarded by the Agentic AI Foundation under the Linux Foundation.25 One AGENTS.md at the repository root can therefore serve a team that uses all four tools.

MCP support differs by protocol revision. All four can use MCP servers, although Pi only added it in 1.0, having previously ruled it out.26 Claude Code negotiates the stateless 2026-07-28 revision by default with HTTP servers and, since 6 October, with local stdio servers.9 A Cursor staff member said on 21 September that Cursor’s MCP client speaks the 2025-11-25 revision and earlier, and on 25 September that there was no timeline for the new one.27 If you run an MCP server for a mixed team, keep the older revision working. Our stateless MCP migration guide covers which clients speak which revision, and making one MCP server work in Claude, ChatGPT, Cursor and VS Code covers the OAuth differences between them.

Default permissions differ sharply. Since version 2.1.283, Claude Code starts interactive terminal and VS Code sessions in auto mode, where a second model reviews actions instead of asking you. Manual mode and a Bash sandbox are also available.28 Codex runs local commands inside an operating-system sandbox by default and asks before going beyond it.29 Cursor’s CLI has a sandbox mode that you can switch on or off.23 Pi’s own security guide says it can read, change and run files with your account’s permissions without asking before each tool call, and it recommends a container or virtual machine.7

What developer surveys say about Claude Code, Codex and Cursor

Three surveys this year asked developers which agents they use. They sample different people, so compare the trends within each survey rather than the numbers across them.

Survey Sample Claude Code Codex Cursor GitHub Copilot
JetBrains, May to July 202630 15,000+ professional developers, reweighted 39% use it at work (18% in January) 16% (3% in January) 12% (18% in January) 21% (29% a year earlier)
Stack Overflow, published 6 October 202631 30,000+ respondents 66% Not reported Not reported 59%
The Pragmatic Engineer, January to February 202632 906 respondents 71% of regular agent users About 60% of Cursor’s usage 39% of regular agent users 46% of regular agent users

Stack Overflow’s page doesn’t say what its 66% and 59% are shares of, so read them as rankings.

Claude Code leads in all three, and in JetBrains’ data it is the most-used AI coding tool for 31% of professional developers. Codex grew fastest in JetBrains’ data, from 3% to 16% adoption at work between January and the summer. OpenAI reported more than 5 million weekly active Codex users on 2 June, up more than six times since its desktop app launched in February.33 Cursor’s adoption fell from 18% to 12% in JetBrains’ data while awareness of it rose from 69% to 75%. JetBrains also found that 90% of professional developers used AI coding agents at work at least weekly, and 68% daily.30

Pi is smaller and newer. Earendil says hundreds of thousands of people use it every week,4 and the 1.0 announcement was the top story on Hacker News’s front page for 1 October.34

Claude Code vs Pi: a managed product or a harness you own

Pi’s README calls it a minimal, extensible agent harness. It ships without sub-agents or plan mode and expects you to add what you need through extensions, skills and packages. It runs interactively, in print or JSON mode, over RPC or inside your own program through a TypeScript SDK.7

The trade-off is control against guardrails. With Pi you can choose any model, route work between providers and change the agent’s source code. Claude Code is extensible through plugins and hooks, but its code is Anthropic’s, and in return you get Anthropic’s permission modes, cloud sessions, IDE extensions and admin controls.598 Pi leaves isolation to you, and its documentation is direct about that.7

Claude subscribers can use Pi too. Anthropic’s help centre says, in an update dated 7 October 2026, that third-party apps can still use a subscription’s limits, and the same update added monthly API credits to Max and Team plans.35 The credits are $100 a month on Max 5x, $200 on Max 20x and $20 or $100 per Team seat, pooled up to $500. They cover the Claude API and the Agent SDK but not interactive Claude Code. Because they apply to any API key in the linked Console organisation, they also pay for Pi calling Claude with such a key. Pi ships a setting that warns when Anthropic subscription sign-in may draw on paid extra usage,7 so check your usage page after the first long session.

Pi is the better fit for engineers who want to own the agent, mix vendors to control cost or embed an agent in their own product. For a team that wants permissions, updates and the surrounding tools handled by the vendor, Claude Code is the safer default.

The best AI coding agent for each kind of team

If your team… Start with Because
Works in the terminal and wants the top independent benchmark scores Claude Code Claude Opus 5.5 and Sonnet 5.5 lead Terminal-Bench 4.0, and Claude models lead CursorBench 4.0
Already pays for ChatGPT, or pays API prices and watches cost per task Codex It is included in ChatGPT plans, and GPT-6.1 Sol scores 55.05% for $1.72 a task
Needs GPT-6 Astra or GPT-6.1 Sol Codex, or Pi with an OpenAI key Cursor has no GPT-6 model, and Claude Code is built around Claude models
Works in an editor and wants Claude, Gemini and Grok in one place Cursor It offers several labs’ models and cloud agents, though OpenAI’s models are due to leave on 12 November
Wants an open-source agent it can modify or embed Codex CLI or Pi They are Apache 2.0 and MIT licensed, and Pi runs nearly any provider’s model
Needs guardrails on by default Claude Code or Codex Claude Code reviews actions with a classifier and Codex sandboxes commands, while Pi has neither built in
Uses a mix of tools Any of them All four read AGENTS.md

Before you commit, work through this list:

  1. Start from where your developers work. Terminal-first teams should trial Claude Code and Codex, and editor-first teams Cursor alongside the VS Code extensions of the other two.
  2. List the models you need. GPT-6 means Codex or Pi, and several vendors in one editor means Cursor.
  3. Price the tier your heaviest users need. Anthropic pitches Pro at short coding sprints, OpenAI says Plus covers a few focused sessions a week, and Cursor recommends Pro+ for daily agent users.101816 Compare Max, Pro or Ultra against what the same work would cost on API keys.
  4. Put shared instructions in AGENTS.md first. The trial is fairer when every tool reads the same project rules.
  5. Run the same real tasks through two tools for two weeks. Count merged changes and review time, and check each vendor’s usage page every few days.
  6. Put two dates in the calendar. GPT-5.5 leaves Codex on 14 October, and OpenAI’s proposed Cursor shutoff is 12 November.

How this comparison was put together

  • Product facts come from each vendor’s documentation, pricing pages and changelogs as of 7 October 2026. Changelog entries dated after 7 October were left out.
  • Benchmarks come from Vals (Terminal-Bench 4.0, updated 7 October), Specific (Real-SWE, September 2026) and Cursor (CursorBench 4.0, updated 7 October). Only Vals is independent of the products it tests.
  • Cost per solved task is our arithmetic, cost per task divided by the pass rate, at the API list prices each benchmark used. Subscribers pay flat fees, so these figures compare efficiency and don’t predict a bill.
  • Surveys are JetBrains (18 August 2026), Stack Overflow (6 October 2026) and The Pragmatic Engineer (3 March 2026).
  • Not covered: our own benchmark runs (we didn’t do any for this post), latency, GitHub Copilot and Gemini CLI beyond the survey figures, enterprise discounts and data-retention terms.

Prices, limits and model line-ups in this market change every few weeks, and corrections are welcome at admin@9io.ai.


  1. Vals AI, “Terminal-Bench 4.0”, updated 7 October 2026, https://www.vals.ai/benchmarks/terminal-bench-4 ↩↩↩

  2. OpenAI, “Our decision on Cursor following its acquisition by SpaceX”, 28 August 2026, https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/ ↩↩

  3. Cursor, “Cursor is now a part of SpaceX”, 14 August 2026, https://cursor.com/blog/joining-spacex, and Cursor blog, “Introducing Grok 4.7”, 21 September 2026, https://cursor.com/blog ↩↩↩

  4. Earendil, “Pi 1.0”, 1 October 2026, https://earendil.com/posts/pi-1-0/ ↩↩

  5. Anthropic, claude-code repository licence, https://github.com/anthropics/claude-code/blob/main/LICENSE.md ↩↩

  6. OpenAI, openai/codex repository, https://github.com/openai/codex ↩

  7. Earendil, Pi 1.0.4 documentation (README, providers, configuration, security and settings), 5 October 2026, https://github.com/earendil-works/pi/tree/v1.0.4/packages/coding-agent, including “Run Pi safely”, https://github.com/earendil-works/pi/blob/v1.0.4/packages/coding-agent/docs/security.md ↩↩↩↩↩↩↩↩

  8. Anthropic, “Platforms and integrations”, Claude Code docs, https://code.claude.com/docs/en/platforms ↩↩↩

  9. Anthropic, “Claude Code changelog”, versions 2.1.257 to 2.1.293 (1 September to 7 October 2026), https://code.claude.com/docs/en/changelog ↩↩↩↩

  10. Anthropic, “Plans & pricing”, https://claude.com/pricing, and “Claude Code”, https://claude.com/product/claude-code ↩↩↩↩

  11. Specific, “Real-SWE”, September 2026, https://withspecific.com/benchmarks/real-swe ↩↩

  12. OpenAI, “DevDay 2026 Recap”, 29 September 2026, https://openai.com/index/devday-2026-recap/ ↩↩↩

  13. OpenAI, “Models”, ChatGPT Work and Codex docs, https://developers.openai.com/codex/models ↩

  14. OpenAI, “ChatGPT & Codex changelog”, entries from 14 September to 7 October 2026, https://developers.openai.com/codex/changelog ↩

  15. OpenAI, “Advanced configuration”, Codex docs, https://developers.openai.com/codex/config-file/config-advanced ↩

  16. Cursor, “Pricing”, https://cursor.com/pricing, and “Models & Pricing”, https://cursor.com/docs/models-and-pricing ↩↩↩↩

  17. Cursor, “CursorBench 4.0”, updated 7 October 2026, https://cursor.com/cursorbench ↩↩

  18. OpenAI, “Pricing”, ChatGPT Work and Codex docs, https://developers.openai.com/codex/pricing ↩↩↩

  19. Anthropic, “Manage costs effectively”, Claude Code docs, https://code.claude.com/docs/en/costs ↩

  20. Anthropic, “Run agents in parallel”, Claude Code docs, https://code.claude.com/docs/en/agents ↩

  21. OpenAI, “AGENTS.md” and “Subagents”, Codex docs, https://developers.openai.com/codex/agent-configuration/agents-md and https://developers.openai.com/codex/agent-configuration/subagents ↩↩

  22. OpenAI, “Model Context Protocol” and “Codex IDE extension”, Codex docs, https://developers.openai.com/codex/extend/mcp and https://developers.openai.com/codex/ide ↩

  23. Cursor docs, “Rules”, “Subagents”, “Model Context Protocol (MCP)”, “Cloud Agents” and “CLI”, https://cursor.com/docs/rules, https://cursor.com/docs/subagents, https://cursor.com/docs/mcp, https://cursor.com/docs/cloud-agent and https://cursor.com/docs/cli/overview, and Cursor changelog, https://cursor.com/changelog ↩↩↩

  24. Anthropic, “How Claude remembers your project”, Claude Code docs, https://code.claude.com/docs/en/memory ↩

  25. AGENTS.md, https://agents.md/ ↩

  26. Earendil, “You said no MCP”, 29 September 2026, https://earendil.com/posts/you-said-no-mcp/ ↩

  27. Cursor community forum, “MCP client cannot connect to modern-only 2026-07-28 Streamable HTTP servers”, 21 to 25 September 2026, https://forum.cursor.com/t/mcp-client-cannot-connect-to-modern-only-2026-07-28-streamable-http-servers-legacy-initialize-rejected/172536 ↩

  28. Anthropic, “Choose a permission mode”, Claude Code docs, https://code.claude.com/docs/en/permission-modes ↩

  29. OpenAI, “Sandbox”, Codex docs, https://developers.openai.com/codex/sandboxing ↩

  30. JetBrains Research, “AI Coding Agents: Adoption Trends”, 18 August 2026, https://blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026/ ↩↩

  31. Stack Overflow, “The results of the 2026 Developer Survey are here!”, 6 October 2026, https://stackoverflow.blog/2026/10/06/the-results-of-the-2026-developer-survey-are-here/, and “AI”, 2026 Developer Survey, https://survey.stackoverflow.co/2026/ai ↩

  32. The Pragmatic Engineer, “AI Tooling for Software Engineers in 2026”, 3 March 2026, https://newsletter.pragmaticengineer.com/p/ai-tooling-2026 ↩

  33. OpenAI, “Codex is becoming a productivity tool for everyone”, 2 June 2026, https://openai.com/index/codex-for-knowledge-work/ ↩

  34. Hacker News, front page for 1 October 2026, https://news.ycombinator.com/front?day=2026-10-01 ↩

  35. Anthropic Help Center, “Use the Claude Agent SDK with your Claude plan”, update of 7 October 2026, https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan, and “Monthly API credits for Max and Team plans”, https://support.claude.com/en/articles/17154008-monthly-api-credits-for-max-and-team-plans ↩

Frequently asked questions

Is Claude Code better than Codex?

On Vals’ independent Terminal-Bench 4.0, Claude Opus 5.5 scores 65.15% and GPT-6.1 Sol, the model OpenAI recommends for Codex, scores 55.05%. Sol costs $1.72 a task against $13.20 for Opus 5.5, and on Real-SWE, which runs each vendor’s own agent, Codex with GPT-6 Astra (46.25%) and Claude Code with Fable 5.1 (45.00%) are close.

What is the difference between Codex and Cursor?

Codex is OpenAI’s coding agent, used from a CLI, an IDE extension, the ChatGPT apps or OpenAI’s cloud, and it runs OpenAI models. Cursor is an AI code editor with a CLI and cloud agents that offers models from several labs, including its own Grok and Composer models, but no GPT-6 model.

Is Cursor better than Claude Code?

They suit different teams. Cursor fits developers who want to work in an editor and switch between model vendors. Claude Code is a terminal-first agent for Claude models, which hold the top places on both Vals’ Terminal-Bench 4.0 and Cursor’s own CursorBench 4.0.

What is the best AI coding agent in 2026?

As of 7 October 2026, Claude Code is the most used in the large surveys, with 39% of professional developers using it at work in JetBrains’ May to July data, and Claude models lead Terminal-Bench 4.0. Codex with GPT-6.1 Sol solves tasks far more cheaply at API prices, so the best choice depends on your editor, budget and model needs.

What is Pi and how is it different from Claude Code?

Pi is an MIT-licensed terminal agent from Earendil that reached version 1.0 on 1 October 2026. It runs models from most providers and leaves out sub-agents, plan mode and a built-in sandbox, while Claude Code is Anthropic’s proprietary agent with subagents, cloud sessions, IDE extensions and permission controls.

How much do Claude Code, Codex and Cursor cost?

As of 7 October 2026, Claude Pro, ChatGPT Plus and Cursor Pro each cost $20 a month. Heavier plans cost $100 or $200 for Claude Max, $100, $200 or $500 for ChatGPT Pro, and $60 or $200 for Cursor Pro+ and Ultra.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.