What a coding agent actually costs
Work out the price of one finished task, find which meter your coding agent runs on, and decide whether the line on your card is earning its place.
on this page · 0 / 0 checked
A coding agent is the easiest software purchase you will ever make and the hardest one to assess afterwards. The seat is cheap, the first session is convincing, and a month later there is a second line on the invoice that nobody authorised in advance. The seat price was never the price. It was the entry fee for a meter, and the meter runs on work the agent decides to do.
This guide is for a solo developer, a freelancer, or a small team paying for a coding agent out of their own budget and unable to say whether it is earning its keep. It is not for an engineering organisation with a FinOps function and a procurement process, which already has better tools for this than a web page. It is also not a comparison review; the aim is to leave you with one number you can defend, not a winner. Every price below was read from the vendor’s own page on 5 September 2026. Vendors move them, and at least one of the units quoted here was replaced during this year [4].
Coding agents bill in three shapes
The first shape is a seat with an allowance behind it and no dollar figure attached. On Claude for Teams and Enterprise plans, each member’s Claude Code usage draws from a per-seat allowance that resets on a rolling five-hour window and a weekly window, shared with Claude chat and Cowork [1]. Usage inside that allowance is not metered in dollars at all [1]. The allowance is the default ceiling, and nothing reaches an invoice until an admin turns on usage credits so members can continue past it [1]. The failure mode here is not a surprise bill. It is a developer who stops at 3pm on a Thursday because the window has not reset.
The second shape is credits with a published dollar value. GitHub Copilot Pro is $10 a month with 1,000 base AI credits and a 500 flex allotment, 1,500 in total; Pro+ is $39 with 3,900 base and 3,100 flex, 7,000 in total; Max is $100 with 10,000 base and 10,000 flex, 20,000 in total [3]. Copilot Business is $19 per granted seat with 1,900 credits and Enterprise is $39 per seat with 3,900, and usage beyond the pool is charged at $0.01 per AI credit [3]. Do the multiplication on those two and the included pool is worth exactly the seat price, $19 and $39, which is a useful anchor: the vendor has priced the allowance at par. Code completions and next edit suggestions are not billed in credits and stay unlimited on every paid plan [3].
The third shape is raw tokens at API rates. Cursor is $20 a month for Pro, $60 for Pro Plus and $200 for Ultra, and splits usage into two pools: a Cursor Models pool covering Grok 4.6, Grok 4.5 and Composer 2.5 with significantly more included usage, and an Other Models pool where third-party models are charged at the model’s API price [5]. Those rates are per million tokens across four dimensions, input, cache write, cache read and output, so Composer 2.5 runs at $0.50 input and $2.50 output while Claude Sonnet 5 runs at $2 and $10 [5]. When the included amount is gone you add on-demand usage at the same rates [5]. Claude Code behaves the same way on the Console or a cloud provider, billed per token to your organisation rather than against a seat [1].
Knowing your shape tells you which surprise to prepare for. Shape one interrupts your work, shape two arrives as an overage line, and shape three has no ceiling at all unless you install one.
The only number that settles it is the cost of one finished task
Vendors publish two kinds of figure and neither is the one you need. Anthropic publishes observed spend across enterprise deployments: around $13 per developer per active day, $150 to $250 per developer per month, with costs staying below $30 per active day for 90% of users [1]. OpenAI publishes the per-task shape instead, saying a typical Codex task on GPT-5.6 Sol may consume between 5 and 30 credits [2]. A six-fold spread on the word “typical” is not a budget. It is a warning that the variance lives inside the task, not inside the month.
The number that settles the argument is your own spend divided by the tasks the agent actually finished. Finished means merged, shipped, or handed to a client, not started. Read one month of spend from the vendor’s usage page, then count the work that came out the other end, including the sessions you abandoned halfway and rewrote by hand. Those cost exactly as much as the successful ones.
The denominator is where this gets honest, and it is the reason almost nobody does it. Counting finished tasks forces you to look at the pile of half-finished ones, and the ratio between the two is the real product review. A tool at $200 a month that finishes 40 tasks is a different purchase from the same tool at $200 finishing six, and no pricing page distinguishes them.
Spend divided by finished tasks. The default spend sits inside the $150-250 per developer per month observed across Claude Code enterprise deployments [1]; replace it with your own invoice. Compare the result with what an hour of your time costs. Computed in the page; nothing is sent anywhere.
What moves the bill inside a session
Context length is the first lever and the one people find last. Claude Code sends your full conversation with every request, and each time the agent uses a tool it sends another request carrying that batch of tool results, so a one-line question in a session that has been open all day still draws usage for the whole conversation [1]. Clearing between unrelated tasks costs nothing and stops you paying to re-read this morning’s dead end all afternoon [1].
Caching is the second. On a Claude subscription the cache lifetime is an hour, dropping to five minutes once you are drawing on usage credits, and five minutes by default on an API key or cloud provider, so the first message after a long break reprocesses your full context [1]. OpenAI’s rate card prices cached input at exactly one tenth of fresh input on every model it lists [2]. Long breaks are expensive in a way that looks like nothing happening.
Output is the third, and it is the stream that costs. Extended thinking tokens are billed as output tokens, and the default budget can run to tens of thousands of tokens per request [1]. On OpenAI’s card, output costs five to six times input on every model listed, with GPT-5.6 Sol at 100 credits per million input tokens against 500 per million output, and GPT-6 Astra at 250 against 1,250 [2]. Anything that shortens the answer moves the expensive number.
Model choice is the fourth and the crudest. Anthropic’s own guidance is that Sonnet handles most coding tasks well and costs less than Opus, which is worth reserving for complex architectural decisions [1]. In Cursor, Composer 2.5 at $0.50 and $2.50 against Claude Sonnet 5 at $2 and $10 is a fourfold difference on both streams for work that may not need the difference [5]. Parallelism is the fifth: agent teams in Claude Code use roughly 7 times more tokens than a standard session when teammates run in plan mode, because each one carries its own context window [1].
The reassuring detail is that waiting is often free. Devin does not accrue consumption while it waits for your response, waits for a test suite to run, or sets up and clones repositories; usage comes from the number and complexity of the actions it takes, with virtual machine time and networking bandwidth typically a small fraction of the total, and Windows sessions consuming about 9% more than equivalent Linux ones [6]. Background processes in Claude Code typically stay under $0.04 per session [1]. The clock is not what you are paying for. The work is.
A vendor’s growth curve is not evidence about your bill
The strongest argument for coding agents in 2026 is commercial rather than technical. Cognition raised at a $26 billion valuation in May 2026 on a $492 million annualised revenue run rate, with enterprise usage of Devin reported growing 50% month over month for the preceding six months, and by August was reported to be in talks at $40 billion on a claimed $1 billion run rate, with Mercedes-Benz, NASA and Goldman Sachs among its listed customers [8]. That is genuine evidence about something. Organisations with security reviews and procurement cycles are signing, and they do not sign for novelty.
It is not evidence that the tool pays for itself in your codebase, and the distinction matters more than the headline. A number built from enterprise contracts tells you what a compliance department approved. It tells you nothing about whether an agent working in a small, idiosyncratic repository with thin test coverage produces work you keep.
A randomised trial points the other way, and the interesting part is not the direction. METR ran one with 16 experienced developers on 246 real issues in open-source repositories averaging more than 22,000 stars and a million lines of code, and found that when AI tools were allowed, developers took 19% longer to complete issues [7]. Those developers had expected AI to speed them up by 24% beforehand, and after experiencing the slowdown they still believed they had been sped up by 20% [7]. The finding worth carrying is not that agents are slow. It is that the impression of speed survives contact with evidence that contradicts it, which is precisely why the cost per finished task has to be arithmetic rather than a feeling.
The unit itself changes underneath you
Whatever you measure, you are measuring against a unit the vendor controls. GitHub has run two in a year. Premium requests are now the legacy system, applying to Copilot Pro and Pro+ subscribers on an existing annual plan who stayed on request-based billing after 1 June 2026, at 300 and 1,500 requests a month with additional requests at $0.04 each [4]. New plans are denominated in AI credits instead [3]. Same product, two units, and a spreadsheet built on the first one silently stops meaning anything.
Prices inside a unit move too. Copilot code review has a model multiplier of 13, so each review consumes 13 premium requests, effective 1 June 2026 [4]. OpenAI states plainly that GPT-5.6 Sol’s promotional pricing is available at least through 21 November 2026, which is a public statement that the price after that date is open, and records GPT-5.4 and GPT-5.4 mini retiring in Codex on 31 August 2026 for users signed in with ChatGPT [2]. A model retirement is a re-pricing event wearing a different hat.
The definitions are worth reading once as well, because they are rarely what you assume. For Copilot agentic features, only the prompts you send count as premium requests and the actions Copilot takes autonomously, such as tool calls, do not [4], which means a long agentic run and a one-line question can cost the same on that meter and wildly different amounts on a token meter. Unused requests from the previous month do not carry over [4]. Put a quarterly reminder in the calendar to re-read whichever page governs your plan. The alternative is finding out from an invoice.
What still goes wrong
Cost per finished task gives you a price, not a value. A task that costs $9 and would have taken you 40 minutes is a good trade; the same $9 for something you would never have bothered doing is a well-measured waste. No dashboard supplies the second half of that comparison, and the temptation once you have a tidy cost figure is to stop before you get to it.
The measurement evidence is thinner than either the optimists or the sceptics admit. The METR trial covered 16 developers working on issues they supplied themselves from large open-source repositories, using early-2025 tooling [7], which makes it a poor guide to what a 2026 agent does in a codebase nobody knows well. The finding that survives is the narrower one about self-report, and it applies to you: your sense of having been sped up is not evidence, and it will not weaken when the numbers disagree with it.
The vendors’ own figures are estimates too. Claude Code computes its dollar figure locally from token counts at list price and Anthropic points to the Console for authoritative billing; if your organisation pays contracted rates, the figures shown to developers will not match the bill unless an admin sets the pricing table in managed settings [1]. OpenAI’s rate card publishes no conversion between credits and dollars at all [2]. Budget in the unit, verify in currency, and do not let a credit count stand in for money in any conversation where the answer matters.
- 01Claude Code Docs — Manage costs effectivelycode.claude.com
- 02OpenAI Help Center — ChatGPT Rate Cardhelp.openai.com
- 03GitHub Docs — Plans for GitHub Copilotdocs.github.com
- 04GitHub Docs — Requests in GitHub Copilotdocs.github.com
- 05Cursor Docs — Pricingcursor.com
- 06Devin Docs — Usagedocs.devin.ai
- 07METR — Measuring the impact of early-2025 AI on experienced open-source developer productivitymetr.org
- 08TechCrunch — AI coding startup Cognition reportedly already in talks to raise at $40B valuationtechcrunch.com