One Indie Hacker Logged 5,200 Agent Coding Runs. The Median Cost Was 42 Cents and the Worst Runs Hit $34.80. Here's the Budget Cap You're Missing.
A widely shared Indie Hackers post logged roughly 5,200 agentic coding runs across Cursor, Claude, and Codex on comparable tasks. The median cost per run was 42 cents. That number alone would make agentic coding sound negligible on a per-task basis. The tail is where it stops being negligible: p95 came in at $7.20, p99 at $34.80, and the same post describes a retry-loop bug that silently drained $30 in eight minutes, plus a developer who burned $1,400 in a single week running agent QA directly against a metered API. I run agentic pipelines against metered APIs myself for this site, so the tail-cost risk in that dataset isn't an abstraction to me.
The numbers, with the right caveat
Worth being upfront about what this data actually is: one operator's self-reported logs from a tool called TokRepo, not an audited, cross-platform industry benchmark. Treat the specific dollar figures as directionally useful rather than a number to build a pricing model around. That said, the shape of the distribution, a low median with a long, expensive tail, matches what anyone who's run agentic coding tools against a metered API has probably felt anecdotally, even without logging every run.
The median, 42 cents per run across Cursor, Claude, and Codex on comparable tasks, is the number that gets quoted when someone argues agentic coding is basically free at indie scale. It's not wrong, exactly. It's incomplete. A distribution with a 42-cent median and a $34.80 p99 isn't telling you "this costs about 42 cents." It's telling you most runs are cheap and a meaningful fraction are wildly not, and you don't get to pick in advance which one you're about to run.
Why the tail is the actual risk
A retry loop is the specific failure mode worth understanding, because it's the mechanism that turns a normal task into a p99 outcome. An agent hits an error, retries, hits a similar error, retries again, and if nothing's stopping it, that loop can burn through API calls fast enough to drain $30 in eight minutes without you noticing until you check the bill. It's not that the agent is doing anything exotic. It's doing exactly what it's built to do, just persistently, against a meter that's ticking the entire time.
The developer in that same post who burned $1,400 in a week running agent QA directly against an API wasn't making an obviously bad decision task by task. Each individual run probably looked reasonable in isolation. The cost came from volume and a few expensive outliers compounding over a week without a ceiling on the total.
The flat-rate-versus-API tradeoff this actually exposes
This is where the "per-task cost comparison" that dominates most AI-tool pricing posts misses the point. Comparing Cursor Pro at $20 a month against metered API pricing on a per-task basis makes the flat-rate plan look expensive most of the time, because most tasks really do cost pennies against the API. What that comparison leaves out is the ceiling. A flat-rate plan trades some speed or rate-limiting during heavy use for a hard cap on total spend. Metered API access trades that ceiling away in exchange for uncapped throughput, and uncapped throughput cuts both ways when a retry loop is the thing doing the throughput.
For a solo operator, the value of a hard ceiling is usually underpriced in these comparisons, because a $20 monthly cap that you hit and get slowed down by is an annoyance. A metered API bill that hits $1,400 in a week because of a bug you didn't catch fast enough is a different category of problem entirely.
What I'd actually do
Set an actual spend cap on any API-billed agent usage, not a mental budget, an enforced one. Most API providers, Anthropic and OpenAI included, support hard usage limits or billing alerts at the account or key level. Set one below what you're comfortable losing to a bad week, not below what you expect to spend on a good one. If you're running agent QA or any workload with retry logic against a metered API, add an explicit retry cap in your own code on top of the provider-side limit, since a provider-level spend cap will still let a retry loop burn through a meaningful chunk of your budget before it trips.
And if your usage pattern looks more like steady daily coding than occasional heavy bursts, a flat-rate plan is probably the safer default even when the per-task math says the API is cheaper. The 42-cent median is real. So is the $34.80 tail. Price the plan you're comfortable being wrong about, not the plan that wins the average case.
The honest take
I don't think the answer is "never use metered APIs," and this site's own pipeline runs on one. The answer is that a spend cap should exist before the first run, not after the first bad week. Where this dataset is weakest: it's one operator's logs across specific tasks, and your own task mix, prompt patterns, and error rates could look nothing like TokRepo's. Treat the specific dollar figures as a prompt to check your own numbers, not as your numbers.
Author
Lukas
@lukcombinator