· 9 min read

OpenAI's new Agents API rents out the Codex harness. Here's what you give up

On September 10, 2026, OpenAI opened its Agents API to public beta. The pitch is straightforward: the same session orchestration, sandboxing, and context-compaction engine that runs Codex is now available behind a single API call, to any developer, for no extra platform fee. You pay for model tokens, tool calls, and sandbox container time, and that's it. I spent a chunk of the last month sketching my own version of exactly this: a loop that keeps agent state alive across turns, retries failed tool calls, and trims context before it blows the window. Watching OpenAI ship it as a product feature made me rethink what I was actually building.

What the Agents API actually gives you

Strip away the marketing and the Agents API is four objects: an Agent (model, instructions, tools, MCP servers), an optional Environment sandbox, a Session that persists across turns, and the stream of events that Session emits. OpenAI's own announcement is specific about what that buys you. Sessions survive across long-running tasks and automatically compact earlier context as you approach the model's context limit, so you're not hand-rolling summarization logic. Multi-agent support lets a parent agent delegate pieces of a task to subagents that run with their own context and get reconciled back into the main thread, a pattern several of OpenAI's launch partners specifically credited for the biggest speed gains: one, Ciridae, said subagent orchestration alone cut their latency by 4x. Tool search loads tool definitions on demand instead of stuffing every schema into the prompt, and programmatic tool calling lets the model chain and filter tool results in code before anything comes back into context.

For execution, you choose the sandbox. OpenAI will host it for you on the same infrastructure that runs Codex and ChatGPT, or you can run it in your own VPC, or hand it to a launch partner like Cloudflare, E2B, Modal, or Vercel. That's a meaningfully different posture than "trust our cloud with everything," and it's the part of this launch I actually respect: OpenAI is selling the harness, not insisting you also buy the compute.

The pricing shape: no markup where you'd expect one

Here's the part worth sitting with. OpenAI says explicitly there are no additional fees for the Agents API itself, you pay standard token and tool rates. That's unusual for a company that could have charged a per-session orchestration premium and probably gotten away with it for a while. Instead the bill comes from three places that stack: model tokens (at whatever rate your chosen model bills), tool calls, and, if you use the OpenAI-hosted sandbox, container time billed by the minute with a five-minute minimum per session.

The token math still matters a lot, because those durable sessions and subagent conversations chew through context fast. If your agent runs on GPT-6 Astra, current public pricing trackers put it at roughly $10 per million input tokens, $1 per million cached input tokens, $12.50 per million for cache writes, and $50 per million output tokens, with a 1.05 million token context window. Past 272,000 tokens in a single request, several of those pricing pages report the rate doubling on input and rising 1.5x on output. None of that is an Agents API fee. It's just what happens when you run long, compacted, multi-agent sessions on a frontier model: the orchestration is free, the thing being orchestrated is not.

Why this matters if you were about to build your own loop

I'm a solo builder. The unglamorous 80% of any agent product isn't the prompt, it's the plumbing: keeping session state alive between requests, retrying a tool call that timed out without duplicating side effects, deciding when to compact context and what to throw away, routing between tools without blowing the token budget. I had a rough version of this working for a client project, and it was already the most fragile part of the codebase. Every edge case (a tool that hangs, a subagent that returns malformed output, a session that needs to resume after a deploy) turned into its own small nightmare.

The Agents API replaces that plumbing wholesale. If your product's value is in the tools you expose, the data you connect, or the workflow you automate, and the orchestration was just overhead standing between you and shipping, this is a genuinely good trade. You get Codex-grade session handling, compaction, and subagent coordination, maintained and improved by OpenAI with every model release, for the cost of the tokens you'd be spending anyway.

The honest take

Here's my discomfort with it. For a lot of solo AI products, the orchestration loop wasn't overhead, it was the moat. If your pitch to users or investors was "we built a reliable agent that manages long tasks without losing the thread," that reliability used to be hard-won engineering: months of tuning retry logic, compaction heuristics, and tool routing that a competitor couldn't trivially copy. OpenAI just packaged that exact capability behind one API call and made it available to anyone with an API key. If your differentiation lived in the harness rather than in the tools, the data, or the domain expertise wrapped around it, that differentiation is now a commodity your competitor can rent for the same price you pay.

There's also a dependency risk that's easy to underweight when a launch is fresh and the pricing looks generous. This is a public beta from a single vendor. The harness is open source on GitHub, which helps, but the hosted session infrastructure, the compaction algorithm, and the multi-agent runtime are not something you control. OpenAI could reprice the sandbox, change compaction behavior in a way that alters your agent's outputs, or deprecate a feature your product depends on, and you'd find out from a changelog rather than a design review. I could be wrong about how much this matters in practice: plenty of successful products are built entirely on top of vendor infrastructure they don't control (every Stripe-based business, for instance), and "rented plumbing" hasn't stopped companies from building real moats elsewhere in the stack.

What I'd actually do

Use the Agents API for prototyping, internal tooling, and getting an MVP in front of users fast. It will save you weeks, and weeks matter more than architectural purity when you're validating an idea. But if the orchestration logic itself was supposed to be your product's defensible core, be honest with yourself about what's left once it's commoditized. Build the parts that are actually yours (the tools, the domain-specific evaluation, the data you've collected, the workflow only you understand) on top of the rented harness, and keep enough internal knowledge of what a fallback orchestration layer would look like that a repricing or deprecation doesn't strand you. I'm using it for a client prototype this month. I'm not rebuilding my own product's core loop on it yet, and I'd want to see a few more months of stable pricing and behavior before I did.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts