· 7 min read

OpenAI, Anthropic and Google Have Been Coordinating on AI Safety Since July. Here's What That Means for Anyone Building on Their APIs.

OpenAI's policy chief told reporters on September 15 that his company has been in talks with Anthropic and Google DeepMind on AI safety since July. Not a press release, not a joint statement, a confirmation dragged out by a TechCrunch question after weeks of the three companies quietly building something together while their labs kept shipping competing models on the usual schedule. If you build anything on top of GPT, Claude, or Gemini, this is the part of the news cycle worth reading past the headline.

What was actually confirmed

Chris Lehane, OpenAI's global policy chief, said the coordination has centered on a handful of concrete pillars: shared technical evaluations, pre-release audits of advanced models, independent testing frameworks, and standardized safety protocols across labs that otherwise spend their marketing budgets telling you why their model is different from the other two. OpenAI has also thrown its weight behind a provision in the proposed FRONTIER Act that would let independent verification organizations into frontier labs to check how models are built, not just what the labs say about them.

None of this is signed yet. As of the report, there's no binding agreement, and antitrust questions (three competitors coordinating on anything invites exactly that scrutiny) are still unresolved. What's confirmed is the direction: labs that spent the last two years outcompeting each other on capability are now spending real time on shared safety infrastructure, at the same moment they're still racing on everything else.

The trigger was one essay, three days earlier

The timing isn't a coincidence. On September 12, Anthropic CEO Dario Amodei published an essay laying out a three-step plan for the industry: independent evaluators, international coordination, and global safety agreements. He argued the industry needs to actively slow the pace of frontier AI development to avoid what he characterized as catastrophic risk, a notable position from the CEO of a company whose entire business model depends on shipping frontier models faster than its two biggest rivals.

Three days later, OpenAI confirms it's been talking to Anthropic and Google about exactly this since July. Either the essay accelerated conversations that were already underway, or it gave OpenAI a clean moment to admit something it had reason to keep quiet about until a competitor said it out loud first. Either read gets you to the same practical place.

The same week, all three shipped safety tooling

This wasn't just talk. In the same window, Google released Gemini 3.8 Flash Cyber, a cybersecurity-focused model built for autonomous vulnerability discovery, positioned as outperforming Anthropic's and OpenAI's comparable offerings. Anthropic built and deployed a classifier that detects and blocks sandbox escape attempts, and separately changed how it structures reward specifications so agents can't shortcut their way to a goal without actually doing the work. Google also rolled out Enterprise Frontier Safeguards, pairing zero data retention with misuse-detection tooling aimed at enterprise customers.

Read individually, these look like three companies shipping normal product updates. Read together, in the same week as the safety-talks confirmation, they read like the visible output of the coordination Lehane described: labs converging on a shared idea of what "safe enough to ship" looks like, then each building their own version of the enforcement layer.

What changes if you build on these APIs

Here's the part that matters if your product calls Claude, GPT, or Gemini in production. Shared safety evaluations and cross-lab audits don't stay abstract. They turn into classifiers that reject more of what your agent tries to do, rate limits that tighten around behavior patterns the labs have decided look risky, and occasional silent changes to model behavior that show up as your eval suite failing on a Tuesday for no reason you can find in a changelog. I've had exactly this happen twice this year: a workflow that worked fine for months started getting flagged by a safety filter with no announcement, and the fix was rewriting the prompt to avoid a trigger I had to reverse-engineer through trial and error.

That's not a complaint about safety work being bad. Sandbox escapes and reward hacking are real failure modes, and a classifier that catches them is doing its job. It's a statement about what "more coordinated industry safety standards" actually looks like from the outside: less visibility into why your agent's behavior changed, and less certainty about which lab's policy update, of three, caused it.

The honest take

I don't think this coordination is bad for the industry. I think Amodei is right that nobody wants a race-to-the-bottom on safety among three companies capable of shipping models that can meaningfully act in the world. But I also think "the three frontier labs are coordinating on safety standards" and "your production agent workflow is now more fragile to changes you don't control" are both true at the same time, and most coverage of this story only wants to talk about the first one.

What I'd actually do if a meaningful share of your revenue depends on one of these APIs: keep a working integration with at least one open-weight model, even if you never route production traffic through it day to day. Not because open weights are safer (they're not necessarily), but because the day a shared safety classifier decides your specific workflow looks like something it should block, you want a fallback that took an afternoon to wire up instead of a week you don't have. I did this for a client-facing agent in August after the second silent-behavior-change incident, and it's paid for the setup time twice since.

If you're not running anything that touches sandboxes, code execution, or agentic tool use, this probably doesn't touch you yet. If you are, this is the week to write down exactly what your agent is allowed to do and why, because you may be explaining it to a rejected API call sooner than you'd like.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts