ChatGPT, Claude, and Grok Went Down at the Same Time Yesterday. Your Fallback Plan Was Never a Plan.
ChatGPT, Claude, and Grok all had outages within the same window on Thursday, September 3. OpenAI's routing error started around 7:43am PT and knocked out ChatGPT and Codex for a chunk of users; Anthropic's Opus 4.8 and Opus 5 models were degraded around the same time while its other models held at baseline error rates; Grok went down too, with xAI confirming it was working the problem. Microsoft Azure had disruptions reported in the same window that may have contributed across more than one of these providers. Three separate companies, three separate model stacks, one bad morning.
If you build on any single one of these APIs, the instinct after reading that is "good thing I'd switch to a different provider if this happened." Yesterday is the specific, dated proof that instinct doesn't hold when the failure is a layer down from the model.
What actually broke
The public reporting points to two different kinds of failure landing close together: OpenAI's own routing layer had a defect that made ChatGPT and Codex return errors for a subset of users starting around 7:43am PT, with a fix rolled out and being monitored by about 8:17am PT the same morning. Separately, Anthropic reported elevated errors specifically on Opus 4.8 and Opus 5, while its other model tiers stayed at normal error rates, suggesting a narrower, model-specific issue rather than a full platform outage. Grok's outage was acknowledged by xAI without a detailed public root cause at the time of writing. Reports also flagged concurrent Microsoft Azure disruptions as a possible shared contributing factor, though the extent of that overlap across all three providers isn't fully confirmed in public reporting as of this writing.
The mechanism matters less than the pattern: none of these are the kind of failure a second model subscription protects you from. If the shared layer is cloud networking, DNS, or a routing tier underneath the API you're calling, "I have a Claude key and an OpenAI key" doesn't help if the thing that broke sits below both of them.
Why "I'll just switch providers" isn't a real fallback
Most solo operators who've thought about this at all have landed on some version of "if my primary API goes down, I'll fail over to a different one." That's a reasonable instinct and a genuinely bad plan as usually implemented, for three reasons. First, most people who say this have never actually wired up the failover path in code, so on the day it matters they're doing it live, under pressure, for the first time. Second, a same-day, same-morning outage across three unrelated providers shows the failure modes aren't as independent as "different company" suggests, shared cloud infrastructure, shared upstream dependencies, and correlated demand spikes during any one provider's incident can take out more than one API at once. Third, even a working failover to a different model provider doesn't help if your product's specific behavior, prompt structure, or fine-tuning is tied to one model's quirks; a hot swap to a different model can silently change your output quality in ways you won't catch until a user does.
What an actual fallback looks like
A real fallback plan for a solo stack has three layers, in order of how often you'll actually need them. Cache and serve stale-but-good responses for anything that doesn't strictly require a fresh generation, this covers more traffic than people assume, especially for content that doesn't change minute to minute. Degrade gracefully instead of hard-failing: a feature that says "AI suggestions are temporarily unavailable" and falls back to a non-AI path beats a 500 error every time, and it's a small amount of extra code most people skip because it doesn't feel urgent until an outage happens. And yes, have a second provider wired up and tested, not just planned, with a scheduled quarterly check that the failover path actually still works against the current API version.
What I'd actually do
Pick whichever of the three above you don't currently have, and build it this week rather than after your next outage. If I had to rank them for a typical solo AI feature: graceful degradation first, because it's the cheapest to build and covers the most outage scenarios regardless of cause; caching second; a tested second-provider failover third, because it's the most work and the least frequently the thing that actually saves you, given how often outages are correlated across providers rather than isolated to one.
The honest take
I don't think most solo operators need multi-cloud, multi-region, five-nines infrastructure for a side project or a small SaaS. That's overbuilding for a risk that, most days, doesn't materialize. What yesterday changes is the assumption that a second API key is meaningfully better than no plan at all. It might be. It might also fail at the exact same time, for a reason that has nothing to do with which model you picked.
Author
Lukas
@lukcombinator