On September 3, 2026, three of the most widely used AI services experienced outages within the same few hours. Claude began returning elevated errors across its website, API, coding tools, and multiple model families. Grok went down across its services. Then ChatGPT and Codex became unavailable for some users.
From the outside, it looked as if the AI internet had failed all at once. The more useful account is less dramatic and more important: three incidents overlapped, but the available evidence does not establish one shared cause. OpenAI attributed its disruption to an internal routing error. SpaceXAI pointed to an outage at its Memphis compute center. Anthropic called its event an infrastructure issue without publicly naming the infrastructure.
That uncertainty is not a footnote. It is the operational lesson. A fallback plan is only as strong as your understanding of what the primary and fallback still share.
What Happened on September 3
Anthropic recorded two incidents that morning. A brief Sonnet 5 event ran from 12:37 UTC (5:37 a.m. Pacific Time) to 12:56 UTC (5:56 a.m. PT). A separate, broader incident began at 13:26 UTC (6:26 a.m. PT), affected claude.ai, the Claude API, Claude Code, and Claude Cowork, and produced elevated errors across several Mythos, Fable, and Opus models. Anthropic said impact ended at 16:16 UTC (9:16 a.m. PT), about two hours and 50 minutes later, and posted its resolved notice at 16:23 UTC (9:23 a.m. PT).
According to reporting based on xAI's status updates, Grok's incident was under investigation from 13:30 UTC (6:30 a.m. PT) until traffic was reported healthy again at 17:05 UTC (10:05 a.m. PT). SpaceXAI later said Grok's disruption followed an outage at its Memphis compute center, apologized to affected compute partners, and said systems had been restored. It did not publish the technical cause of the facility outage.
OpenAI's incident began at 14:43 UTC (7:43 a.m. PT). An OpenAI spokesperson said a routing error made ChatGPT and Codex unavailable for some users. The company reported that a solution had been implemented by about 15:17 UTC (8:17 a.m. PT); its status page remained open for monitoring and marked the incident resolved at 16:55 UTC (9:55 a.m. PT).
The failures also travelled downstream. Cursor's status records tied separate service disruptions to upstream Anthropic and OpenAI issues, and recorded degradation across its Grok-powered features. An API outage rarely stays inside the company whose logo appears on the status page.
What We Know — and What We Do Not
There is a documented relationship behind some of the speculation. Before the outage, Anthropic announced an agreement to use all of the compute capacity at SpaceX's Colossus 1 data center. Reporting identified Colossus 1 as a Memphis facility. On September 3, SpaceXAI said its Memphis compute center had gone down and mentioned unnamed compute partners.
That does not prove Claude and Grok shared the same failure. The Register wrote that SpaceXAI's statement suggested a possible explanation for Anthropic's trouble, not that a common root cause had been confirmed. Anthropic's own incident page said it identified a cause and deployed a fix, but did not name a compute partner or facility. No fetched company statement connected the Claude incident to the Memphis outage.
OpenAI's explanation was different: a routing error, with no reported connection to Memphis. Nor did the major cloud and network status records cited in coverage establish a broader internet infrastructure failure. Based on the public record, the responsible conclusion is that the timing overlapped, one physical facility was involved in Grok's outage, and the relationship between that event and Claude's outage remains unconfirmed.
None of the three providers had published a formal postmortem when this article was prepared. Status updates tell us when services degraded and recovered. They do not give customers a complete dependency map.
Why a Second API Is Not Automatically Redundancy
After a provider outage, the standard recommendation is to add another provider. That is directionally right and operationally incomplete. Two API keys can still lead to the same region, facility, network path, orchestration layer, or capacity partner. Even when the infrastructure is independent, both integrations may depend on one gateway, one authentication service, or one piece of application code.
This is correlated failure: components that look separate from your side of the API boundary can fail together because they share something you cannot see. The September 3 incidents do not prove that correlation caused all three outages. They do show why an architecture diagram that ends at vendor names is not enough.
Real redundancy removes the failure domain that matters. If your primary model stops responding, a second endpoint at the same provider may help with a model-specific problem but not with provider-wide routing. A different provider may help with that, but only if your own gateway and the providers' disclosed infrastructure do not recreate the same dependency underneath.
Design the Failure Mode Before the Failover
The goal is not to promise uninterrupted AI. That promise is rarely credible. The goal is to keep one vendor incident from turning into an application incident with no controlled recovery path.
- Map dependencies beyond the model name. Record the provider, model family, region options, gateway, identity system, data stores, and any disclosed compute or cloud partners. Where a provider will not disclose a layer, mark it as unknown rather than assuming independence.
- Fail over across a genuinely different model family. A backup endpoint is useful, but a tested adapter for another provider and model family covers more failure modes. Run representative evaluations regularly so the fallback is known to meet the minimum job, not merely known to accept a request.
- Degrade gracefully. Decide which features can switch to a smaller model, return cached or rules-based output, accept a manual review, or become temporarily read-only. A controlled reduction in capability is better than a spinner that waits forever.
- Queue work that does not need an immediate answer. Document extraction, classification, enrichment, and report generation often tolerate delay. Put durable jobs on a queue with idempotency and bounded retries so they can resume after recovery instead of disappearing or multiplying.
- Use timeouts and circuit breakers. Stop sending traffic into a failing dependency, protect your own worker capacity, and expose an honest service state to users. Failover logic that waits on the primary until every request times out is not failover.
Test the Path, Not the Diagram
A fallback that has never carried production-shaped work is an intention. Test it with the prompts, context sizes, structured outputs, tool calls, safety requirements, and latency limits your application actually uses. Confirm that credentials are valid, quotas are available, observability follows the request, and operators know whether the switch is automatic or manual.
Then test the degraded experience. What does a customer see when every model is unavailable? What happens to a task already in progress? Can an operator replay queued work safely? Does the system explain that processing is delayed, or does it report success before an answer exists?
These questions make this incident different from another warning about vendor concentration. The hard part is not adding a second model. It is defining what the business does during the minutes or hours when models are unreliable.
What To Do This Week
Choose one AI-dependent workflow that would interrupt operations if it stopped. Trace every dependency from the user action to the final result. Label each one as owned, independently replaceable, shared, or unknown. Then run one exercise with the primary provider disabled.
- Does traffic reach a fallback that has passed the same minimum evaluations?
- If it cannot, does the product preserve the request and communicate a useful degraded state?
- Can the team recover queued work without duplicates or lost jobs?
- Which infrastructure assumptions still depend on a vendor answer you do not have?
September 3 was not proof that every major AI provider rests on one hidden machine. It was proof that simultaneous-looking failures can have different causes, incomplete disclosures, and the same business consequence. Build for that consequence: understand what you can, isolate what you control, and make failure a mode your system can enter deliberately rather than a surprise it can only endure.