Tech

Major AI Outage Hits ChatGPT, Claude, and Grok Simultaneously

The outage began in the early morning hours Eastern Time. By 7:53 a.m. ET, outage-tracking data indicated that more than 5,000 reports related to ChatGPT had been logged. This number continued to climb rapidly, passing 22,000 reports within two hours. By mid-morning, combined reports for ChatGPT and OpenAI’s coding tool, Codex, topped 66,000. OpenAI’s own status page confirmed “elevated errors” across both services at approximately 10:58 a.m. UTC (6:58 a.m. ET), noting that the disruption affected 15 separate components on ChatGPT and four on Codex.

While ChatGPT bore the brunt of the public backlash, rival systems also suffered. Anthropic’s Claude peaked at 1,324 user reports just before 11 a.m. ET, and xAI’s Grok recorded 1,365 reports around 10 a.m. ET. Microsoft’s Copilot assistant was also flagged for stability and downtime issues during the same window. In contrast, Google’s Gemini, which operates on Google Cloud rather than Microsoft Azure, remained largely operational, with only about 500 reports at its peak. Cursor, an AI agent that relies on Grok and Claude, also confirmed an outage resulting from the upstream failures.

Infrastructure Concentration Risk

The simultaneous nature of the failures points to a shared underlying cause in cloud infrastructure. Reports indicate that the trigger was a regional failure in Microsoft Azure’s East US infrastructure. OpenAI, Anthropic, and xAI all rely on Azure-hosted compute and networking for at least part of their production stacks. When that regional infrastructure degraded, the failure propagated to each of the competing services at roughly the same time.

Photo by Brett Jordan / Pexels

This incident underscores a critical vulnerability in the current AI landscape: despite fierce competition over model quality and pricing, major providers share significant infrastructure exposure. A single-region cloud failure, which is typically a contained event that degrades service for specific workloads, became a systemic issue because multiple independent companies had meaningful production dependencies in the same region of the same cloud provider.

Other infrastructure providers also showed symptoms during the window. Cloudflare, which routes and caches traffic for a large share of the internet, and Amazon Web Services were reported to be experiencing issues, though neither was identified as the root trigger. The primary cause remained attributed to the Azure East US breakdown.

Recovery and Official Response

By 12:38 p.m. PT (3:38 p.m. ET), all services were reported to be back to normal across ChatGPT, Claude, and Grok. OpenAI stated that a fix had been applied and services were recovering. Anthropic noted that most models had recovered, with the exception of specific versions, while xAI confirmed the outage had resolved.

Photo by Ivan Chumak / Pexels

However, the official explanation for the root cause remains contested. While multiple tech outlets and industry analysis pointed to the Azure East US failure as the trigger, Microsoft has denied that Azure was the cause. Microsoft’s Azure status history logged no East US incident for September 3, and the company stated that the outage was not related to its cloud services. This discrepancy leaves the precise technical origin of the failure unresolved, though the pattern of simultaneous disruption among Azure-dependent services remains a clear indicator of shared infrastructure risk.

The event serves as a stark reminder of the fragility of the global AI ecosystem. As businesses and individuals increasingly rely on these platforms for critical tasks, the lack of diversified infrastructure among top-tier AI developers poses a significant operational risk. The incident has prompted renewed scrutiny of how major tech companies manage their cloud dependencies and whether the current concentration of critical AI workloads in specific regions presents an unacceptable level of systemic vulnerability.

Emma Watson

Emma Watson reports on technology with interests spanning artificial intelligence, consumer technology, online security, digital platforms, and major industry developments. She follows new products and services alongside the policies and business decisions influencing them. Emma's approach emphasizes clear explanations, reliable sourcing, and practical context for readers trying to understand how technology is changing.

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button