One afternoon, three AI outages: what the OpenAI, Claude and Grok status pages told customers
Robin Geall
·12 min read
Retained public incident entries, checked at 09:00 UTC on 4 September 2026. Timestamp bases differ between providers.
On 3 September 2026, the public status pages for Claude, OpenAI and Grok all reported service problems during the same afternoon.
For customers trying to finish a task, the first question was simple: is the problem on my side or theirs?
The three status sites answered that question. But they offered very different levels of help after that.
At 16:06 UTC, Claude said it had deployed a fix and was monitoring recovery. OpenAI said it had applied a mitigation. Grok’s latest message still said it was investigating an outage first reported at 13:30.
Claude gave customers a changing picture of the incident, including which products and models were affected, what had recovered and when the impact ended. OpenAI provided a broad service list and finished with useful advice for some Codex users, but its middle update were limited. Grok acknowledged the outage, then published no intermediate update for more than three and a half hours.
This is a comparison of public incident communication, not a reliability ranking. A status page cannot tell us what happened in internal channels, and there is no evidence that these incidents shared a cause. They simply overlapped.
How we compared the status pages
We reviewed the first-party public records at 09:00 UTC on 4 September 2026.
We looked at whether each status page helped customers answer five practical questions:
Is the provider aware of the problem?
Which products, features or users are affected?
What has changed since the previous update?
Is there anything I need to do?
When will I hear more?
We did not attempt to score acknowledgement speed because the providers did not consistently state when customer impact began. We also treated their severity labels separately. “Major,” “degraded performance” and “outage” are not standardized measures across three different status systems.
The initial notice and final resolution are included in the update count.
Public incident communication on 3 September 2026. All times are UTC.
Mitigation applied, monitoring, residual user action
No intermediate progress update
Firm next-update time
No
No
No
Our assessment
Clearest changing picture of scope, recovery and when impact ended
Broad consolidated scope, but a sparse and hard-to-reconstruct timeline
Clear opening and closing, with no intermediate progress update
The providers expose different timestamp information, so the interval row is not a controlled comparison of live publication cadence. Claude’s retained update times match its record creation times. OpenAI exposes different creation and display times. Grok’s service pages show the displayed timeline.
Claude’s incident began with an investigating update at 13:26. The notice named three affected models and attached the incident to claude.ai, the Claude API, Claude Code and Claude Cowork.
Fifteen minutes later, the team said it had identified the cause and was working on a fix. At 13:50, it expanded the model list to provide what it called an exhaustive account of the affected models.
That additional scope was useful. Customers could compare the list with the model supporting their own work rather than assuming every part of Claude was equally affected.
At 14:49, Claude said work on a fix was continuing. The message did not add much, but it stopped the timeline from going completely quiet.
At 15:25, Claude reported that most models had returned to their normal error rate and that only Opus 4.8 and Opus 5 remained affected. It reported a deployed fix at 16:06, followed by resolution at 16:23. The final message also stated that impact had ended at 16:16.
What worked
Seven updates would mean little if all seven said the same thing. Claude’s record was useful because the information changed.
The timeline moved from an initial model list to wider scope, partial recovery, a deployed fix and an explicit end to customer impact. A customer using a recovered model could make a different decision from someone still relying on an affected version.
What was missing
There were still gaps.
Claude said it had identified the cause without explaining what that cause was. Withholding an unverified or sensitive technical explanation during an active incident can be responsible, but the final message could have said whether further analysis would follow.
The team also promised updates “as soon as possible” rather than giving customers a specific time. The longest interval between retained updates was 59 minutes, which is a long wait during an incident classified as major.
Even with those limitations, Claude gave customers the clearest account of what was changing.
OpenAI: broad scope, but a harder timeline to reconstruct
What customers saw
OpenAI’s incident covered elevated errors across ChatGPT and Codex. Its page grouped 15 affected ChatGPT components and four Codex components under one record.
The visible timeline shows an investigating message at 14:43, a mitigation and monitoring message at 15:17 and resolution at 16:55. OpenAI classified the event as degraded performance, while the status API recorded its impact as minor.
The resolution message included the most directly actionable instruction in any of the three incidents: some Codex remote-control users might need to pair their mobile device again.
“Resolved” does not always mean every customer can immediately carry on. Calling out a remaining step can prevent repeated failures and unnecessary support tickets.
What worked
Grouping the affected components under one incident was a good structural decision. Customers could see the breadth of the event without searching through separate incident pages.
What was missing
The public record left two gaps that made the incident harder to follow.
The timestamps tell two different stories
There is an important wrinkle in OpenAI’s record.
The status page displays the first two updates at 14:43 and 15:17. However, the public incidents API records them as created at 14:58 and 15:50. The resolution was created and displayed at 16:55.
The earlier records were also edited later.
There may be a reasonable explanation. The displayed values could represent when the team later determined that an event occurred or a mitigation was applied. But the customer-facing incident page does not explain why those timestamps differ.
That matters because the retrospective page appears to show a 34-minute interval before mitigation, followed by 98 minutes until resolution. Using the record creation times, the intervals between entries were approximately 52 and 65 minutes.
A customer looking at the page afterwards cannot easily tell when each update became available to read.
Status pages should separate three different moments:
When customer impact began
When an operational event occurred
When the public update was published
Those times will not always be the same, but they should be labelled clearly.
Where OpenAI’s updates were thin
The first message said OpenAI was investigating the issue for the listed services. The second said a mitigation had been applied and recovery was being monitored.
Neither explained which customer experiences were failing, what was recovering or what remained affected. The component list supplied breadth, but the messages offered little detail about customer impact.
A stronger middle update would have combined the mitigation milestone with observable information: which services were recovering, what problems customers might still encounter and when the team would report again.
Grok: a clear start and finish, with little between
What customers saw
The records linked from Grok’s affected service histories use a common 13:30 start time.
Most of the opening messages said Grok was experiencing issues and that the team was working to restore service. The Android service-history record used shorter wording, saying only that there were issues with its models.
The next entries said the situation had been resolved and traffic was healthy again.
Depending on the service, those messages arrived between 17:04 and 17:09. The Grok web record shows a duration of 3 hours 37 minutes. Grok in X shows 3 hours 35 minutes, while the US West API record shows 3 hours 39 minutes.
The opening and closing messages were clear. But there was nothing between them.
What worked
Grok deserves credit for acknowledging the outage and closing the records linked from its service histories.
What was missing
Customers were not told whether the problem had been identified, whether a mitigation was underway or whether parts of the service were recovering.
Saying that the team is working to restore service provides some reassurance. It does not provide progress.
Even if there is nothing new to report, an update can still reduce uncertainty:
That message does not reveal sensitive technical information or promise a recovery time. It simply gives customers a reason not to refresh the status page every few minutes.
Separate records fragmented the story
Grok recorded the event in eight service-history records covering web, X, mobile apps, Grok Build, workspace plugins and two US API regions.
Service-specific records can support targeted subscriptions. During a broad event, however, they can make customers reconstruct the overall picture from several separate histories. A parent incident with affected services attached would provide a clearer source of truth while still allowing targeted notifications.
The Android history created an additional complication.
Its service-history record showed recovery at 17:04. But when we checked the following morning, a second official Android incident record, created at 13:43, still displayed a single unresolved outage update.
The Android service history showed the service as resolved and currently operational, so this appears to be a duplicate or superseded record rather than evidence of a continuing outage. Customers should not have to make that interpretation themselves.
Of the three providers, though, it did the least to reduce uncertainty while the incident was active.
Five lessons for incident communicators
1. Describe customer impact, not just internal activity
“We are investigating” is a useful first message. Later updates should add what customers can observe: failed requests, unavailable conversations, affected models or recovering regions.
You do not need to know the root cause before you communicate early during an incident. You only need to confirm that customers are experiencing a real problem.
2. Show what changed
Each update should add at least one useful piece of information. Claude’s narrowing model list is a good example. OpenAI’s mobile re-pairing instruction is another.
If nothing has changed, say so. A no-change update still reassures customers that the incident has an owner.
3. Promise the next communication, not the resolution
Resolution times are difficult to predict. The next public update is under your control.
“We will update again by 15:30 UTC” gives customers a clear expectation without inventing an estimated recovery time. None of the three providers made that commitment.
4. Keep one trustworthy history
If an incident affects several products, attach them to one shared record where possible.
If separate records are necessary, link them and use consistent times and wording. Clearly distinguish when impact began, when each message was published and when impact ended.
5. Give root-cause analysis time
Live updates and post-incident reviews have different jobs.
During an incident, customers need acknowledgement, scope, progress and practical guidance. A later post-incident review can explain the cause, lessons and follow-up actions with greater confidence.
None of these incident records contained a post-incident review when we checked. We have not marked the providers down for that. A careful explanation published later is more valuable than fast speculation.
Frequently asked questions
How often should a status page be updated during a major incident?
There is no interval that fits every incident, but teams should decide their cadence in advance. For a high-impact event, 15 minutes can be a useful starting point. Meet the expectation you set, even if the next message only says the investigation continues.
What should the first incident update include?
Confirm the customer-visible problem, name the affected services, include the approximate start time if known and give a time for the next update. You do not need to wait for a root cause.
Should a status page publish the root cause?
Not before the cause has been verified. During the incident, focus on impact and recovery. At resolution, publish a short confirmed explanation or say whether a fuller review will follow.
How should incident timestamps be displayed?
Label the start of customer impact, each update’s publication time and the end of impact separately. If a message refers to an earlier operational event, put that event time in the message rather than silently changing its publication time.
The verdict: Claude gave customers the most useful record
Claude’s updates helped customers follow a changing incident. OpenAI provided valuable scope and a useful final instruction, but its timeline was sparse and difficult to reconstruct. Grok offered clear opening and closing messages, but little help between them and a fragmented history across affected services.
This conclusion concerns three providers’ public incident communications during one afternoon. It does not prove that one provider is generally more reliable, transparent or effective at incident response than another.
The wider lesson is more durable: the best status update is written for the person whose work has stopped. Tell them what is affected, what has changed, what they can do and when you will speak again.
For a practical framework your team can prepare before its next outage, read Incident Communication 101.
Incident communication without the chaos
Incident updates are written under pressure. Sorry™ helps teams use prepared templates, select affected components, notify the relevant subscribers and move an incident clearly through investigating, identified, recovering and resolved.