Platform incident - Service disruption
Updates
Root Cause Analysis
Incident: Platform incident - Service disruption
Incident summary
Platform users started to experience degraded performance and intermittent failures between October 6, 2026, 13:12 UTC - October 6, 2026, 13:39 UTC. The disruption occurred on both BSS and Storefront sites.
Customer impact
- Connection timeouts while loading BSS pages.
- Connection timeouts while loading Storefront pages.
- Certain operations, such as order submissions, or subscription updates/cancellations failed or remained in a pending status
What happened
A node disk failure caused disruption to the core platform gateway API endpoint. This endpoint serves requests that are related to the platform’s message broker system which is utilized by both Storefront and BSS web applications. Although the endpoint did not fail completely, most requests served within the incident duration failed or entered a pending state, which in turn caused application errors and timeouts while using the platform sites.
Root cause
Infrastructure-level failure disrupted availability of platform’s core gateway API endpoint.
Contributing factors
Resolution and recovery
The response team identified the errors originating from the problematic endpoint and proceeded immediately with re-provisioning and recovery to a healthy node. These actions restored the API endpoint functionality and proper connectivity was re-established between the BSS and Storefront web applications.
Corrective and preventive actions
| Workstream | Commitment | Status |
|---|---|---|
| Increase API endpoint redundancy | Added multiple active API endpoint instances across different nodes to increase redundancy and not rely solely on failover process during a hardware/infrastructure failure | Completed |
Timeline
| Time | Event |
|---|---|
| 6 Oct, 13:12 UTC | A node disk failure occurs |
| 6 Oct, 13:14 UTC | First performance and connectivity alerts are received |
| 6 Oct, 13:19 UTC | Technical investigation begins |
| 6 Oct, 13:30 UTC | Failed API endpoint is re-provisioned to another node |
| 6 Oct, 13:39 UTC | Platform sites connectivity is restored |
On October 6 2026, users in the EU region may have experienced a performance degradation and connectivity issues impacting platform operations. The issue has now been resolved, and the service is operating normally for all affected customers.
Mitigations have been deployed and we are seeing positive signs of recovery across all affected platform operations. We are continuing to monitor system health to ensure full stability and will share our next update within the next hour.
We have identified the cause of the disruption and our teams are diligently working on mitigations. We’re currently seeing signs of improvements. We’ll continue to share additional updates here as more information becomes available.
We are investigating an issue that affects a subset of Storefront and BSS sites in the EU region. Affected users may experience performance degradation, connectivity issues, and delayed responses or errors in certain operations.
← Back