BSS Portal degraded performance issues
Updates
Root Cause Analysis
Incident: BSS Portal degraded performance issues
Ref. issue no: DOPS-9882
Incident summary
Platform users started to experience degraded performance and intermittent failures between August 26, 2026, 23:21 UTC August 26, 2026, 06:20 UTC. Degradation was particularly noticeable when using the BSS Dashboard and Add Payment functionality. The underlying database server experienced sustained high resource utilization, which caused delayed responses and partially affected service availability. Normal performance was restored after the affected database sessions were gracefully terminated.
Customer impact
- BSS Dashboard pages responded slowly or failed to complete some requests.
- Add Payment requests experienced slow responses or failures during the incident window.
- The service was degraded for a duration of approximately 6 hours and 59 minutes. The number of affected users and unsuccessful payment attempts has not yet been confirmed.
What happened
Database queries supporting the BSS Dashboard and Add Payment functionality did not use the available database access paths efficiently. As request volume accumulated, these queries consumed increasing database resources and caused broader performance degradation. Once the team identified the affected sessions, terminating them relieved the immediate resource pressure and restored normal response times.
Root cause
Inefficient execution of database queries used by the BSS Dashboard and Add Payment functionality gradually led to sustained high CPU utilization and resource contention on the shared database server.
Contributing factors
- Existing monitoring did not generate an early internal alert for the sustained database resource condition as the rate of requests was normal and degradation increased gradually.
- The incident was not reported as critical; thus, the escalation workflow did not trigger an immediate after-hours response, extending the time before investigation began.
- Additional safeguards, such as query timeouts or resource controls, were not in place to limit the impact of unusually expensive queries.
Resolution and recovery
The response team identified the database sessions consuming the most resources and began graceful termination of the affected sessions, while validating that terminated sessions did not cause additional failures. Database CPU utilization and query response times returned to normal levels, restoring the affected functionality by 06:20 UTC. Performance was monitored following recovery and remained stable based on the information available for this report.
Corrective and preventive actions
| Workstream | Commitment | Status |
|---|---|---|
| Query and database optimization | Review execution plans and optimize the affected queries, indexes, and statistics so requests use efficient access paths | In progress / target date: Sep. 11, 2027 |
| Proactive monitoring | Add alerts for sustained CPU utilization, in conjunction with long-running queries to support earlier detection | Planned / target date: Sep. 4, 2027 |
| Resource safeguards | Assess query timeouts and database resource controls to reduce the impact of runaway workloads | Under review |
| Validation and readiness | Performance-test the Dashboard and Add Payment paths under production-representative load and document a response runbook | Under review |
Timeline
| Time | Event |
|---|---|
| 26 Aug, 23:21 UTC | Performance degradation begins |
| 27 Aug, 01:41 UTC | Customer support ticket reports an urgent service issue |
| 27 Aug, 05:15 UTC | Technical investigation begins |
| 27 Aug, 06:20 UTC | Affected sessions are terminated and normal performance is restored |
Service performance has been restored for the affected platform functionality. Our investigation found that inefficient database query execution caused elevated resource utilization and degraded response times. We mitigated the issue by terminating the affected database sessions, and performance returned to normal levels at 06:20 UTC.
Monitoring/Recovered: performance issues recovered; continuing investigation into root cause
Our team identified degraded performance issues on BSS Portal started by 26 August 2026, 23:21 UTC
← Back