BSS Portal degraded performance issues

Minor incident BSS AU region
2026-08-27 02:21 EEST · 1 day, 13 hours, 23 minutes

Updates

Post-mortem

Root Cause Analysis

Incident: BSS Portal degraded performance issues
Ref. issue no: DOPS-9882

Incident summary

Platform users started to experience degraded performance and intermittent failures between August 26, 2026, 23:21 UTC August 26, 2026, 06:20 UTC. Degradation was particularly noticeable when using the BSS Dashboard and Add Payment functionality. The underlying database server experienced sustained high resource utilization, which caused delayed responses and partially affected service availability. Normal performance was restored after the affected database sessions were gracefully terminated.

Customer impact

  • BSS Dashboard pages responded slowly or failed to complete some requests.
  • Add Payment requests experienced slow responses or failures during the incident window.
  • The service was degraded for a duration of approximately 6 hours and 59 minutes. The number of affected users and unsuccessful payment attempts has not yet been confirmed.

What happened

Database queries supporting the BSS Dashboard and Add Payment functionality did not use the available database access paths efficiently. As request volume accumulated, these queries consumed increasing database resources and caused broader performance degradation. Once the team identified the affected sessions, terminating them relieved the immediate resource pressure and restored normal response times.

Root cause

Inefficient execution of database queries used by the BSS Dashboard and Add Payment functionality gradually led to sustained high CPU utilization and resource contention on the shared database server.

Contributing factors

  • Existing monitoring did not generate an early internal alert for the sustained database resource condition as the rate of requests was normal and degradation increased gradually.
  • The incident was not reported as critical; thus, the escalation workflow did not trigger an immediate after-hours response, extending the time before investigation began.
  • Additional safeguards, such as query timeouts or resource controls, were not in place to limit the impact of unusually expensive queries.

Resolution and recovery

The response team identified the database sessions consuming the most resources and began graceful termination of the affected sessions, while validating that terminated sessions did not cause additional failures. Database CPU utilization and query response times returned to normal levels, restoring the affected functionality by 06:20 UTC. Performance was monitored following recovery and remained stable based on the information available for this report.

Corrective and preventive actions

Workstream Commitment Status
Query and database optimization Review execution plans and optimize the affected queries, indexes, and statistics so requests use efficient access paths In progress / target date: Sep. 11, 2027
Proactive monitoring Add alerts for sustained CPU utilization, in conjunction with long-running queries to support earlier detection Planned / target date: Sep. 4, 2027
Resource safeguards Assess query timeouts and database resource controls to reduce the impact of runaway workloads Under review
Validation and readiness Performance-test the Dashboard and Add Payment paths under production-representative load and document a response runbook Under review

Timeline

Time Event
26 Aug, 23:21 UTC Performance degradation begins
27 Aug, 01:41 UTC Customer support ticket reports an urgent service issue
27 Aug, 05:15 UTC Technical investigation begins
27 Aug, 06:20 UTC Affected sessions are terminated and normal performance is restored
August 28, 2026 · 18:03 EEST
Resolved

Service performance has been restored for the affected platform functionality. Our investigation found that inefficient database query execution caused elevated resource utilization and degraded response times. We mitigated the issue by terminating the affected database sessions, and performance returned to normal levels at 06:20 UTC.

August 27, 2026 · 18:09 EEST
Investigating

Our team is continuing investigation into root cause

August 27, 2026 · 13:53 EEST
Investigating

Our team is continuing investigation into root cause

August 27, 2026 · 11:46 EEST
Investigating

Our team is continuing investigation into root cause

August 27, 2026 · 10:48 EEST
Update

Monitoring/Recovered: performance issues recovered; continuing investigation into root cause

August 27, 2026 · 09:41 EEST
Investigating

Our team identified degraded performance issues on BSS Portal started by 26 August 2026, 23:21 UTC

August 27, 2026 · 09:00 EEST

← Back