BSS Application Performance Degradation Due to Backend Server Overload

Major incident BSS US region
2026-07-24 17:53 EEST · 42 minutes

Updates

Issue

Τhe BSS application experienced performance degradation caused by one of the backend application servers becoming overloaded. The increased load resulted in elevated response times and a period of approximately two minutes during which the affected server was unable to serve client requests.

The issue was detected by the operations team’s monitoring systems, which identified abnormal response times and reduced service availability. As an immediate mitigation, client traffic was redirected to a healthy backend server, restoring normal application performance.

This incident occurred while the Azure Load Balancer remained unavailable following the previously reported load balancer incident. Because the normal load balancing functionality was not available, the environment had reduced resilience, making it more susceptible to performance issues when a backend server experienced excessive load. Restoring the Azure Load Balancer remains a critical priority to reduce the likelihood of similar incidents in the future.

Business Impact

  • Increased response times for BSS application users.
  • Temporary service interruption of approximately two minutes for requests handled by the affected backend server.
  • Degraded user experience during the incident window.
  • No data integrity issues were identified.

Root Cause

The incident was caused by excessive load on one backend application server, resulting in degraded application performance and a brief service interruption.

Although the workload was successfully shifted to a healthy server, the environment was operating without the normal Azure Load Balancer functionality due to the previously reported load balancer routing incident. As a result, traffic distribution and automatic balancing capabilities were reduced, increasing the likelihood that individual backend servers could become overloaded.

The absence of a fully operational load balancing service contributed to reduced operational resilience during this incident.

Detection

The incident was proactively detected through the organization’s infrastructure and application monitoring systems, which generated alerts indicating:

  • Elevated application response times.
  • Reduced backend server responsiveness.
  • Temporary loss of service from the affected backend server.

Early detection enabled rapid investigation and mitigation before a prolonged service outage occurred.

Mitigation

The following actions were performed:

  • Investigated monitoring alerts and confirmed backend server overload.
  • Redirected client traffic to a healthy backend application server.
  • Continued monitoring to verify service recovery and system stability.
  • Maintained the temporary routing configuration established following the previous Azure Load Balancer incident.

Corrective Actions

Completed

  • Redirected client traffic to a healthy backend server.
  • Verified restoration of normal application response times.
  • Continued enhanced monitoring of backend server health and application performance.

Pending

  • Restore the Azure Load Balancer to normal operation once Microsoft Azure completes resolution of the previously reported load balancer incident.
  • Reinstate automatic traffic distribution across backend servers.
  • Review backend server capacity thresholds and alerting to improve early detection of resource saturation.
  • Validate failover procedures through regular operational testing.

Relationship to Previous Incident

This incident is directly related to the previously reported Azure Load Balancer Traffic Routing Failure. While the temporary routing configuration restored application availability, the environment continued operating without its intended load balancing capabilities.

Recovering the Azure Load Balancer is critical to restoring automatic traffic distribution, improving overall system resilience, and reducing the likelihood of backend server overload resulting in similar performance degradation incidents.

July 29, 2026 · 18:03 EEST

← Back