What happened
On July 30, our monitoring detected elevated error rates on an optical link within our core infrastructure. To prevent further degradation, our team replaced the faulty component the same afternoon. The replacement itself lasted under a minute.
That brief interruption had an unexpected side effect on the routing between several of our aggregation nodes. Some network sessions came back up in an inconsistent state that was not visible through standard monitoring: they appeared healthy while no longer carrying traffic.
Why the incident lasted longer than initially reported
We resolved this incident on July 30 at 18:00 CEST, having restored connectivity for the customers who had reported issues. That assessment turned out to be premature.
Because the affected sessions still appeared healthy, we had no reliable way to identify the remaining ones. They surfaced individually over the following hours as customers reported them, causing further interruptions during the night of July 30 and the morning of July 31 for a limited number of services. We should have kept this incident open until we had verified the state of the entire affected scope, and we apologize for the confusion this caused.
Resolution
An emergency maintenance window on the night of July 31 to August 1 addressed both the symptoms and their root cause. All core equipment was updated to the latest long-term software release, and configuration inconsistencies identified during our investigation were corrected. Services have been stable since.
What we are changing
Apology
We sincerely apologize for the disruption, and particularly for the repeated interruptions experienced between July 30 and August 1. We understand the impact this had on our customers and their own end users.