Status impact color legend
  • Black impact: None
  • Yellow impact: Minor
  • Red impact: Major
  • Blue impact: Maintenance

[Instances][POST MORTEM] - FR-PAR-1 - Issue with some PRO2 instances

Incident Report for Scaleway

Postmortem

July 28, 2026 — 09:02 CEST: Our monitoring system detected connectivity issues affecting multiple PRO2 hypervisors in FR-PAR-1. Several servers appeared unreachable.

July 28, 2026 — 09:14 CEST: The Scaleway Instances team began investigating the issue.

July 28, 2026 — 09:14 CEST: Initial findings showed that the affected servers were still operational, but network traffic was being disrupted.

July 28, 2026 — 09:16 CEST: An internal incident was opened, and the Network team joined the investigation.

July 28, 2026 — 09:32 CEST: One of the two leaf switches showed uplink issues. However, part of the traffic was still being forwarded through the affected switch.

July 28, 2026 — 09:35 CEST: We disabled the hypervisor ports connected to the faulty leaf switch to force all traffic through the redundant switch.

July 28, 2026 — 09:35 CEST: Connectivity was restored for the affected instances.

July 28, 2026 — 09:37 CEST: Logs from the faulty leaf switch showed that several ports went down, followed by one of its power supply units and then its uplinks. This sequence of events caused traffic black hole.

July 28, 2026 — 09:38 CEST: We also identified multiple power supply alerts affecting several hypervisors in the same rack.

July 28, 2026 — 09:44 CEST: A ticket was opened with our data center provider to request an electrical inspection of the rack.

July 28, 2026 — 11:19 CEST: While reviewing the equipment logs, we identified a protective shutdown on one switch following a temperature alert. However, no abnormal temperature readings were detected by the other sensors in the rack.

July 28, 2026 — 11:21 CEST: The data center provider performed a visual inspection and found no visible issues with the rack.

July 28, 2026 — 15:09 CEST: The faulty component responsible for the incident was identified. A defective power supply unit caused electrical disruptions across the rack, triggering temperature alerts, power supply alerts on other devices, and the loss of several links on one leaf switch.

July 29, 2026 — 14:21 CEST: All faulty components were replaced. The network connections that had been disabled were re-enabled. The service was fully restored and remained stable.

Posted Jul 30, 2026 - 14:00 CEST

Resolved

July 28, 2026 — 09:02 CEST: Our monitoring system detected connectivity issues affecting multiple PRO2 hypervisors in FR-PAR-1. Several servers appeared unreachable.
Posted Jul 28, 2026 - 09:00 CEST