There is an issue which does not disrupt the main hosting service
-
RFO & timeline
Yesterday, there were network issues in BIT-2C at 20:28 Amsterdam time. For most customers, the issue was resolved at 20:45. For specific customers, the issue sustained until 21:39.
We hereby publish the RFO (Reason for Outage) with the exact events & timeline and points of improvement.
08:55: Rack at the BIT-2C facility largely lost power
On September 8, 2026, at 08:55, a rack at the BIT-2C facility largely lost power. The ATS in the rack had failed. While no Cyberfusion equipment is located in this rack, and therefore there was no impact at this time, it does house a spine switch that connects our equipment at BIT-2C.
20:28: BIT-2C goes completely offline
While our supplier put a new PDU into service at 19:00—an action that should not have affected Cyberfusion equipment, as it is not located in this rack—human error resulted in a spine switch—which is located in this rack—being unplugged from the PDU, causing it to lose power. This caused the BIT-2C location to go completely offline.
- Core nodes running in BIT-2C offline.
- Core clusters (with multiple nodes) route outbound traffic through a router in BIT-2C. During the incident, a short (expected) hiccup occurs during the automatic router failover to BIT-2A.
20:45: BIT-2C back up
The spine switch in BIT-2C is back online, restoring network connectivity.
- Core nodes running in BIT-2C back up.
21:15: Remaining network issues for Core clusters (with multiple nodes)
Core clusters (with multiple nodes) route outbound traffic through a router in BIT-2C. During the incident, an automatic failover to a backup router in BIT-2A occurred.
Even though BIT-2C is back up, outbound traffic still does not work.
21:39: Outbound network connectivity for Core clusters (with multiple nodes) restored
Routers typically run in pairs: master and backup. Nodes are configured with a 'gateway', which is an IP address pointing to the router that outbound traffic should be routed through. That gateway IP address is held by whichever router of the pair is currently the 'master'. If that router fails, another router automatically takes over the gateway IP address. This procedure worked as intended: during the incident, an automatic failover to a backup router in BIT-2A occurred.
Each router keeps its own MAC address - the pair does not share one. A MAC address is the hardware address a server actually sends its packets to. So when the master changes, the gateway IP address suddenly belongs to a different MAC address.
Every server keeps a small table (the neighbour cache) that says: 'gateway IP address X is at MAC address Y'. To keep that table correct, the new master announces the change: Keepalived sends unsolicited neighbour advertisements for IPv6, and gratuitous ARP for IPv4. Servers receive those and update their table.
During a period of flapping, those announcements can be lost. The servers then keep the old entry, and keep sending all their outbound traffic to the router that is no longer the master. Traffic doesn't come back if other parts of the network do send traffic to the actual master (asymmetric routing).
Once our engineers found the cause, manually triggering another VRRP failover—thereby triggering new advertisements and 'refreshing' the neighbour caches on all nodes—restored outbound traffic.
Points of improvement
Subject Explanation Status Replacement of ATS and PDUs During the emergency maintenance, our supplier removed the faulty ATS—which caused the initial power outage. In addition, new PDUs were installed in the rack. Finished Spine switch redundancy An ongoing project aims to implement redundant spine switches at each location. This eliminates the single point of failure associated with having only one spine switch per location; if one spine switch fails, such as due to a power issue, the location remains operational via the second spine switch. In progress Document announcement loss during flapping It took some time to restore the outbound connectivity of the Core clusters due to a specific issue involving MAC addresses. This issue and the associated procedure have been documented to ensure a faster resolution in the future. Finished -
Issues resolved
There were issues with outgoing IPv6 traffic and DNS resolving, those should be resolved now.
-
Still issues
We're still seeing issues, we're investigating.
-
Impact
All services should be back up. We saw a major hiccup at 20:27, followed by some minor ones.
-
Cause
Our network provider is having issues with a network switch.
-
Network issues
We just saw a network hiccup, we're investigating.
Experiencing an issue?
Question
Response within 4 working hours (working days from 9:00 AM to 5:30 PM Amsterdam time)
Free
Ask your questionUrgent problem
Response within 1 hour (24/7)*
*Charges may apply for this service.
Call +31 (0)40 711 44 96