TL;DR

On August 3, between 18:30 and 23:00, we carried out maintenance on our infrastructure. The maintenance ran into a series of unexpected issues, which significantly extended the maintenance window beyond what we had planned. The following day, August 4, at approximately 16:00, our hypervisor which hosts all services went down entirely. We dispatched emergency remote hands to investigate on-site, and after extensive troubleshooting identified faulty RAM as the cause, which we replaced. All services were fully restored by 19:00 on August 4.

We recognize this was a long and disruptive incident, and we’re publishing this postmortem to explain what happened and what we’re changing as a result.

Timeline

All times CEST.

August 3

  • 18:30 — Maintenance begins on our infrastructure.
  • 18:30–23:00 — Maintenance encounters a series of unexpected complications, extending the work well beyond the originally planned window.
  • 23:00 — Maintenance work concludes; all infrastructure returns to normal operation.

August 4

  • ~16:00 — The hypervisor goes down.
  • ~16:00 — Emergency remote hands are dispatched to the site to begin investigation.
  • 16:00–19:00 — Extensive on-site debugging is carried out to isolate the cause of the failure.
  • 19:00 — Faulty RAM is identified and replaced. All affected services are confirmed recovered.

Root Cause

The outage on August 4 was ultimately traced to a faulty RAM module in the main hypervisor, which was identified and replaced by our remote hands team. We are continuing to review whether this hardware fault was connected to the extended maintenance work performed the previous day, or whether it was an unrelated, coincidental hardware failure. We’ll update this postmortem if that investigation surfaces a confirmed link.

Impact

All services hosted in Amsterdam were unavailable from approximately 16:00 to 19:00 on August 4 (roughly 3 hours), following an already extended maintenance window the night before. We understand this represents a significant disruption for affected clients, and we apologize for the inconvenience and impact this caused.

Resolution

  • Emergency remote hands were engaged to physically inspect the affected hardware.
  • Through methodical debugging, the faulty RAM module was isolated as the source of the failure.
  • The RAM was swapped, and all services were verified as fully recovered by 19:00.

What We’re Doing Differently

To reduce the impact of incidents like this going forward, we are introducing the following changes:

  1. Proactive client notifications — We will notify all affected clients by email as soon as we identify unplanned downtime, rather than clients having to discover it themselves.
  2. Public status page — We are publishing a status page so clients can check real-time system status and incident updates independently.
  3. Maintenance window review — We are reviewing our maintenance procedures to better anticipate and plan for complications, to avoid maintenance windows running significantly over their planned duration.
  4. Hardware health monitoring — We are evaluating improvements to our hardware monitoring so that failing components can be flagged and addressed before they cause an outage.

Closing

We know outages like this are frustrating, especially when they follow an already extended maintenance window. We’re committed to the changes above and to being more transparent with clients when things go wrong. Thank you for your patience while we worked through this.

If you have questions about this incident, please reach out to [email protected].