Skip to content
Back to Journal
Engineering3 min read28 March 2022

The three-hour server migration that cost us a client

The three-hour server migration that cost us a client

It was supposed to take forty minutes. We had done server migrations before and they were routine. New server, DNS propagation, a quick test, done. We had moved this particular client from shared hosting to a VPS because their traffic had grown and the old environment was struggling.

What we had not accounted for was that their email was running through cPanel on the same server, and migrating it properly would take longer than we expected because their mailbox was 14GB and we had not provisioned for the transfer time. Email went down at 11am on a Monday. It came back at 2:15pm.

Three hours and fifteen minutes of downtime. On a Monday morning.

What we did wrong

The technical problem was fixable and we fixed it. The real failure was in the three steps before that Monday.

First, we had not done a thorough audit of what was running on the server before planning the migration. We had checked the website. We had not checked what else lived in that cPanel account.

Second, we had scheduled the migration on a weekday morning without telling the client. Our logic was that we would be quick and they would not notice. This was wrong on multiple levels. The client should have known the migration was happening and when. The client should have been given a choice about timing. A Monday morning, when their team is processing orders and sending invoices, is not the right window for an infrastructure change.

Third, when things went wrong, we spent the first forty minutes trying to fix it before telling the client what was happening. By the time we messaged them, they had already noticed and sent us three WhatsApp messages and an email.

What this cost

The client did not renew. They had been with us for eight months. They were not unreasonable people. But three hours of email downtime on a business day, with no warning and slow communication while it was happening, broke the trust we had built.

We wrote a post-mortem, which is a thing we did not do before this incident. We now write post-mortems for every failure, technical or otherwise, that affects a client. The post-mortem for this one identified: pre-migration audits are mandatory, client sign-off with a preferred maintenance window is required, and any downtime gets a client notification within ten minutes regardless of whether we think we can resolve it quickly.

The technical quality of the migration was actually good. The server performed better. The site was faster. None of that mattered because the experience around the migration was careless. That is the lesson we carry.

Published 28 March 2022
Start a Project