Omegaswift
Cloud

Backups You Can Actually Restore From

Backups are easy to buy and easy to believe in. Recovery is the part that has to be proved, and the only proof is a restore.

The Omegaswift engineering teamCloud and infrastructure8 min read

Start at the Recovery and Work Back

The question is not how to take backups. It is how quickly each system has to be working again, and how much recent work you can afford to lose. Everything else follows from those two answers, and both of them belong to the business rather than to whoever runs the servers.

Ask the owner of each significant system two questions. If this were down, when does it stop being annoying and start being serious. And if we brought it back as it was a few hours ago, what would somebody have to redo by hand.

The answers vary more than people expect. A document store can often lose a day without much pain. An order system frequently cannot lose an hour, because rebuilding that hour is worse than the outage was. One rule for everything means paying too much for the systems that do not need it and too little for the ones that do.

The Gaps Are in the Newest Systems

Old fashioned servers are usually well covered, because the process was built when servers were all there was. The holes are elsewhere.

Subscription platforms split the responsibility. They keep the platform running. Your data inside it is often your problem, and what looks like retention is usually a recycle bin with a short window. Something deleted or corrupted and noticed months later is frequently gone for good.

Laptops are the second hole. Where people keep work on the machine, it is protected only if something is deliberately protecting it. The third hole is settings: the accounts, the network rules, the pipelines and the policies that make the whole thing what it is. Restoring data into an environment somebody then has to rebuild from memory turns a short recovery into a long one.

Isolation Is What Makes a Backup a Backup

A copy that can be deleted by whoever got into the main system is not a backup. It is a second copy of the same risk. Ransomware crews look for the backup system first, because destroying it turns a technical problem into a commercial negotiation.

In practice that means at least one copy that cannot be changed or deleted for the whole of its retention period, a login for the backup system that is separate from your normal one and protected by two step sign in, and administrator access that does not flow out of a compromised account.

Copies in more than one place and in more than one form is still good advice, and the reason has not changed. A copy in the same platform, under the same login, in the same region is exposed to everything the original is exposed to.

What a Real Restore Test Looks Like

Checking that a backup job finished proves that a job finished. It does not prove the data is usable, that the application starts, or that anybody knows the procedure.

A real test restores a whole system into a separate environment, starts the application, gets somebody who uses it every day to confirm the data looks right, and times the whole thing from the moment the decision was made. That elapsed time is your actual recovery capability. It is nearly always longer than the number written in the plan.

Run some tests without warning and with the usual owner unavailable. A recovery that works only when one particular engineer answers the phone is not a process. It is a dependency on somebody who may be on a plane.

Recovery Is an Order, Not a Pile

Restoring twenty systems at once is not possible, and would not help if it were. They lean on each other, and the order is not obvious at four in the morning.

Write the order down now, while nothing is broken. Logins and name lookups first, because almost nothing signs in without them. Then network paths and anything that has to phone home for a licence. Then the systems the business named as most urgent, then the ones that feed them, then the rest.

Treat bandwidth as part of the order. Pulling large volumes back down a line sized for a normal day takes far longer than the technology suggests. That is much easier to plan around than to discover on the day.

Documentation and Keys Under Pressure

The recovery notes have to be readable when the thing they describe is unavailable. A procedure stored on the file server you are restoring, or in a wiki behind the login system that is down, is not available at the one moment it is needed.

Keep a copy somewhere separate carrying the steps, the order, the contacts and the escalation route for every supplier involved. Include account numbers and support entitlements, because a support call that begins with hunting for a contract reference begins slowly.

Emergency administrator credentials need the same treatment. Stored properly, reachable without the systems being recovered, and checked now and then, so that you find out a password has expired on a quiet Tuesday rather than during an incident.

Written by

The Omegaswift engineering team

Cloud and infrastructure at Omegaswift. Filed under Cloud.

Ask us about this

Ready to talk about your IT?

We are happy to answer any questions you have and help you work out which of our services fit your needs.