
The Question Is Order, Not Speed
Nobody seriously argues that patching matters. The argument is always about the server that cannot be restarted during trading hours and the system whose vendor withdraws support if you touch it. Both of those have answers, and neither answer is to give up.
For almost everything else, the answer is order. Updates go to a small group first, then to a wider one, then to everybody, with enough time between each step for a problem to appear somewhere cheap. That structure is usually called rings, and it is what lets you patch quickly without gambling the business on a Tuesday.
The other half of the job is knowing what you have. A patching process that covers the machines your tool can see, while several servers and a rack of equipment sit quietly outside it, will report green every month and protect nothing at all.
What follows is how to build both, and what to do about the systems that genuinely cannot follow the process.
Rings, and Who Goes in the First One
The first ring is a handful of machines, including some belonging to the people who build and support your systems. They take the update first, and they know to report anything odd rather than quietly working around it.
The second ring is a wider pilot chosen to cover the real variety of your estate. Different hardware, different departments, the machine running the design software, the laptop with the ancient plug in that finance depends on. A pilot made entirely of IT machines proves only that updates work on IT machines.
The third ring is everybody else, and it should happen automatically once the earlier rings have been quiet for an agreed period. Write that period down for each category of update and hold to it, because the deadline is what stops a rollout stalling during a busy week and never starting again.
Do not put the finance team in an early ring during month end, or the warehouse during the pre holiday peak. Ring membership should follow the business calendar rather than the org chart, and it is worth revisiting once a year instead of being set once and forgotten.
Not Every Update Carries the Same Risk
Treat updates by category rather than as one long queue. Browser and operating system security updates are low risk and high value. They should move fast, with a short deferral window and a hard deadline behind it.
Firmware, drivers and anything touching storage or networking are the ones that produce a machine which will not start. They deserve slower ring progression and a tested way back. The same care applies to database engine upgrades and to anything a supplier describes as a major version.
Line of business applications need their own path, usually with the vendor involved and a functional test written by the people who use the software daily. Your IT team can confirm the application opens. Only the finance team can tell you whether the year end report still balances.
Then there is the exception category. An actively exploited weakness in something facing the internet skips the rings, goes out today, and gets handled as an incident rather than as maintenance. Decide in advance who is allowed to make that call, so nobody spends the afternoon hunting for approval while a scanner sweeps your address range.
Maintenance Windows People Believe In
A regular, boring, well announced window beats an occasional heroic one. The same evening each month, published a year in advance, with a standing agenda. When it is predictable, departments plan around it instead of objecting to it every time.
Announce what is changing, what might be affected and who to ring if something looks wrong the next morning. Then send a short note when it is finished, including anything that did not go ahead and why. People forgive an outage they were warned about far more readily than a surprise of the same length.
Decide before you start who signs off the go and no go, and what the back out step is. Where the back out is restore from backup, somebody should have verified that backup the same day rather than assuming it. Leave the following morning clear. The problems that appear are usually reported by users at nine o'clock rather than by monitoring at midnight, and having the person who did the work available is what keeps a small regression from becoming a week of confusion.
The Machine That Cannot Be Restarted
Every business has one. The server running the production line. The till system during trading hours. The application whose licence server sulks after a reboot. The database that everything else quietly depends on.
The best answer is architectural. Two of them, and you patch one at a time. Where that is affordable it removes the problem completely and tends to pay for itself in other ways. Where it is not affordable, write that down in the risk register in plain language rather than quietly failing to patch and hoping nobody asks.
Where a restart really is limited to a rare window, use what is available in between. Live patching for kernels exists on several platforms. Plenty of updates apply without a reboot. Compensating controls carry more weight here than anywhere else: tighter network restrictions around the machine, closer monitoring, no general user access, no browsing or mail from it.
And schedule the window you do get properly. Book it a year ahead with the business, treat it as immovable, and use the whole of it. The teams that struggle most are the ones who get one window a year and spend the first hour of it deciding what to do.
When the Vendor Forbids Updates
This is real and it is common. Medical equipment, industrial controllers, laboratory instruments, older accounting packages: plenty of vendors will withdraw support if you patch the underlying operating system yourself.
Get the restriction in writing, with a name and a date attached. Verbal advice from an engineer on a support call is not something you can show anybody later, and it has a habit of being denied at exactly the moment it matters.
Then build a fence instead of arguing. Put the machine on its own network segment. Block it from the internet entirely unless there is a specific documented need. Restrict what can reach it down to the named systems that must. Remove browsing and email from it completely. Watch it more closely than anything else you own, because you have accepted that it will stay vulnerable.
Then put two questions to the vendor in writing. Which updates do they test and approve, and on what schedule. And when does support for this version end. Put that end date into your roadmap with a budget beside it, because an unpatchable system with no replacement date is a decision to keep it forever.
Rollback, and the Test That Is Actually Useful
Take a snapshot or an image before anything significant, and confirm it exists rather than assuming the job ran. On a virtual machine that is a couple of minutes. On a physical one it is a verified backup plus known good boot media, and both take longer to arrange than people allow for.
Write the back out steps down before you start, in enough detail that somebody else could follow them. The person who applied the update is not necessarily the person available at seven the next morning when the warehouse cannot print labels.
The functional test list is the part most teams skip and the part that catches almost everything. Ask the people who use each system for the handful of actions that would tell them it is working. Sign in. Print a delivery note. Raise an invoice. Run the overnight job. Open the report the director reads on a Monday. That list is worth more than any technical health check.
Then keep a record of what was applied where and when. A fortnight later, when something breaks, the first useful question is what changed, and a patch log answers it in a minute instead of an afternoon.
Measure Coverage, Not Effort
The report worth reading shows how many machines are behind, by how long, and which ones they are. A count of updates applied only tells you that the process ran somewhere.
Watch for the silent failure. A machine whose agent has stopped reporting drops off the list of noncompliant devices and starts looking like a success. Reconcile the patching tool against your asset inventory and your directory on a schedule, and treat anything that has not reported for a while as out of date until somebody proves otherwise.
Keep an exceptions register with an owner, a reason, a compensating control and a review date for every system that cannot follow the standard process. Review it on a fixed date with the owners present. Exceptions without expiry dates stop being exceptions and become the permanent shape of your estate.
Then report it upwards in business terms. Which systems are behind, what somebody could do with the weakness, and what it would take to fix or replace them. Nobody funds a patching project on its own merits. People do fund keeping the thing that pays the bills running.
The Omegaswift engineering team
Security and operations at Omegaswift. Filed under Cyber Security.



