Applying the patch was never the hard part
A patch is a file. Getting it onto a fleet that is doing paid work, in an order that keeps failures affordable, inside windows somebody else controls, and then persuading the machines to actually run it — that is a logistics problem, and it is where remediation programmes go to stall. These essays are about the moving of the thing rather than the thing itself.
All essays
- Nobody owns this server
Remediation backlogs are not full of hard engineering. They are full of items waiting on a decision that nobody is positioned to make, and the fix is to change who has to act rather than to argue harder.
- The restart you keep postponing
A patch that is installed is not a patch that is running. Deferred restarts compound into reboot debt, and the debt is self-reinforcing: the longer you wait, the more frightening the restart becomes, so you wait longer.
- Downtime you are not allowed to schedule
In plants, hospitals and trading floors the patch takes minutes and the permission takes months. When the maintenance window is the binding constraint, the useful engineering is not making the patch faster — it is changing what has to happen inside the window.
- Patch the fleet in waves, and mean it
Staged rollout is only worth its overhead if the stage boundaries are real — if the first batch is small and representative, if the soak is watching for something specific, and if you can actually reverse.