PLC and HMI Backups: The Forgotten Part of Resilience
When a control system goes down, whether from ransomware, a failed hard drive or a well-meaning change gone wrong, the speed of recovery depends almost entirely on one thing: whether you have a good, recent, tested copy of everything you need to rebuild. In many plants that copy is incomplete, years out of date or sitting on the same network as the systems it is meant to protect. Backups are not glamorous, but they are one of the most reliable ways to turn a crisis into an inconvenience.
Why OT backups get forgotten
In IT, backups are usually automated, centrally managed and checked by someone whose job it is. In OT, the picture is often different. Controller programs live on a laptop that belongs to whoever last made a change. HMI projects were left with the integrator who commissioned the system. Network switch configurations were never saved after the original install. Nobody is quite sure which firmware version is running on the safety controller.
None of this is negligence. It is the natural result of systems that run for decades, change hands between vendors and staff, and rarely fail. The problem only becomes visible when you need to recover.
What to back up
A useful OT backup set covers far more than PLC programs. At a minimum, consider:
PLC, RTU and controller programs, including the comments, tag databases and documentation that make the logic readable.
Firmware versions for controllers, I/O modules, drives and communication cards, along with copies of the firmware files themselves where vendors allow.
HMI and SCADA projects, including graphics, alarm configurations, scripts and the runtime settings needed to deploy them.
Historian configuration, such as tag lists, collection settings and archive structure. Historical data may also need protecting where it supports compliance or investigations.
Network device configurations for switches, routers, firewalls and wireless or radio equipment.
Safety instrumented system (SIS) configurations, handled under your functional safety management process and with the involvement of the people responsible for it.
Software installers and licences, including licence files, activation keys and dongle details. A perfect backup is useless if you cannot reinstall the software to open it.
Operating system images for HMIs, servers and programming workstations, so a failed machine can be rebuilt quickly.
Documentation: network drawings, I/O lists, IP address plans, vendor contacts and the recovery procedures themselves.
Keep copies offline and out of reach
A backup that ransomware can reach is not a backup. Attackers routinely look for and destroy backup copies before deploying ransomware, and OT backups stored on a shared drive in the corporate domain are an easy target.
Good practice includes:
Keeping at least one copy offline or on immutable storage that cannot be altered or deleted from the network.
Storing copies in more than one physical location, so a fire or flood at one site does not take out both the system and its backup.
Restricting who can access and change backup repositories, using accounts separate from everyday corporate credentials.
Recording what each backup contains, when it was taken and what version of the system it reflects.
Know what is actually running

Backups only help if they match reality. Logic gets changed during troubleshooting at 3 a.m., and that change may never make it back to the master copy.
Change detection closes that gap. The idea is simple: keep an approved "golden" version of each controller program and configuration, then regularly compare it with what is actually running. Some vendors and third-party tools can automate this comparison and flag differences. Where automation is not possible, a scheduled manual comparison is still worth doing.
Change detection serves two purposes. Operationally, it keeps your backups current. From a security standpoint, an unexpected change in controller logic is one of the most important signals that something is wrong, whether it is an unapproved modification or something worse.
Test your restores
A backup you have never restored is a hope, not a plan. Testing is where most programs fall short, and it is also where the most valuable lessons emerge.
Testing does not have to mean restoring onto a live system. Options include:
Restoring HMI and server images to spare hardware or a virtual environment.
Opening controller backups in the programming software to confirm they are complete and readable.
Walking through a recovery procedure step by step with controls staff during a planned outage.
Running a tabletop exercise focused on rebuilding a specific system from scratch.
Each test tends to reveal something: a missing licence, an unreadable file format, a dependency on a vendor who no longer supports the product, or a procedure that only one person understands.
Recovery time and spare hardware
Even with good backups, OT recovery takes time. Rebuilding an HMI server, reinstalling software, reapplying licences, restoring projects and verifying communication with field devices can take hours or days per system. Controllers may need to be reloaded and then checked before the process restarts. Some steps may require vendor support that is not available on short notice.
Leadership should understand these timelines before an incident, not during one. Agreeing on realistic recovery time objectives for each critical system helps set priorities and shows where investment in faster recovery is justified.
Many control systems rely on hardware that is no longer manufactured. If a controller, a communication module or an industrial PC fails, a replacement may take weeks to source, or may only be available second-hand with uncertain provenance.
Review which components are critical, which are hard to replace, and whether keeping verified spares on site is worthwhile. Spares should be stored properly, recorded in your inventory and, where practical, loaded with known-good firmware so they are ready to use.
How QBits Networks can help
Backup and recovery readiness is a natural part of our OT incident response readiness work and our broader OT security assessments. We can help you identify what needs protecting, review how backups are stored and tested, and build recovery procedures your team can follow under pressure. Tell us what you need through our contact page and we'll scope it with you.