Incident Response for OT: Why Your IT Playbook Won't Work
Most organizations already have an incident response plan, and most of those plans were written for laptops, servers and email. Apply that same playbook to a control system and you can turn a cyber incident into a safety incident or an unplanned shutdown. OT incident response needs its own plan, its own people and its own decision rules, agreed long before anything goes wrong.
Why the IT playbook breaks down
A typical IT response follows a familiar rhythm: detect, contain, eradicate, recover. Containment usually means pulling a machine off the network, and recovery often means wiping and reimaging it. On the plant floor, each of those steps carries consequences IT teams rarely have to weigh.
Isolation can stop the process. Disconnecting an HMI, a historian or a programming workstation may cut operators off from the information they need to run safely. Pulling the wrong cable can trip a unit.
Wiping destroys evidence and configuration. Reimaging an HMI may erase the only copy of a project file, a licence key or a vendor-specific driver.
Rebooting is not harmless. Some controllers and legacy systems do not come back cleanly, and some only restart with vendor help.
Uptime and safety are the priorities. In IT, confidentiality is often the top concern. In OT, the questions are whether people are safe, whether the process is under control and whether the environment is protected.
OT incidents still need a firm response, but every action has to be weighed against its operational impact by people who understand the process.
Safety first, and operations decides

The single most important principle in OT incident response is that safety comes before everything else, including evidence collection and speed of recovery. The second is that operations, not IT or security, has the final say on actions that affect the running process.
That requires clear decision authority, written down in advance:
Who can authorize isolating a network segment or a control system?
Who can order a controlled shutdown of a unit or a whole site, and on what criteria?
Who can decide to keep running in a degraded state, and for how long?
Who speaks to regulators, customers and the media?
Answered during an incident, these questions get answered slowly and under pressure. A good plan names roles rather than individuals, with alternates, so it still works at 2 a.m. on a holiday weekend.
At a minimum, the plan should define these roles:
An incident lead who coordinates the overall response.
An operations lead with authority over process decisions.
A controls or automation lead who understands the systems involved.
An IT and security lead for corporate systems and the shared boundary.
A communications lead for internal and external messaging.
An executive sponsor who can approve major business decisions.
Preserving evidence without making things worse
Forensics in OT is different too. Much of the most important evidence lives in places IT tools don't reach.
Controller logic and configuration. Capture the running program from affected PLCs, RTUs and safety controllers and compare it with a known-good copy. Unexpected logic changes are among the most serious findings in any OT investigation.
Firmware versions. Record what is actually running, not what the documentation says.
Programming workstation and HMI images. Where possible, take a disk image before rebuilding anything.
Network captures. A passive capture from a SPAN or mirror port can show what talked to what, without touching the devices themselves.
Logs and alarms. Historian data, alarm journals and operator logs help build the timeline.
Preservation should be done carefully, ideally by people who know both the equipment and the evidence requirements. When in doubt, preserve first and rebuild second.
Plan for manual operations
Every OT response plan should answer one blunt question: if we lose our control system, or decide we can't trust it, can we keep running safely, and for how long?
For some sites the answer is a controlled shutdown. For others, such as a water utility or a pipeline, stopping isn't always the safest option, and the plan has to cover running by hand. That means:
Documented manual procedures, kept on paper and accessible without a computer.
Operators who have actually practised them, not just read them.
Enough staff, radios, keys and field instruments to run that way for an extended period.
Clear criteria for when to switch back to automated control.
Know who to call before you need them
Incident response in OT almost always involves people outside your organization. Build the call list now.
Your control system vendors and integrators. They may be the only ones who can safely recover certain equipment. Know their emergency support process and what your contract actually covers.
National cybersecurity agencies. In the US, CISA provides incident response support and guidance for critical infrastructure. In Canada, the Canadian Centre for Cyber Security plays a similar role. Both accept incident reports and can share threat information.
Law enforcement, where criminal activity such as ransomware is involved.
Regulators and sector bodies, where reporting obligations apply. Some sectors have short, mandatory reporting windows.
Insurers and legal counsel, who may need to be involved early, particularly before any engagement with an attacker.
External incident response support with genuine OT experience.
Practise with operations in the room
A plan that has never been exercised is a theory. Tabletop exercises are the cheapest way to find the gaps, and in OT they only work when operations, controls staff, IT, security and leadership take part together.
Useful scenarios include ransomware that spreads from the corporate network and threatens HMIs, a suspicious change discovered in PLC logic, a compromised remote access account, and the loss of a historian or the control room network. Ask the hard questions: who makes the shutdown call, how operators would know the HMI is lying to them, and what happens when the vendor can't be reached.
Keep early exercises short, capture every gap, assign an owner and run them again.
How QBits Networks can help
Our OT incident response readiness service reviews your existing plan, maps decision authority and roles, and runs tabletop exercises with your operations and controls teams so that gaps surface in a conference room rather than during a real event. We can also help build the supporting procedures your people will rely on. Tell us what you need through our contact page and we'll scope it with you.