Blog

From Plan to Practice: Preparing for OT Incidents with Response Plans, Playbooks, Runbooks, and Tabletop Exercises

PICERL, the incident response framework popularized by SANS, breaks incident handling into six phases: Preparation, Identification, Containment, Eradication, Recovery, and Lessons Learned. Each phase matters, but in operational technology (OT) environments, Preparation is the most important step. Every action a responder takes during a live industrial incident depends on decisions and documentation that had to exist before the first alarm. A team cannot contain a threat without pre-approved authority to isolate a process network, cannot escalate to the appropriate people without defined criteria and current contacts, and cannot swiftly and systematically execute technical steps without having them documented and tested. This article covers the core components of OT incident response preparedness: the incident response plan, playbooks, runbooks, and the tabletop exercises that validate them. In OT, the response you execute is the response you prepared.

OT Incident Response Plans

The first major step any organization should take is building an OT-specific incident response plan. Organizations generally have two options when drafting the plan. The first is an all-hazards OT incident response plan that addresses cybersecurity incidents alongside events such as equipment failures, natural disasters, and safety incidents, which allows cyber response to integrate with existing emergency management and business continuity processes that operations staff already know. The second is an independent OT cybersecurity incident response plan, which provides more depth on cyber-specific activities such as evidence preservation and threat containment, and should clearly reference how it coordinates with the corporate IT incident response plan and site emergency procedures. 

The incident response plan is the overarching strategic document used during an incident. It can be structured around the PICERL framework to clearly define each stage of the incident response process. Key components include roles and responsibilities for all of the internal stakeholders involved in handling an incident, along with external parties such as vendors, integrators, and any incident response retainer firm. The plan should contain current contact information for all parties, including out-of-band methods in case corporate email or phone systems are unavailable. Clear escalation criteria and procedures should define when an event becomes an incident and when leadership, regulators, or law enforcement must be engaged. An incident severity matrix should rate incidents based on OT-relevant impacts such as loss of view, loss of control, or loss of safety. The plan should define an escalation path for each severity level.

The most important elements of an effective plan are keeping it current and testing it regularly. Updates should occur at least annually and after any significant incident, exercise, organizational change, or system modification. Testing through tabletop and functional exercises verifies that the plan works for the organization and that staff understand their roles. The incident response plan should serve as the foundation that all other incident documentation, including playbooks and runbooks, feeds into.

Playbooks

Once the incident response plan has been created, the next step is to build out playbooks. Playbooks go one step further than the plan, serving as scenario-specific operational documentation that translates the plan’s policies, roles, and escalation paths into a defined course of action for a particular type of incident. Organizations should identify the top three to five scenarios they are most likely to encounter, such as ransomware, human-machine interface (HMI) manipulation, and vendor or supply chain compromise, and develop a response for each one. Each playbook should follow the same PICERL structure as the incident response plan so that responders move between documents without confusion.

A well-built OT playbook defines the indicators or triggers that activate it, the roles responsible at each stage, and the key decision points that are unique to industrial environments. Examples include when to sever connectivity between IT and OT, when to transition the process to manual or degraded operations, and who holds the authority to make those calls. Playbooks should also capture notification requirements for leadership, regulators, insurers, and government agencies, along with references to the specific runbooks responders will execute for technical tasks.

Runbooks

The next set of documents to develop is the runbooks, the most in-depth component of the incident response documentation collection. Runbooks are task-specific tactical and technical documents that spell out the exact commands, scripts, tool settings, and fixes responders need. Where the incident response plan and playbooks define who is involved, what decisions must be made, and when actions occur, runbooks provide the step-by-step “how” for a specific asset or situation. In an OT environment, typical runbooks cover tasks such as isolating a compromised network segment at a firewall or conduit boundary, collecting volatile and log data from an engineering workstation, verifying programmable logic controller (PLC) logic and firmware against known-good baselines, restoring an HMI or historian server from backup, and reloading controller projects using vendor tools. 

Each runbook should list its prerequisites, including required access, credentials, licenses, software versions, and the location of verified backups and integrity hashes. Steps should be written so that a qualified responder under pressure can follow them without interpretation, and they should include validation checks that confirm each step succeeded, rollback instructions if a step fails, and safety hold points where operations staff must confirm the process is in a safe state before work continues. Runbooks should be version controlled, reviewed whenever systems or firmware change, and maintained in both digital and printed form in a location that remains accessible if corporate systems are unavailable.

Tabletop Exercises and Testing

Developing an incident response plan, playbooks, and runbooks is only the beginning. Documentation that has never been exercised often contains outdated contacts, unclear decision authority, missing steps, and assumptions that fall apart under pressure. Regular testing validates that each document works as written, builds familiarity among the people who will rely on it, and exposes gaps before an adversary or real-world failure does. 

Tabletop exercises are well suited to testing the incident response plan and playbooks. A facilitator walks participants through a realistic scenario, introducing injects that force decisions at key points, such as whether to disconnect the OT network from IT, when to move to manual operations, and who notifies regulators. Participation should extend beyond the security team to include operations, OT engineering, safety, plant leadership, legal, communications, and key vendors or integrators, since an OT incident will require all of these groups to coordinate. Scenarios drawn from peer incidents and the organization’s playbook library keep exercises relevant. Runbooks require a more hands-on approach and should be validated through drills and functional exercises in a lab, test environment, or planned maintenance window, including timed restorations from backup. 

Every exercise should conclude with a hotwash, a formal after-action report, and an improvement plan that documents findings, assigns owners, and sets deadlines for corrective actions. Updates should then be made to the affected plan, playbooks, and runbooks. A sustainable cadence might include testing the incident response plan at least annually, rotating through playbook scenarios over the course of the year, and testing critical runbooks and backup restorations on a scheduled basis and after significant system changes. Over time, this cycle of testing and improvement turns the documentation into a working capability that responders trust.

Conclusion

Preparation is the phase of PICERL that an organization fully controls, and the investment made there determines how every subsequent phase unfolds. An OT-specific incident response plan establishes authority, communication paths, and severity criteria. Playbooks translate that structure into defined actions for the scenarios most likely to affect the organization. Runbooks give responders the precise technical steps to isolate, verify, and restore critical assets. Regular tabletop exercises, drills, and restoration tests confirm that all three layers work together and that the people executing them know their roles. Organizations that build and maintain this documentation collection can make faster decisions during an incident, reduce downtime, and return the process to a safe and reliable state with confidence. When the first alarm sounds, the quality of the response will reflect the work completed long before it.