Blog

Mind the OT Recovery Gap: How Recovery Became the Weakest Link in OT Incident Response

When an industrial cyber incident makes the news, the story usually covers how the attackers got in and how responders pushed them out. Those details matter, and they shape how the industry defends itself. The phase that decides how much an incident costs gets far less attention: recovery. Every hour a line sits idle means lost production, missed shipments, and pressure on the suppliers, customers, and communities that depend on the operation. The time between detecting an incident and returning an industrial process to safe, verified operation is the OT recovery gap. Shrinking that gap is one of the most practical ways an organization can reduce the business impact of a cyber incident.

OT Incident Response Overview

SANS has outlined a framework for responding to cybersecurity incidents that covers every phase of the incident response lifecycle. It is known as PICERL, which stands for Preparation, Identification, Containment, Eradication, Recovery, and Lessons Learned. The framework can be applied to OT cybersecurity incident handling, giving organizations a clear guide to the steps they need to take to work through an incident effectively.

Each phase takes on added weight in an industrial setting:

  • Preparation covers the plans, roles, tools, and backups an organization puts in place before an incident. In OT, this includes defining which processes are most critical, who has authority to shut down or restart a line, and how operations, engineering, and security teams will coordinate.
  • Identification is detecting and confirming that an incident is underway. OT teams have to separate malicious activity from equipment faults, process upsets, and routine engineering changes. Without network visibility and a known baseline, that distinction is hard to make.
  • Containment limits the spread and impact of the incident. In industrial environments, containment decisions carry safety and production consequences, such as isolating network segments, severing IT and OT connections, or shutting down a process as a precaution.
  • Eradication removes the threat from affected systems. That can mean cleaning or rebuilding engineering workstations and servers, removing unauthorized accounts and remote access paths, and confirming that controllers are running only approved logic.
  • Recovery returns systems and processes to normal operation. This is often the longest phase in OT, because it involves physical equipment, safety validation, and staged restarts.
  • Lessons Learned captures what worked, what failed, and what needs to change. The findings feed back into Preparation and strengthen the response to the next incident.

OT environments set priorities differently from IT. Safety of the people, processes, and environment comes first, followed by availability and integrity. Many industrial assets run for decades, rely on vendor-specific software, and cannot be patched or rebooted on demand. These realities shape how every phase of PICERL is carried out on the plant floor.

In an industrial environment, recovery means restoring a physical process. That includes the logic running on PLCs and safety controllers, HMI and SCADA systems, historian data, engineering workstation images, device firmware, and network device configurations. Each one of those assets must be returned to a known-good state and validated before operators can safely restart production. The time and confidence required to complete that work depend almost entirely on how well the organization prepared before the incident began.

The OT Recovery Gap

Public accounts of most incidents focus on the identification, containment, and eradication steps taken by the organization that experienced the incident. Within the PICERL cycle, the time between Identification and the completion of Recovery in an industrial environment can span from days to weeks to months depending on the organization. That gap of time is where a large slice of the cost from an incident takes place, and where the downtime causes damaging effects to the secondary and tertiary organizations around it. 

The Jaguar Land Rover (JLR) incident illustrates how extensively the consequences of a cyber disruption to manufacturing can cascade. In late August 2025, a cyber attack began that forced JLR to halt vehicle production at plants in Solihull, Halewood, and Wolverhampton. Manufacturing remained offline for almost six weeks at sites that together build about 1,000 vehicles a day. A report from the Cyber Monitoring Centre (CMC) assessed the overall impact at £1.9 billion in losses with over 5,000 UK organizations impacted, including suppliers and dealerships. The UK Government provided JLR with a £1.5 billion loan guarantee to “be paid back over five years and bolster JLR’s cash reserves so it can support its supply chain which has been greatly impacted by the shutdown.” The CMC report found that lost manufacturing output at JLR and its suppliers accounted for the vast majority of those losses, a cost that grew with each week production stayed down.

The same pattern appears across industrial incidents of every size. Once production stops, costs build from several directions at once: lost output, idle labor, missed delivery commitments, contractual penalties, emergency response services, and damage to customer relationships. Many of these costs grow with every day a process stays offline, which makes time the most important variable in the total impact of an incident. Of all the factors that shape an incident, organizations have the most influence over how prepared they are to restore operations once the threat is contained. That makes the recovery gap one of the few parts of an incident an organization can plan to shorten. 

Why OT Recovery Takes So Long

Across the industry, practitioners see several common obstacles that stretch out OT recovery:

  • Informal and untested procedures. Many organizations have no written procedure for backing up and restoring OT assets. Where procedures do exist, they are rarely exercised, so responders discover gaps in the middle of an incident, when time and production are already being lost.
  • Legacy controllers without native backup. Many PLCs, controllers, and other control systems in service today were installed decades ago, before automated backup was a design consideration. Capturing their logic and configuration often requires vendor-specific software, a connected engineering workstation, and manual effort. As a result, backups happen infrequently or not at all.
  • Incomplete asset inventories. An organization cannot back up what it does not know it has. Gaps in OT asset inventories leave controllers, drives, HMIs, and network devices outside the backup scope, and those missing assets often surface only when a line refuses to restart.
  • Uptime pressure limits recovery testing. Industrial processes are built to run continuously, and planned outages are scarce and heavily scheduled. That leaves little opportunity to take systems offline and prove that a restore will work, so many recovery plans stay theoretical until an incident forces the test.
  • Air-gapped networks complicate storage and transfer. Segmented and air-gapped networks protect critical processes, and they also make it harder to move backup files to secure, offsite, or immutable storage. Manual transfers with removable media add effort and delay, and they introduce their own security risks.
  • Recovery lives in people’s heads. In many facilities, knowing how to restore a controller or rebuild an engineering workstation rests with a few experienced engineers or an outside integrator. When those people are unavailable or have left the organization, recovery slows dramatically. Documented runbooks and centrally stored project files keep that knowledge available to whoever is on shift.

Closing the Gap

The recovery phase is won or lost during Preparation. Organizations that recover quickly tend to have these practices in place:

  1. Prioritize by process. Identify the assets each critical process depends on and set a recovery order and target time for each.
  2. Automate versioned backups. Capture PLC, HMI, robot, and drive configurations on a schedule, and compare them against what is running on the device.
  3. Keep offline copies. Store backups where ransomware on the business network cannot reach them.
  4. Maintain known-good baselines. Change detection shows responders exactly what was altered and which version to restore.
  5. Stage the rebuild kit. Keep engineering workstation images, software installers, license keys, and firmware ready to deploy.
  6. Exercise and time restoration. Run tabletop and hands-on restore drills, measure how long each step takes, and feed the results into Lessons Learned.

Conclusion

Identification, containment, and eradication stop an attacker. Recovery restores production, revenue, and the trust of the customers and communities that depend on the operation. Organizations that treat OT backup and recovery as a core part of their incident response plan shrink their recovery gap, which limits the damage to their own operations and to the organizations that depend on them.