Back to insights
Playbooks & Best PracticesThreat Intelligence & Incident ResponsePoultryJun 1st, 2026 · 5 min read

Incident Response Testing

The Plan Is Not the Response

Steve Mustard
President & CEO

Share this insight

Every organization says they have a plan. It may sit in a binder on a shelf, a folder on SharePoint, a policy management system, or a carefully formatted PDF attached to the annual audit evidence. It may have been reviewed, approved, and assigned an owner. It may even contain the right phases, the right escalation paths, the right contact lists, and the right language.

But none of that proves the organization can respond.

In operational technology environments, especially in food manufacturing, incident response is not a paperwork exercise. It is a practical capability. It depends on people making decisions under pressure, understanding the plant, communicating across functions, protecting safety, preserving evidence, restoring operations, and learning from what happened.

That capability cannot be assumed. It must be tested. An untested incident response plan is not much better than no plan at all. In some cases, it may be worse because it creates the illusion of preparedness. The organization believes it has a response capability because the document is in place. The real test only comes when something goes wrong. And the worst time to discover that the plan does not work is during a real incident.

The Difference Between Having a Plan and Being Ready

Incident response planning often starts in the right place. Organizations identify roles and responsibilities. They define escalation paths. They document communication channels. They describe how incidents should be assessed, contained, recovered, and reviewed.

The real question is whether people can use it when conditions are uncertain, information is incomplete, and the operational consequences are changing by the minute. In an OT environment, a cybersecurity incident is rarely just a cybersecurity problem. It can affect production, product quality, worker safety, environmental compliance, equipment integrity, customer commitments, and public trust.

That means the response cannot be cyber-led in isolation. It must involve operations, engineering, maintenance, automation, safety, quality, legal, communications, vendors, and leadership. Each group may see the incident differently. Each group may have different priorities. Each group may use a different language. Testing helps expose those differences before they matter.

A plan may say “isolate the affected system.” But what does that mean on a production line? Who has the authority to do it? Which cable needs to be pulled? Which switch port should be disabled? What happens to the packaging line, the refrigeration system, the clean-in-place sequence, or the batching process if that system is disconnected?

A plan may say “restore from backup.” But where are the backups? Are they current? Are they usable? Has anyone restored them onto real hardware? Are the software licenses available? Are the configuration files complete? Does the vendor need to be involved? Is the restore process safe to perform while the plant is operating?

A plan may say “notify management.” But who needs to know what, and when? What should be said to operators? What should be said to customers? What should be said to regulators? Who decides whether production continues?

These are not details to be discovered during the incident. They are exactly why exercises matter.

Exercising the Whole Response, Not Just the First Hour

Many incident response discussions focus heavily on the initial response: detect the incident, assess the situation, contain the spread, and coordinate the first decisions. That is important, but it is not enough. Testing should cover the full incident response lifecycle, including response, recovery, restoration, remediation, and reassessment.

  • Response is the immediate action. It includes identifying what is happening, establishing command and communication, protecting people and operations, and containing the event.
  • Recovery is the process of bringing essential capability back. In OT, this may involve temporary procedures, manual operations, partial production, spare equipment, vendor support, or alternate control paths.
  • Restoration is the return to normal or near-normal operations. It requires confidence that systems are stable, configurations are correct, and the process can run safely.
  • Remediation addresses the underlying causes. This may include patching, access control changes, network segmentation, vendor access changes, backup improvements, monitoring improvements, or procedure updates.
  • Reassessment is the learning phase. It asks what the incident revealed about risk, architecture, dependencies, training, governance, and assumptions.

This last phase is often overlooked. But it is where organizations improve. The exercise is not only a test of the plan. It is a test of the organization’s understanding of itself.

Tabletop Exercises: Useful, But Limited

The tabletop exercise is the most common form of incident response testing. It is popular for good reasons. It is relatively easy to organize, inexpensive, and useful for bringing people together. A facilitator presents a scenario, introduces new information through scripted injects, and asks the team how it would respond. For many organizations, this is the right place to start.

A tabletop can reveal unclear roles, outdated contact lists, missing decision criteria, weak escalation paths, and conflicting assumptions. It can help executives understand that an OT incident is not simply an IT outage. It can help operations teams understand how cybersecurity decisions affect plant continuity. It can help cybersecurity teams understand that containment actions have physical and production consequences.

But tabletops also have limits. They are often too neat. The scenario is pre-scripted. The injects are controlled. The pace is artificial. Participants know they are in a meeting, not an incident. No one must find a laptop charger, locate a network cabinet, call a vendor at 2 a.m., identify the correct switch, restore a controller configuration, or explain to production why a line may need to stop. A tabletop can test discussion. It cannot fully test actions.

A Tabletop Exercise is a good place to start, but it has it's limitations.

Real-World Exercises: The Value of Friction

Real-world exercises are more demanding because they introduce practical friction. In a real-world drill, the team may have to access the equipment room, locate the affected asset, follow the actual procedure, use the real communication channels, test backup restoration, validate documentation, or demonstrate how a system would be isolated. Even when carefully limited to avoid impact on live operations, this type of exercise reveals things that a conference room cannot:

  • The cable is not labeled.
  • The backup server cannot be reached.
  • The vendor contact has changed.
  • The technician with the required knowledge is on vacation.
  • The drawing does not match the installation.
  • The remote access path is not documented.
  • The spare workstation is missing a required software license.

These are the kinds of discoveries that make real-world exercises uncomfortable and valuable.

The challenge is cost and complexity. Real-world exercises require planning, coordination, safeguards, and careful boundaries. They may not be practical every quarter. They may need to be limited in scope. But organizations that never test the practical elements of their response are leaving a major part of their capability unproven.

Virtual Exercises: A New Middle Ground

A third option is emerging: virtual incident response testing.

Virtual exercises can combine scenario generation, dynamic injects, immersive environments, and repeatable performance measurement. Generative AI can help create configurable scenarios based on industry, facility type, technology, team structure, known controls, and common incident patterns. Instead of following a fixed script, the scenario can evolve based on the team’s decisions. This matters because real incidents do not follow scripts.

A team may choose to isolate a system, only to discover that the action affects another process dependency. It may delay communication, causing confusion in operations. It may restore from backup before understanding whether the backup is clean. It may focus on malware while missing a remote access weakness. Dynamic scenarios can expose these decision consequences in a way static exercises often cannot.

VR headsets allow for an immersive response test.

Virtual environments can also introduce practical realism missing from tabletops. Participants may navigate a simulated plant, control room, network cabinet, engineering workstation, or vendor access portal. In more advanced formats, this may involve a 3D immersive environment viewed through VR glasses, or an immersive room where teams can stand together inside a simulated facility. Instead of merely discussing what they would do, participants can look around, locate assets, follow visual cues, interact with simulated equipment, and experience some of the spatial and situational pressure of an incident.


Immersive rooms provide an alternative way to conduct realistic response tests

The attraction is not that virtual exercises replace real-world drills. They do not. The attraction is that they can be used more frequently, with lower cost and lower operational risk.

They also allow repeatability. Teams can run similar scenarios over time, compare performance, and see whether training and process improvements are working. That turns incident response testing from an annual compliance event into a learning system.

Realism Matters Because Pressure Changes Behavior

Incident response exercises should be as realistic as practical. That does not mean reckless. No one should create unnecessary risk to live operations in the name of training. But the exercise should reflect the real conditions under which decisions will be made.

Real incidents are fast-moving. They involve incomplete information. They create tension between containment and continuity. They require judgment. They expose weak handoffs. They punish unclear authority. They reveal whether people know the difference between what the plan says and what the plant can tolerate.

This is especially important in OT environments because the correct cybersecurity action may not be the correct operational action if taken without context. Disconnecting a system may stop an attack path, but it may also disrupt monitoring, alarming, sequencing, safety visibility, or production quality. Restoring a system quickly may reduce downtime, but it may also reintroduce compromise if the root cause is not understood. Continuing to operate may protect production, but it may increase risk if process visibility or control integrity is degraded.

Exercises allow teams to practice these tradeoffs before they are real.

Frequency Should Match Consequence

There is no single correct frequency for incident response testing. The right cadence depends on the organization, the risk profile, the complexity of the operation, the maturity of the response capability, and the consequences of failure.

Many critical infrastructure organizations conduct drills one to four times per year. That is a useful benchmark, but the deeper question is whether the frequency is enough to maintain readiness.

A facility with high-consequence operations, complex OT dependencies, frequent changes, multiple vendors, remote access, aging systems, or limited internal expertise may need more frequent exercises. A stable environment with mature procedures may need fewer full-scale drills, but it still needs periodic testing, refreshers, and targeted exercises when systems or personnel change.

The trigger should not only be the calendar. Exercises should also follow major changes: new remote access architecture, new control system upgrades, new vendors, new production lines, new cloud connectivity, new backup platforms, new monitoring tools, new corporate incident response procedures, or lessons from incidents elsewhere in the sector.

Every major change alters the response environment. The plan should be tested against the world as it now exists, not the world as it was when the document was written.

What Good Testing Reveals

A good incident response exercise does more than confirm that people attended a meeting. It reveals whether procedures are usable. It shows whether documentation has enough detail. It tests whether contact lists are current. It identifies whether roles are understood. It exposes whether decisions can be made quickly. It clarifies whether cybersecurity, operations, engineering, safety, and leadership can work as one team.

It also reveals hidden dependencies. The organization may discover that only one person knows how to restore a controller. It may discover that a vendor must be contacted through procurement rather than directly. It may discover that engineering workstations are not backed up. It may discover that plant diagrams are outdated. It may discover that the incident response bridge line excludes the people who know the process best.

These discoveries are not failures of the exercise. They are the point of the exercise. The purpose is not to prove that the plan is perfect. The purpose is to find out where reality is different from the plan while there is still time to fix it.

From Compliance Exercise to Operational Capability

Incident response testing is sometimes treated as evidence for auditors, insurers, regulators, or customers. That evidence may be necessary, but it is not the real value.

The real value is operational confidence:

  • Can the organization recognize when something is wrong?
  • Can it make decisions under pressure?
  • Can it protect people, product, equipment, and the environment?
  • Can it restore operations safely?
  • Can it learn and improve?
  • Can it do all of that when the incident involves both cyber systems and physical processes?

In OT environments, resilience is not created by documents alone. It is created by practiced coordination, tested procedures, trusted communication, realistic assumptions, and repeated learning. The plan matters, but the plan is only the beginning.

Preparedness is not what the organization says it will do. Preparedness is what the organization has practiced doing.

About the leader

Steve Mustard
Steve Mustard
President & CEO

Steve Mustard is an industrial automation consultant with more than 35 years of engineering experience across multiple sectors. He is a licensed Professional Engineer (PE) in Texas and Kansas, a Liveryman of the Worshipful Company of Engineers, an ISA Certified Automation Professional® (CAP®), a UK registered Chartered Engineer (CEng), a European registered Engineer (Eur Ing), a GIAC Global Industrial Cyber Security Professional (GICSP), and a Certified Mission Critical Professional (CMCP). He was the 2021 President of the International Society of Automation (ISA) and is a Life Fellow of the Society. He is a Fellow of the Institution of Engineering and Technology, and a member of the Water Environment Federation (WEF) Safety and Security Committee. Mustard writes and presents on a wide array of technical topics and is the author of “Industrial Cybersecurity, Case Studies and Best Practices” and ‘Mission Critical Operations Primer”, both published by ISA and “A Guide to Cybersecurity for Water and Wastewater Utilities”, published by WEF. He has also contributed to other technical books, including “Project Management: A Technician’s Guide”, published by ISA, WEF’s “Design of Water Resource Recovery Facilities, Manual of Practice No.8, Sixth Edition” and “The Digital Twin” book., published by Springer Nature. Mustard’s previous and current client list includes: the UK Ministry of Defence; NATO; major utilities, such as Anglian Water Services and Sydney Water Corporation; major oil and gas companies, such as bp, BG Group and Shell; Fortune 500 companies, such as Quintiles Laboratories; and other leading organizations.

Share this insight

Report a copyright concern

Incident Response Testing · CSAFI