
The Manual Override Problem
In early August 2026, two municipal water systems in New Jersey were hit by cyberattacks that disrupted their ability to remotely monitor and manage operational control systems.
The utilities were not publicly identified. What happened next is more important.
Operators shifted to manual operations. Water service continued. Customers retained access to safe drinking water. There was no reported disruption to service.
On its face, this is a cybersecurity success story. The technology was compromised. The fallback worked. Operations continued.
But from a governance perspective, the incident raises a different question: having a manual fallback is one thing. Does the person who needs to invoke it know when they are authorized to do so?
That is the Manual Override Problem. Because resilience is not simply having another way to operate. It is having the technical capability, operational competence, and decision authority to use that capability when normal systems fail.
A fallback system only creates resilience if the organization can activate it under pressure
The New Jersey utilities appear to have done exactly what organizations should be able to do during an operational technology disruption.
According to state officials, vulnerable internet-exposed control systems temporarily limited operators' ability to monitor or manage them remotely. Staff shifted quickly to manual operations, and service continued uninterrupted.
There is no public evidence that either utility experienced a decision-authority failure. Quite the opposite. Their rapid transition suggests that the operational response worked. That is precisely why the incidents are useful. They allow leadership teams to ask: what had to already be in place for that response to work?
Someone had to recognize that remote operations could no longer be trusted. Someone had to know that manual operation was technically possible. Operators had to know how to do it safely. And someone had to have authority to make the transition. The last requirement is easy to overlook.
Manual capability is not the same thing as operational readiness
The FBI and EPA have been unusually explicit about the role manual operations can play during the current wave of attacks against water and wastewater utilities.
Since July 27, utilities in at least seven states have reported incidents involving malicious access to operational technology. Attackers have remotely accessed internet-facing programmable logic controllers, changed IP addresses and passwords, and caused organizations to lose monitoring and control functionality. Reported operational effects have included pressure loss and flooding.
And according to the FBI, the extent of the impact depended partly on something deceptively simple: the organization's capability to switch to manual operations. The FBI and EPA now recommend that critical infrastructure operators practice and maintain the ability to operate OT systems manually, including routinely testing continuity plans, fail-safe mechanisms, backups, standby systems and other capabilities required for safe manual operation.
That is a technical and operational requirement. But there is another layer.
An organization can have technically sound manual procedures and still introduce delay if nobody knows when the threshold for using them has been crossed. The question is not only "can we operate manually," it is "under what conditions are we authorized to stop relying on the automated environment and take manual control."
What specific condition gives an operator authority to move a critical process into manual operation?
The decision may need to happen before the technical investigation is finished
During an OT cyber incident, cybersecurity teams understandably want to understand what happened. Which systems were accessed? What did the attacker change? Can the control environment still be trusted? Has malicious access been contained?
Those questions matter. But operations may be dealing with a different clock. Water still needs to flow. Pressure still needs to be maintained. Treatment processes still need to operate safely. Customers still expect service.
The two New Jersey utilities temporarily lost remote monitoring or management capability and moved quickly to manual operations. That illustrates an important principle: operational decisions and forensic decisions do not always run on the same timeline.
During a critical-infrastructure incident, the organization may need to change operations before cybersecurity teams can provide technical certainty. That means authority thresholds should be based partly on operational conditions, not solely on confirmed cyber findings. Loss of visibility, loss of remote control, unexpected equipment behavior, unsafe operating conditions, or inability to verify system integrity could each justify a change in operating mode before investigators know exactly what caused the problem.
Can Operations act when the consequence is clear but the cause is still being investigated?
Who actually owns the decision?
This is where many continuity plans become less precise. A procedure may say "switch to manual operations." But that describes an action. It does not necessarily establish authority.
Who can order the switch? The operator? A plant supervisor? The utility director? The incident commander? The city manager? Does cybersecurity need to approve it? Does executive leadership need to be notified first? And what happens at 2:00 a.m. when the person normally responsible cannot be reached?
The New Jersey incidents do not publicly tell us how those decisions were structured internally. Nor should we assume there was a problem. Instead, their successful response gives other organizations a useful test.
Decision authority should be as engineered as the fallback technology itself. If a critical function can operate in more than one mode, the organization should establish who can authorize the transition, what conditions trigger that authority, whether prior approval is required, who assumes authority when the designated decision-maker is unavailable, and how the decision is documented and communicated. Without those rules, a technical fallback can exist while still being difficult to activate under pressure.
If your automated environment became unreliable right now, could an operator tell you exactly who can authorize the fallback?
A successful fallback can hide the governance dependency underneath it
Because New Jersey customers experienced uninterrupted water service, these incidents could easily be summarized as: attack occurred, utility switched to manual, problem solved. That misses the more valuable lesson.
Successful resilience often looks uneventful from the outside precisely because preparation worked inside the organization. The technology may have provided a manual mode. But people still had to recognize the situation, make a decision and execute the transition. That makes manual operations a human governance capability, not simply a technical feature.
Consider what could happen if any one of those elements were missing. The fallback exists, but operators have not practiced it. Operators have practiced it, but nobody is certain when it can be invoked. Authority exists, but it sits with someone who cannot be reached. Operations wants to transition, but Cybersecurity wants systems preserved for investigation. The person with authority assumes someone else has already made the decision. None of these scenarios requires another vulnerability. They are organizational failure modes.
Resilience is not demonstrated by having a backup. It is demonstrated by the organization's ability to activate the backup quickly, safely and with clear authority.
Which of your organization's fallback capabilities has never actually been activated under real pressure, and how confident are you that it would work the first time it needs to?
Manual operations are becoming part of the cyber resilience conversation
The FBI's July 30 warning makes the operational stakes clear. Attackers targeting internet-exposed PLCs have been able to alter device configurations, interfere with visibility and control, and create physical operational effects.
Federal guidance therefore does not stop with stronger passwords, firewalls or removing PLCs from direct internet exposure. It specifically tells operators to practice and maintain the ability to operate OT systems manually, describing that capability as vital to restoring operations immediately after an incident.
For executive leadership, that recommendation should trigger a second exercise. Don't just test whether operators can perform the technical procedure. Test the decision environment around it. Give the team an ambiguous scenario: remote visibility disappears, equipment may still be operating, cybersecurity cannot yet confirm what has happened, the designated executive is unavailable. Then ask: who decides what happens next?
That is where a continuity plan becomes an operating capability.
A technical capability that has never been paired with a decision exercise is untested in the way that actually matters. The procedure has been rehearsed. The authority to invoke it has not.
The next time you test a fallback system, are you testing whether it works, or testing whether the right person knows they're allowed to use it?
Manual resilience requires more than manual controls
The New Jersey incidents point to five capabilities critical-infrastructure organizations should test.
1. Define the trigger
Teams should know which operational conditions justify abandoning normal automated or remote operation. The trigger should not depend exclusively on knowing the root cause.
2. Assign authority before the incident
Someone must have explicit authority to initiate the transition. That authority should be understood across Operations, Cybersecurity and leadership.
3. Build succession into the authority model
Critical decisions cannot depend on one person answering their phone. Organizations should know who assumes decision authority when the primary owner is unavailable.
4. Exercise the decision, not only the procedure
A successful tabletop should test more than whether operators remember which switch to flip. It should test whether participants recognize the threshold, know who owns the decision and can act without unnecessary delay.
5. Connect cyber recovery to operational continuity
Cybersecurity may be investigating the compromise while Operations is keeping the physical system running. Those responsibilities need to work together without creating decision paralysis.
A manual fallback is only as resilient as the decision system surrounding it
The two New Jersey utilities provide an encouraging example. Their remote operational capabilities were disrupted. Operators shifted to manual operations. Water service continued. Customers retained access to safe drinking water. That is what resilience is supposed to look like.
But it should also prompt every critical-infrastructure leader to ask what would happen inside their own organization. The question is not merely "can we operate manually." It is "who can make that call, under what conditions, and can they do it without waiting for the organization to achieve certainty."
Because the moment automation becomes unreliable is the wrong moment to discover that the fallback procedure is documented but the authority to use it is not.
The technology may have a manual override. Your governance needs one too.
Critical-infrastructure organizations routinely test backup systems, continuity procedures and technical recovery. They should test decision authority with the same discipline.
CrisisOS5™ examines how Operations, Cybersecurity, Legal, Communications and executive leadership function when normal systems become unreliable and consequential decisions need to happen before the full technical picture is available.
- Reveal which operational conditions your organization treats as authorization to switch to manual control
- Identify who can invoke a fallback procedure, and who assumes that authority when they're unavailable
- Test whether your continuity plans have been exercised as decisions, not just as procedures
- Clarify the handoff between whoever is investigating the compromise and whoever is keeping the physical process running
Sources
- Federal Bureau of Investigation and Environmental Protection Agency, "Malicious Cyber Actors Targeting Water and Wastewater Sector Internet-Facing Programmable Logic Controllers, Causing Operational Disruptions," July 30, 2026. View source
- ABC News, "2 New Jersey municipal water systems targeted in cyberattacks," August 5, 2026. View source
Source note: This case study is based on the FBI and EPA's July 30, 2026 public service announcement on attacks against water and wastewater operational technology, and ABC News reporting on the August 2026 New Jersey municipal water system incidents.