Operational Resilience: Why It Must Be a Daily Practice
Operational resilience is often discussed at the point where disruption has already begun.
A system has gone down. Communications are unavailable. An incident response team has been activated. Leaders need answers quickly, while regulators, customers and partners may all be waiting for updates. The organisation may already be operating against regulatory reporting deadlines and continuity obligations.
By that stage, however, many of the factors that will determine the quality of the response are already fixed. Teams either know how to work together under pressure, or they do not. Escalation paths either exist, or they need to be improvised. Access, permissions and workflows are either familiar and tested, or they become another problem to solve while the incident is unfolding.
Operational resilience is not created in the moment of crisis. It is built every day.
Cyber incidents become operational incidents
A major cyber incident rarely remains confined to the security team.
An attack may begin with the compromise of data, systems or identities, but the operational impact can spread quickly. Communications may become unreliable. Access to critical applications may be restricted. Teams may lose visibility into what is happening across the organisation.
Containment actions can add to that disruption. Security teams may deliberately isolate parts of the network to prevent an attacker from moving further through the environment. In doing so, they can also remove access to the systems employees normally rely on to communicate and coordinate.
Attackers may target communications and identity infrastructure, while defenders may need to shut systems down to contain a breach. Teams can therefore find themselves trying to respond at exactly the moment their normal ways of working have disappeared.
At that point, the problem is no longer simply technical. The organisation still has to make decisions, coordinate people and resources, maintain essential services and understand what is happening quickly enough to act.
Resilience is becoming a board and regulatory issue
Operational resilience is no longer only a concern for cyber and technology teams.
For regulated and nationally important organisations, there is increasing scrutiny on whether critical services can continue through technology failures, cyber incidents and third-party disruption. That shifts resilience from a technical preparedness issue into a broader governance, risk and business continuity priority.
In Singapore, the Monetary Authority of Singapore places business continuity and technology risk management expectations on financial institutions, while the Cyber Security Agency sets cybersecurity requirements for owners of Critical Information Infrastructure. For organisations operating across government, financial services and critical infrastructure, the question is therefore not simply whether systems can eventually be recovered, but whether essential operations can continue while disruption is still unfolding.
The same direction is visible internationally. Frameworks such as the EU’s Digital Operational Resilience Act place greater emphasis on ICT resilience, testing, incident management and third-party technology risk.
For cyber leaders, these requirements can also help frame the business case internally. Resilience investment is not only about reducing the likelihood or impact of an attack. It is about demonstrating that the organisation can continue operating, maintain governance and meet its obligations when technology is under pressure.
Meeting those expectations requires more than recovery planning. It requires secure, sovereign coordination that is already embedded in the way teams operate.
Operational resilience is increasingly something organisations need to demonstrate, not simply assume.
The platform you rely on during an incident should not be unfamiliar
Many organisations have fallback communications plans or dedicated emergency platforms. But a tool that sits unused for most of the year can introduce its own problems when it is finally needed.
Users may struggle to access it. Permissions or contact information may no longer reflect the current organisation. Workflows that exist in a plan may never have been tested under real conditions.
When the environment is already uncertain, every additional unfamiliar process adds friction.
The stronger model is not one where teams suddenly switch to an entirely different way of working when something goes wrong. It is one where the environment they rely on during disruption is already part of how they operate.
The platform trusted during an incident should already be part of daily operations.
Everyday operations are where resilience is built
Resilience is often associated with extraordinary events, but much of it is created through ordinary work.
Cyber and operational teams routinely share intelligence, coordinate investigations, manage escalations, connect with technical systems and make decisions across organisational boundaries. Those activities create immediate value, while also building the habits and operational context that become critical when conditions deteriorate.
A team that already knows how to coordinate across security, IT, operations and leadership does not need to invent that coordination model during an incident. Workflows that already route information, actions and decisions to the right people can continue to provide structure as response tempo increases.
Daily operational efficiency and crisis readiness are therefore not separate objectives. The same systems and workflows that improve coordination today can form the foundation of response tomorrow.
But resilience is not only about availability. It is also about control.
For government, financial services and critical infrastructure organisations, the ability to determine where operational data resides, how systems are deployed and which third-party dependencies they accept can be as important as whether a service remains online.
A response environment that stays available during disruption but introduces new dependencies, limits control over sensitive data or relies on infrastructure outside the organisation’s governance model does not fully solve the resilience problem.
Sovereign control needs to be designed into daily operations so that the same expectations around data handling, access, deployment and governance continue to apply when pressure rises.
Resilience has to hold as conditions change
Operational conditions do not move instantly from normal to crisis. There is often an intermediate period where risk is rising, more teams are becoming involved and the organisation needs greater situational awareness.
That makes resilience easier to understand as a continuum.
Everyday operations
Teams coordinate normal work, share threat intelligence, manage operational workflows and collaborate across functions. The focus is on efficient, secure execution while maintaining control over sensitive information and operational processes.
This is where daily value is created and where the relationships, workflows and familiarity required later are established.
Heightened pressure
An emerging threat, significant vulnerability, fraud event or operational issue changes the tempo. More stakeholders become involved, escalations accelerate and leadership requires clearer, more frequent information.
Established systems and practiced workflows allow teams to increase operational tempo without simultaneously having to create new communications channels, redefine responsibilities or learn unfamiliar tools.
Active disruption
Primary systems may now be degraded, compromised or deliberately isolated. The organisation needs secure fallback and out-of-band communications, clear decision authority and reliable ways to maintain a common operating picture.
These should not be three disconnected operating models. Resilient organisations need to move between them without losing context, control or trusted coordination.
Out-of-band communications are part of the architecture
Out-of-band communications become essential when the primary environment can no longer be trusted.
If corporate messaging, email or identity systems are unavailable, compromised or intentionally taken offline, teams still need a secure way to coordinate. But an effective out-of-band capability is more than an emergency chat channel.
Roles and permissions need to be established. Escalation paths need to be understood. Workflows need to be tested. Information still needs to be secure, governed and auditable.
Without that preparation, teams may fall back on consumer messaging applications. While that can restore basic communication, it can also introduce new risks around sensitive information, auditability, governance and chain of custody.
The goal is therefore not simply to have another communications channel available. It is to maintain a trusted operational environment when the primary one is under pressure.
That environment also needs to be sufficiently independent from the systems it is designed to replace. If a fallback channel relies on the same identity, infrastructure or vendor dependencies as the primary environment, a single incident may affect both.
This is where sovereignty becomes an operational requirement, not simply a data residency discussion. Organisations need control over where the environment runs, how it is administered, who can access it and which external dependencies they are prepared to accept.
Decision-making is an enterprise resilience capability
Communication matters because decisions depend on it.
During normal operations, leaders are accustomed to receiving structured information. They can bring together the relevant teams, establish what is known and decide what needs to happen next.
During a major cyber incident, that structure can deteriorate quickly. Trusted updates may become harder to obtain, different teams may hold different parts of the picture, and leaders may need to make consequential decisions before every question has been answered.
This is why operational resilience should not be measured only by system availability. A more useful question is whether the organisation can continue making sound decisions when the environment around those decisions becomes unreliable.
That capability extends well beyond cybersecurity.
IT may be restoring systems while operations works to maintain essential services. Risk and legal teams may be assessing exposure. Fraud teams may be responding to new criminal activity. Communications teams may need to update customers or the public, while executives make decisions about business continuity and regulatory obligations.
In financial services, for example, a disruption to payment or settlement infrastructure does not remain a technology problem for long. It can quickly become a customer, regulatory, operational and reputational issue.
Resilience therefore depends on trusted coordination across the organisation. Teams need shared context, clear decision authority and a persistent record of actions and decisions. They need to be able to move information across functions without losing accountability or creating further uncertainty.
For cyber leaders seeking support and investment, this is also where regulatory expectations can help translate technical risk into a wider business requirement around continuity, accountability and service availability.
Governance does not stop during disruption
The regulatory and governance clock continues to run even when internal systems do not.
Organisations may still need to gather facts, report incidents, preserve evidence and communicate with regulators or other stakeholders while normal communications channels are unavailable.
Continuity planning therefore has to preserve more than basic communication. It also has to preserve governance.
Who authorised a decision? What information was available at the time? What actions were taken? Who received an update? What evidence needs to be retained for regulatory or legal review?
This is where regulatory expectations established before an incident become operational requirements during one. Resilient organisations need secure access, appropriate controls and a persistent record of decisions throughout the incident lifecycle.
Finding another way to send messages is not enough.
Seven questions leaders should ask now
Operational resilience becomes easier to evaluate when it is treated as part of the operating model rather than as a contingency plan.
Leaders can start by asking:
- Are the systems we rely on during an incident used and tested regularly?
- Can teams still coordinate if corporate communications or identity systems are unavailable?
- Are incident roles, permissions and escalation paths already established?
- Can security, operational and executive teams work from the same trusted information?
- Do we retain control over where sensitive operational data resides and how the response environment is deployed?
- Have we identified the technology and third-party dependencies that could affect our ability to operate through disruption?
- Can we maintain an auditable record of decisions and meet regulatory obligations throughout the incident?
If the answer to any of these is unclear, that identifies an area where resilience may still depend on assumptions rather than practiced capability.
Resilience is built before it is needed
No organisation can predict every disruption, but organisations can decide how prepared they will be to operate through one.
That preparation is not limited to incident response plans, recovery targets or emergency communications tools. It is reflected in the systems people use every day, the workflows they understand, the relationships they have established and the control they retain over the environments supporting critical work.
Out-of-band communications remain an important part of that architecture. When primary systems are unavailable or untrusted, organisations need an independent way to coordinate securely. But the strongest resilience strategy starts earlier by making trusted, sovereign coordination part of daily operations.
Operational resilience ultimately depends on three things: continuity in the ability to operate, coordination in the ability to make trusted decisions, and control over the environment supporting both.
The organisations best prepared for disruption are not those that begin building resilience when an incident starts. They are the ones that practice it every day.
Meet Mattermost at GovWare 2026
Join Mattermost at GovWare 2026, 13–15 October at Marina Bay Sands, Singapore. Visit us at Booth E36 to explore how secure, sovereign coordination can help organisations keep critical operations moving every day and under pressure.