When people talk about resilience in cybersecurity, the conversation often begins at the most visible moment: the incident.

What we do when something fails. Who responds. How we contain the problem. How we recover. How long it takes to restore operations.

All of that matters. But an organization does not become resilient on the day of an incident. That day simply reveals how resilient it already was.

The ability to respond and recover is built beforehand, often through less visible decisions: understanding dependencies, reducing fragility, preparing alternatives, defining responsibilities, and practicing how to act when normal conditions no longer apply.

Resilience does not mean preventing every incident

No organization can guarantee that nothing will go wrong.

We can reduce exposure, strengthen controls, and improve detection, but failures, mistakes, unexpected changes, and threats that cross some defenses will always exist.

That is why I find it more useful to think of resilience as an organization’s ability to keep making sound decisions when something important degrades.

That includes resisting, adapting, responding, and recovering, but also recognizing quickly which capabilities are essential and what temporary compromises are acceptable.

What is prepared before changes what happens during

Several questions appear operational, but they are actually design decisions.

What has to keep working?

Not every service has the same value or urgency.

A resilient organization needs to know which processes, systems, identities, data, and providers support what truly must continue. Without that clarity, recovery can turn into a competition between competing urgencies.

What do we depend on?

Important capabilities rarely depend on one system alone.

Applications, networks, identities, people, providers, data, and manual processes are interconnected. A dependency that is barely visible during normal operations can end up defining the real recovery time.

Understanding those relationships before an incident makes it easier to design alternatives and identify single points of failure.

How do we operate if part of the environment becomes unavailable?

Resilience requires options.

Sometimes that means a secondary platform. In other cases, a manual procedure, an alternate communication path, a trusted backup, or an explicit decision to operate with reduced capacity.

The goal is not to duplicate everything. It is to avoid allowing a single failure to turn a technical problem into a total and unnecessary disruption.

Who can make decisions under pressure?

Incidents do not only test technology. They also test governance and coordination.

When responsibilities are unclear, decisions slow down. When everyone waits for someone else’s approval, uncertainty grows.

Defining in advance who evaluates, who communicates, who authorizes changes, and when escalation is required can be as important as a technical control.

Recovery is not simply turning systems back on

Fast recovery is not always good recovery.

Returning a system to service without understanding what happened, validating integrity, or correcting the conditions that enabled the problem can restore operations to a state that is still fragile.

That is why recovery should answer at least three questions:

  • can we operate again?;
  • can we trust what we are restoring?;
  • have we reduced the chance of repeating the exact same problem?

That third question connects recovery with learning.

Learning is also a capability

Incidents generate many opportunities for improvement. The challenge is turning them into actual change.

A useful closeout should not produce only a report. It should produce decisions: controls that need strengthening, dependencies that need review, procedures that did not work, information that was missing, and capabilities that proved more important than expected.

If those decisions do not become part of normal work, an organization can restore systems without improving resilience.

Resilience shows up in the details prepared beforehand

Backups that can actually be restored. People who understand their role. Known dependencies. Tested alternatives. Clear priorities. Available communication channels. Criteria for degraded operations. Decisions that are recorded and reviewed.

None of these elements is especially dramatic.

Together, however, they determine whether an incident becomes a manageable disruption or a prolonged crisis.

That is why I believe resilience starts long before the incident.

It starts when an organization stops asking only how to prevent something from failing and also asks how to keep operating and making decisions when something inevitably does not go as planned.