Business Recovery After a Cyberattack Requires Dependency Mapping, Evidence, and Coordination
- During sophisticated intrusions, defenders should assume that management systems may have been compromised.
- Across large-scale recoveries, undocumented dependencies usually create more downtime than damaged hardware.
- Renfrow feels that every hour spent discovering relationships during an incident is an hour customers remain offline.
- The Fenix24 team often says that organizations don't recover applications; they recover ecosystems.
- The organizations that recover fastest are those that continuously test complete business workflows rather than isolated infrastructure.
Heath Renfrow, CISO and Co-Founder of Fenix24, says recovery from ransomware or another large-scale cyberattack should begin with the business capability that reduces risk fastest, not the most important-looking server.
Renfrow’s security leadership spans the U.S. Navy, Department of Defense, Army Information Management Command, Army Corps of Engineers, and Army Healthcare, where he served as its first CISO.
He explains how undocumented dependencies can leave restored systems unable to serve customers. Providers must map relationships, assign responsibilities, and give response and restoration teams the same operational picture.
Renfrow adds that forensic evidence may require preserving some systems while clean restoration begins elsewhere. Before production resumes, technical teams and business leaders must verify system integrity, working business processes, and the remaining risk.
Read on to learn why Renfrow distinguishes “system restored” from “business restored” and how investigation and recovery can advance together.
Vishwa: Across recoveries you have observed, which overlooked condition has caused delay after containment? Can you explain one through a real example?
Heath: The biggest misconception is that containment marks the beginning of recovery. In reality, containment often exposes a much larger problem: organizations rarely understand the dependencies that allow critical services to function.
One financial services client believed restoring several hundred virtual machines would bring operations back online. While the infrastructure recovery went exactly as planned, the organization quickly discovered that authentication services relied on certificate authorities, DNS, Active Directory replication, application service accounts, and integrations that had never been documented. The individual systems were functioning, but the business still could not operate.
We often say organizations don't recover applications; they recover ecosystems. In this case, the delay wasn't caused by technology failure, but from missing knowledge about how the ecosystems worked together.
Across large-scale recoveries, undocumented dependencies usually create more downtime than damaged hardware or encrypted files. Every hour spent discovering relationships during an incident is an hour customers remain offline.
Recovery planning should focus less on simply maintaining inventories and more on understanding operational dependencies. If organizations cannot clearly explain what must exist before an application can serve a customer, they likely don’t understand what must be recovered first.
Vishwa: Beyond immutability, what tests establish that a backup can support a clean, operationally viable restoration?
Heath: Immutability answers the important question of whether a backup can be altered, but it does not answer the question organizations ultimately care about: whether the business can operate after restoration.
Organizations should regularly perform full recovery exercises instead of relying only on simple restore tests. Successful validation requires more than confirming systems are restored.
It includes:
- verifying Active Directory health
- application functionality
- database consistency
- authentication services
- Networking
- DNS
- certificate services
- endpoint protection
- monitoring, and
- integrations with external systems
Another critical component is measuring operational objectives by asking whether the recovery meets required recovery time objectives (RTOs) and recovery point objectives (RPOs), whether users can authenticate, whether transactions can complete, and whether business processes can resume.
Additionally, security validation is another critical step. Restoring malware or compromised identities simply recreates the incident.
The organizations that recover fastest are those that continuously test complete business workflows rather than isolated infrastructure. A successful backup is not measured by whether a server boots, but is measured by
- whether customers can place orders,
- clinicians can access records, or
- employees can perform their jobs safely
Vishwa: How do you determine whether security consoles and logs can be trusted after an intrusion? When their integrity cannot be established, which telemetry helps?
Heath: During sophisticated intrusions, defenders should assume that management systems may have been compromised until their integrity can be verified. A security console being online does not automatically make it a trusted source of truth.
We begin by:
- Validating administrative access history,
- Configuration changes,
- Log continuity,
- Clock synchronization, and evidence of tampering.
- Missing events, unexpected gaps, disabled logging, or unexplained administrative actions all reduce confidence in the data.
When that confidence cannot be established, independent telemetry becomes critical for reconstructing what occurred. Elements such as
- network flow records,
- firewall logs,
- cloud audit logs,
- identity provider events,
- DNS activity,
- endpoint forensic artifacts,
- virtualization platforms,
- storage systems, and
- backup metadata can all provide independent evidence to help piece together attacker activity.
It’s important to note that no single source provides the complete picture. Confidence comes from corroborating evidence across multiple independent systems.
One lesson repeated throughout large-scale incident response is that attackers frequently erase evidence in one location while overlooking another. As a result, recovery decisions should rely on converging evidence rather than trusting any individual console.
Vishwa: How should responders reconcile business interruption costs with regulatory obligations when setting restoration priorities?
Heath: Restoration priorities should never be determined solely by financial impact or by compliance requirements. Both are critical components of business resilience.
- The first priority should be protecting life, safety, and legally required operations.
- From there, organizations should focus on restoring the capabilities needed to meet contractual, regulatory, and operational obligations.
- Revenue-generating systems certainly matter, but restoring them before foundational identity, networking, or security services are available and validated can create additional delays.
I encourage organizations to define restoration tiers before an incident occurs by combining business impact analyses with regulatory requirements and technical dependencies. During an active incident, executives should continuously reassess priorities as new intelligence emerges. What appears financially critical may depend on infrastructure that has not yet been secured or validated.
The organizations that recover most effectively do not ask, “Which server should we restore first?” They ask, “Which business capability reduces organizational risk the fastest while allowing safe, compliant operations to resume?”
Vishwa: What technical evidence should be documented before a restored system returns to production, and who should participate in that decision?
Heath: Returning a system to production should be treated as an evidence-based decision, not just a technical milestone. At a minimum, responders should document
- malware validation results
- vulnerability status
- configuration integrity
- identity verification
- logging functionality
- backup provenance
- patch levels
- application testing, and
- confirmation that required business workflows execute successfully
Just as important is documenting what remains unknown, since recovery always involves measured risk, and leadership needs to understand those assumptions before making a decision.
Additionally, the approval process should never belong to a single technical team. Security, infrastructure, application owners, business leadership, compliance teams, and executive decision-makers all bring different perspectives on acceptable risk.
I often describe this as moving from “system restored” to “business restored.” Those are not the same event. Production should resume only when technical evidence supports the decision and business leadership accepts any remaining operational risk.
Vishwa: Which early deliverables demonstrate that a provider can coordinate and execute an enterprise-scale restoration? What indicators should be used to assess progress?
Heath: The earliest deliverables reveal whether a provider has a defined execution framework or is simply reacting to events as they unfold. Within the first several days, organizations should expect to see an operational governance structure, restoration priorities, dependency mapping, infrastructure inventories, communication cadence, executive reporting, and a documented recovery plan with clearly assigned ownership.
Progress should be measured through objective operational metrics rather than activity counts. This should include the percentage of validated identities restored, critical applications brought back online, dependency blockers removed, infrastructure rebuilt, backup validation completed, and successful business process testing.
Large recoveries succeed because of disciplined coordination as much as technical expertise. Hundreds of specialists may participate, but progress depends on everyone working from the same operational picture and moving toward the same recovery objectives.
Vishwa: When incident response and restoration providers recommend different approaches, what evidence should be used to resolve it?
Heath: Incident responders and restoration teams often optimize for different objectives, with one focused on understanding the attack and preserving investigative integrity, while the other is focused on restoring business operations. Neither perspective is sufficient on its own.
Decisions should be driven by evidence including forensic findings, attacker persistence mechanisms, identity compromise, infrastructure integrity, backup validation, dependency analysis, and quantified business impact.
For example, forensic evidence may support preserving certain systems for investigation, while recovery evidence may demonstrate that clean, validated restoration can safely begin in parallel. These objectives are often complementary rather than conflicting.
Executive leadership should require both teams to present measurable evidence supporting their recommendations instead of relying on subjective assessments or competing priorities.
The most successful recoveries occur when investigation and restoration operate as coordinated workstreams, sharing a common operational picture that allows organizations to reduce business downtime without compromising investigative integrity.
Vishwa: When ransomware disrupts multiple systems, how do you decide what to restore first without a complete technology map? Which overlooked service commonly delays recovery?
Heath: Without an accurate dependency map, organizations often restore infrastructure in the wrong order. Servers become available, but applications remain unusable because foundational services are still unavailable.
When documentation is incomplete, responders should reconstruct dependencies using network relationships, authentication patterns, application communications, infrastructure telemetry, interviews with application owners, and historical operational knowledge.
Identity services remain the most frequently underestimated dependency, yet Active Directory, certificate services, DNS, identity federation, service accounts, and privileged access underpin nearly every modern application. Until identity functions correctly, many other restored systems cannot operate reliably.
This is why technology inventories alone are insufficient, and organizations need operational dependency maps that explain how business capabilities actually function.
Recovery priorities should therefore follow dependency chains rather than asset lists. Restoring what appears most important first often delays recovery more than restoring what enables everything else.
Vishwa: What evidence gives responders confidence to retain and restore a compromised identity environment rather than rebuild it?
Heath: Rebuilding identity should never be the default response. Modern identity environments are deeply integrated with applications, cloud platforms, certificates, and operational processes, and rebuilding unnecessarily can significantly extend downtime.
Confidence comes from evidence showing the environment remains structurally trustworthy. We examine privileged account activity, domain controller integrity, authentication logs, replication health, administrative changes, persistence mechanisms, certificate infrastructure, service accounts, trust relationships, and indicators of attacker control.
Just as important is understanding the attacker's objective. A ransomware operator seeking encryption presents different risks than an adversary focused on long-term identity persistence.
When evidence demonstrates compromise has been identified, persistence removed, privileged access re-established, and monitoring restored, retaining the existing identity environment is often both faster and safer than rebuilding.
The decision should always be based on validated forensic evidence, not assumptions or fear.
Vishwa: What did your tenure as CISO for U.S. Army Healthcare teach you about assigning decision authority when technical recovery, operational continuity, and safety requirements conflict?
Heath: One of the most valuable lessons I learned was that technical authority and operational authority are not the same thing. Security teams understand cyber risk, engineers understand technical recovery, and operational leaders understand mission impact, but none of them should make enterprise decisions independently during a crisis.
In healthcare environments specifically, every decision potentially affected patient care, regulatory obligations, and operational continuity simultaneously. That required a structured governance model where technical experts presented evidence, operational leaders evaluated mission impact, and executives accepted organizational risk.
I continue to use that philosophy today. Decisions should be evidence-driven, roles should be clearly defined before incidents occur, and authority should be delegated according to responsibility, not hierarchy.
The objective is not consensus for its own sake. The objective is ensuring the right people make informed decisions based on verified information while maintaining accountability.
Whether supporting healthcare, financial services, manufacturing, or government, that governance model consistently produces faster, safer, and more defensible recoveries.





