Senior Disaster Recovery and Resilience Engineer

Posted 20 Hours Ago
Be an Early Applicant
3 Locations
Hybrid
Senior level
Blockchain • Energy • Cryptocurrency
The Role
Hands-on role to assess, implement, test, and document backup, restore, failover, and recovery capabilities. Inventory critical systems, design and automate backup and restoration, run recovery exercises, produce runbooks, validate recoverability, measure RTO/RPO, and train system owners. Collaborate with Security, SRE, DevOps, QA, and application teams to harden shared recovery capabilities and transfer operational ownership.
Summary Generated by Built In

Senior Disaster Recovery and Resilience Engineer

About the Role

SunCore Digital is seeking a hands-on Senior Disaster Recovery and Resilience Engineer to assess, implement, test, and document service and data recovery capabilities.

This is an engineering position, not solely a policy, compliance, or coordination role. The successful candidate will identify recovery risks, build and harden shared recovery capabilities, automate recovery-specific processes where practical, and prove that critical systems and data can be restored. Once system-specific procedures are tested and documented, the owning teams will be trained to execute and maintain them.

The engineer may draw on application developers when recovery improvements require changes to application code, databases, data flows, or system-specific behavior. Recovery procedures must ultimately be understood and maintainable by the teams that own the affected systems.

Responsibilities

Recovery assessment

  • Inventory critical applications, services, databases, storage systems, queues, infrastructure, and external recovery dependencies.

  • Determine what is currently backed up and how those backups are created, retained, protected, and monitored.

  • Identify systems with missing, incomplete, unverified, or person-dependent recovery processes.

  • Assess risks related to data loss, service loss, infrastructure failure, configuration loss, credential availability, and third-party dependencies.

  • Distinguish configured backups from proven recoverability.

  • Identify manual recovery procedures and undocumented knowledge.

  • Document confirmed capabilities, untested capabilities, and unknowns.

Backup and restoration

  • Design and implement backup improvements for critical data and configuration, then transfer ongoing operation to the designated long-term owner.

  • Translate approved recovery, retention, security, and access requirements into backup controls and monitoring.

  • Create, test, and harden restoration procedures with system owners, then transfer routine execution and system-specific maintenance to those teams.

  • Execute restoration tests using representative environments and data.

  • Validate backup completeness and integrity.

  • Measure recovery duration and potential data loss during exercises.

  • Establish monitoring and escalation so failed or incomplete backups are detected by the accountable long-term owner.

  • Automate backup verification, restoration validation, recovery evidence collection, and other repeatable recovery processes where practical; transfer routine operation of system-specific automation to the owning teams after it is tested and documented.

A backup will not be treated as reliable solely because a scheduled job reports success. Restoration must be tested.

Service recovery and failover

  • Design and implement shared service-recovery and failover capabilities, transfer supporting shared components to the designated long-term owner, and ensure system-owning teams retain responsibility for application-specific recovery behavior and procedures after handoff.

  • Identify the infrastructure, application, database, network, access, and vendor dependencies required during recovery.

  • Establish safe recovery sequencing.

  • Develop procedures for partial outages, regional failures, infrastructure loss, data corruption, and service dependency failures.

  • Implement shared recovery mitigations directly and coordinate application-specific changes with system owners, who retain responsibility for those changes.

  • Test restored services for technical functionality.

  • Work with QA to validate that critical business workflows and data remain correct after recovery.

  • Document conditions under which recovery, rollback, or failover may be unsafe.

Recovery objectives and planning

  • Work with engineering and business leadership to document recovery requirements for critical systems.

  • Translate approved business requirements into technical recovery capabilities.

  • Measure actual recovery performance against defined expectations.

  • Identify where current architecture cannot meet required recovery expectations.

  • Provide technical options and evidence to support leadership decisions.

  • Maintain a recovery dependency map for critical systems.

Final recovery priorities, risk acceptance, and business requirements remain with leadership.

Disaster-recovery exercises

  • Plan and execute recovery exercises.

  • Develop test scenarios for service, infrastructure, data, credential, and dependency failures.

  • Ensure exercises produce objective evidence.

  • Record results, defects, recovery times, data-loss observations, and unresolved risks.

  • Implement shared recovery corrective actions and coordinate system-specific corrective work with the owning teams.

  • Repeat exercises after material changes.

  • Ensure procedures can be followed by qualified staff who did not author them.

Documentation and knowledge transfer

  • Establish recovery runbook standards and create initial system-specific runbooks with the owning teams.

  • Document required access, tools, credentials, dependencies, procedures, validation steps, and escalation paths.

  • Clearly label untested procedures.

  • Train system owners and relevant operational staff to execute and maintain the system-specific procedures they own.

  • Reduce reliance on undocumented individual knowledge.

  • Ensure application teams understand their ongoing recovery responsibilities.

  • Maintain shared recovery evidence and exercise results; system owners maintain the accuracy of their service-specific runbooks.

Security and collaboration

  • Work with Security Operations to review recovery implementations that affect production access, sensitive data, credentials, networks, infrastructure, or security controls.

  • Work with SRE on service dependencies, failure behavior, observability, and operational response.

  • Define recovery requirements for infrastructure-as-code, environment recreation, artifacts, and delivery tooling; work with DevOps and platform staff to implement those requirements within their shared capabilities.

  • Work with application engineers on application-specific recovery changes.

  • Work with QA on post-recovery functional and data validation.

  • Escalate recovery risks that cannot be mitigated within current architecture or resources.

Initial Priorities

  • Inventory existing backup and recovery capabilities.

  • Identify critical systems with no confirmed recovery path.

  • Verify who owns each backup and recovery process.

  • Determine when critical backups were last restored successfully.

  • Establish repeatable restore testing.

  • Document initial recovery requirements and dependencies.

  • Implement the highest-priority recovery mitigations.

  • Create and test recovery runbooks.

  • Identify person-dependent and manual recovery processes.

  • Produce a factual recovery-readiness assessment before broader external use.

Required Qualifications

  • Five or more years of experience in disaster recovery, infrastructure resilience, backup and recovery engineering, SRE, cloud infrastructure, systems engineering, or a related technical field.

  • Hands-on experience implementing backup, restore, recovery, and failover capabilities.

  • Experience performing recovery exercises rather than only writing recovery plans.

  • Experience recovering databases, cloud infrastructure, applications, configurations, and storage systems.

  • Experience with recovery automation and scripting.

  • Understanding of recovery-time and recovery-point concepts.

  • Experience documenting and validating recovery dependencies.

  • Familiarity with cloud security, identity, secrets, networking, and access requirements during recovery.

  • Experience troubleshooting complex system failures.

  • Ability to work with application developers on recovery-related code and data changes.

  • Strong runbook and technical-documentation skills.

  • Ability to identify unknowns and avoid treating untested procedures as proven.

  • Comfort working in a small organization where recovery practices are still being established.

Bonus Points

  • Experience preparing a platform for its first external users.

  • Experience establishing a disaster-recovery capability from an early stage.

  • Experience with data-intensive, financial, digital-asset, telemetry, or operational systems.

  • Experience with infrastructure as code.

  • Experience with controlled disaster-recovery or resilience exercises.

  • Experience recovering event-driven or distributed systems.

  • Experience in a remote, asynchronous environment.

What Success Looks Like

  • Critical systems and data have documented recovery requirements.

  • Backup ownership, schedules, retention, and locations are known.

  • Critical backups have been restored successfully.

  • Recovery procedures are documented, tested, and repeatable.

  • Actual recovery times and data-loss exposure are measured.

  • Systems without viable recovery paths are visible to leadership.

  • High-priority recovery mitigations are implemented.

  • System-owning teams understand and can maintain their recovery procedures.

  • Recovery capability does not depend entirely on one individual.

Compensation & Benefits

●        Base Compensation: Competitive, based on experience, portfolio strength, and geographic location.

●        Milestone Bonuses: 10% milestone bonus awarded to every team member assisting with the MVP build-out.

●        Annual Performance Bonus: Represents a significant percentage of total compensation.

●        Benefits for Domestic Hires: Health, dental, and vision plans.

●        Annual Paid Offsite: Team retreats in Caribbean, Hawaii, ski destinations, and other exciting locations.

Why Join SunCore Digital?

●        High-impact role in a rapidly scaling digital company.

●        Fully remote team with async flexibility.

●        Direct collaboration with executive leadership and best-in-class marketing/design partners.

●        Opportunity to shape Security processes and mentor junior team members.

●        Competitive compensation with performance upside.

Skills Required

  • Five or more years of experience in disaster recovery, infrastructure resilience, backup and recovery engineering, SRE, cloud infrastructure, systems engineering, or related technical field
  • Hands-on experience implementing backup, restore, recovery, and failover capabilities
  • Experience performing recovery exercises (not just writing plans)
  • Experience recovering databases, cloud infrastructure, applications, configurations, and storage systems
  • Experience with recovery automation and scripting
  • Understanding of recovery-time and recovery-point objectives (RTO/RPO)
  • Experience documenting and validating recovery dependencies
  • Familiarity with cloud security, identity, secrets, networking, and access requirements during recovery
  • Experience troubleshooting complex system failures
  • Ability to work with application developers on recovery-related code and data changes
  • Strong runbook and technical-documentation skills
  • Comfort working in a small organization where recovery practices are still being established
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2025

What We Do

SunCore Digital is a blockchain services company specializing in the development of digital infrastructure for cryptocurrency mining. The company provides a dependable energy platform to power mining equipment and offers users long-term rights to mine, focusing on accelerating the build-out of digital infrastructure to support the blockchain ecosystem.

Similar Jobs

Octus Logo Octus

Business Data Associate

Fintech • News + Entertainment • Software • Database • Financial Services
Easy Apply
Hybrid
London, Greater London, England, GBR
808 Employees

Notion Logo Notion

Customer Success Manager

Artificial Intelligence • Productivity • Software
Hybrid
London, Greater London, England, GBR
1000 Employees

CSC Logo CSC

Global Head HR Service Delivery Colleague Experience

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Hybrid
2 Locations
8500 Employees

Navan Logo Navan

Product Director, Service Operations

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
London, Greater London, England, GBR
3300 Employees

Similar Companies Hiring

Runwise Thumbnail
Greentech • Hardware • Real Estate • Software • Energy • PropTech
New York, NY
199 Employees
Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
108 Employees
Rain Thumbnail
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
New York, NY
100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account