Jobs · Engineering

Systems Architect, Disaster Recovery

Donatech Corporation · United States · 2 wk ago
RemoteRemoteEngineeringContract

This position requires the candidate to be a W2 employee of Donatech. U.S. citizenship is required.

About the Role

Designs, develops, and oversees system architectures for enterprise-wide disaster recovery (DR) and high availability (HA) solutions. Serves as the technical authority for DR during design reviews and program risk assessments, ensuring alignment with security, compliance, and cost optimization policies.

Responsibilities

  • Define end-to-end DR and HA architectures for enterprise workloads, incorporating multi-region cloud, hybrid, and on-premises solutions.
  • Develop architectural blueprints, reference designs, and pattern libraries aligned with client security, compliance, and cost optimization policies.
  • Design and implement automated failover, replication, and failback mechanisms (e.g., Site Recovery Manager, Kubernetes-based HA, database mirroring, storage-level replication).
  • Evaluate and integrate emerging technologies (e.g., Immutable Infrastructure, Chaos Engineering, Serverless DR) to improve resiliency and reduce mean time to recover (MTTR).
  • Ensure all DR solutions comply with corporate policies (CRX 301, CRX 302) and regulatory requirements (e.g., NIST 800-34, ISO 22301, FedRAMP).
  • Create and maintain DR documentation, runbooks, and test plans; conduct periodic reviews and updates.
  • Lead full-scale DR exercise planning, execution, and post-mortem analysis for multi-site, multi-cloud environments.
  • Define success criteria, metrics, and KPIs for DR/HA solutions; report findings to senior leadership and stakeholders.
  • Partner with IT Infrastructure, Cloud Engineering, Application Development, Security, and Governance teams to embed DR/HA considerations early in the SDLC.
  • Serve as the technical authority for DR during design reviews (SRR, PDR, CDR, TRR) and program risk assessments.
  • Conduct risk assessments, threat modeling, and capacity planning to anticipate emerging resiliency challenges.
  • Drive adoption of Model-Based Systems Engineering (MBSE) and automated documentation tools to keep architecture artifacts current.
  • Review existing DR plan architectures across the enterprise, assessing alignment with current resilience standards and organizational Recovery Objectives (RTO/RPO).
  • Collaborate with internal teams (Application Owners, IT Service Managers, Engineering) to update and refine DR plans.
  • Develop and implement remediation plans to modernize legacy systems and ensure compliance with corporate policies and regulatory requirements.
  • Track progress and report status to senior leadership, providing insights into plan modernization efforts and risk mitigation strategies.

Requirements

  • 5+ years of experience designing and implementing DR/HA solutions for enterprise-scale workloads in cloud, hybrid, and on-premises environments.
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical discipline (or equivalent experience).
  • Proven hands-on experience with cloud platforms (AWS, Azure, GCP) and related services (e.g., Disaster Recovery, Site Recovery Manager, cross-region replication, networking, IAM).
  • Strong understanding of networking, storage, virtualization, container orchestration (Kubernetes), and database technologies as they relate to resiliency.
  • Excellent written and verbal communication skills; ability to translate complex technical concepts for both technical and non-technical audiences.
  • U.S. citizenship required.

Skills

  • Familiarity with automated disaster recovery (DR) solutions, including:
    • Amazon Web Services (AWS) Disaster Recovery Service (DRS): Configuration and management of replication, failover, and failback processes for AWS workloads.
    • Microsoft Azure Site Recovery (ASR): Setup and management of site recovery between on-premises environments, Azure, and other clouds.
    • Zerto: Hands-on experience with continuous data protection (CDP) and near-zero RPO replication across VMware, Hyper-V, and cloud environments.
    • Veeam Backup & Replication: Agentless backup, replication, and automated failover testing for virtual, physical, and cloud workloads.
    • IBM Resiliency Services (formerly IBM Disaster Recovery as a Service): Managed DR services for hybrid cloud environments, including integration with IBM Cloud and on-premises infrastructure.
  • Experience with DR automation, including:
    • Scripting and integration with IaC tools (Terraform, CloudFormation) and CI/CD pipelines (Jenkins, GitLab) to automate DR workflows.
    • DR exercise planning and execution, including defining success criteria, metrics, and KPIs for recovery processes.
  • Strong analytical skills: ability to perform risk assessments, impact analysis, and cost-benefit modeling for DR solutions.
  • Hybrid/multi-cloud deployments, with ability to manage DR across multiple cloud providers and on-premises environments.
  • Advanced certifications (e.g., AWS Certified Solutions Architect - Professional, Azure Solutions Architect Expert, VMware VCAP DCV).

Similar jobs

IT Disaster Recovery Manager

Valvoline Global OperationsUnited States· 1 mo ago
RemoteInformation Technologyapply on jobs.valvolineglobal.com