Cloud Platforms Engineer
ECS · Fairfax, VA · 1 wk ago
$180k–$210k/yrFull-time
Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in reliability and resiliency to design, operate, and continuously improve an Azure Government cloud infrastructure supporting mission-critical workloads for multiple coalition Mission Partner Network enclaves in support of the DoW community. This hybrid role is based in Fairfax, VA.
Responsibilities
- Design, deploy, and maintain highly available, fault-tolerant cloud infrastructure using Terraform and Infrastructure as Code principles.
- Architect and implement High Availability (HA) and Disaster Recovery (DR) solutions, including infrastructure redundancy, automated failover, backup and restoration, geographic resiliency, and recovery procedures.
- Manage the complete lifecycle of cloud infrastructure, including provisioning, configuration, operating system and platform maintenance, patching, upgrades, vulnerability remediation, and decommissioning.
- Develop automation for infrastructure deployment and routine operational activities to reduce manual administration, configuration drift, and the potential for human error.
- Implement and maintain monitoring, logging, alerting, and observability capabilities to identify infrastructure degradation, capacity constraints, performance issues, and potential service disruptions before they impact users.
- Participate in incident response, troubleshooting, and root-cause analysis for production infrastructure events and develop corrective actions to prevent recurrence.
- Define and maintain infrastructure reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, recovery time objectives (RTOs), and recovery point objectives (RPOs).
- Develop, maintain, and regularly validate disaster recovery procedures through recovery exercises, failover testing, and infrastructure restoration testing.
- Evaluate cloud infrastructure capacity, performance, availability, and scalability and recommend architectural or operational improvements.
- Partner with application development, cybersecurity, DevOps, and platform engineering teams to establish standardized deployment patterns and resilient cloud architectures.
- Maintain infrastructure documentation, operational procedures, architecture diagrams, runbooks, and recovery procedures required to support production environments.
- Provide technical leadership and guidance regarding cloud infrastructure reliability, resiliency, automation, and operational best practices.
- Other duties, as assigned.
Requirements
- U.S. Citizen.
- Active DoD Secret security clearance.
- Bachelor’s degree with 5+ years of related work experience.
- Ability to obtain a DoD 8140 IAT Level II Security+ (or higher) within 60 days of hire.
- Ability to work in a hybrid capacity in Fairfax, VA (up to 3 days in office).
- Ability to travel <20% throughout the lifespan of the Program to CONUS / OCONUS customer sites and government installations.
- Strong experience designing, deploying, and supporting highly available, fault-tolerant cloud infrastructure.
- Hands-on experience with Terraform and Infrastructure as Code (IaC) principles for provisioning and managing cloud resources.
- Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture, including redundancy, automated failover, backup and restoration, geographic resiliency, and recovery planning.
- Experience managing the full cloud infrastructure lifecycle, including provisioning, configuration, patching, upgrades, vulnerability remediation, maintenance, and decommissioning.
- Strong infrastructure automation skills with an emphasis on reducing manual administration, configuration drift, and operational error.
- Experience implementing and operating monitoring, logging, alerting, and observability solutions for production infrastructure.
- Strong troubleshooting and diagnostic skills, including incident response, root-cause analysis, and corrective action development.
- Understanding of Site Reliability Engineering concepts, including SLIs, SLOs, availability targets, RTOs, and RPOs.
- Experience developing and validating disaster recovery procedures, including failover exercises, recovery testing, and infrastructure restoration.
- Ability to evaluate infrastructure capacity, performance, scalability, availability, and resiliency and recommend architectural or operational improvements.
- Experience working collaboratively with application development, cybersecurity, DevOps, and platform engineering teams.
- Ability to develop and maintain technical documentation, including architecture diagrams, operational procedures, runbooks, and recovery documentation.
- Strong understanding of cloud infrastructure security, vulnerability management, and operational best practices.
- Demonstrated ability to provide technical leadership and guidance in infrastructure reliability, resiliency, automation, and cloud operations.
- Experience supporting production, enterprise, regulated, or mission-critical environments is highly desirable.
- Strong problem-solving and decision-making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.
- Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end-users to executive management).
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Computer Networking, Computer Security, or other Science, Technology, Engineering and Mathematics (STEM) discipline.
- Demonstrated experience administering and engineering cloud infrastructure in production environments.
- Strong hands-on experience with Terraform and Infrastructure as Code (IaC) practices.
- Experience developing and maintaining infrastructure automation for provisioning and operational tasks.
- Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture and implementation.
- Experience with monitoring, logging, alerting, and observability solutions.
- Strong operational troubleshooting, incident response, and root-cause analysis skills.
- Experience supporting enterprise-scale, regulated, or government environments.
- Familiarity with security, compliance, vulnerability management, and governance requirements in cloud environments.
- Ability to work effectively across infrastructure, cybersecurity, DevOps, platform engineering, and application teams.
- Strong understanding of cloud reliability, resiliency, scalability, and operational best practices.
Pay
$180,000 - $210,000
Benefits
Comprehensive benefits package available.