Platform Engineer IV
About the role
On our Cloud Reliability Engineering team, you’ll architect, build, and evolve the fault-tolerant infrastructure and application patterns the enterprise depends on to meet its availability and recovery commitments. Working autonomously and with wide latitude, you’ll serve as a strategic technical lead for the most complex, cross-functional resiliency solutions running in AWS—partnering with product, application, and platform leaders to define reliability strategy, from SLOs and error budgets to disaster-recovery posture.
Responsibilities
- Design multi-region, multi-AZ architectures and publish the reference patterns teams adopt across the enterprise
- Lead chaos experiments and game days that prove out resilience
- Automate self-healing and orchestrated failover to hit strict RTO/RPO targets
- Coordinate incident response for high-severity outages, authoring the blameless post-mortems that permanently fix failures
- Mentor senior and lead engineers and raise the bar for reliability practices org-wide
Requirements
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical cloud experience
- At least 10 years in Site Reliability Engineering (SRE), DevOps, or Cloud Architecture roles, including demonstrated experience as a technical lead on enterprise-scale reliability initiatives
- Deep expertise in enterprise backup, recovery, disaster recovery, and cyber resilience on AWS—AWS Backup, cross-region recovery, immutable backups, recovery orchestration, application-consistent recovery, RTO/RPO planning, recovery testing, and backup governance
- Expertise in AWS resiliency services (e.g., Application Recovery Controller and Resiliency Hub), multi-region architectures, modern application architecture, and Kubernetes/container orchestration (e.g., EKS scaling and networking)
- Proven experience defining SLOs/SLIs and error budgets and automating self-healing and orchestrated failover to meet strict RTO/RPO targets
- Hands-on experience leading chaos/failure-injection experiments with tools such as Chaos Mesh, Gremlin, or AWS Fault Injection Simulator (FIS)
- Advanced Infrastructure-as-Code skills (e.g., Terraform for immutable infrastructure) and observability expertise with Prometheus, OpenTelemetry, or Datadog
- Working knowledge of workflow/streaming platforms such as AutoSys, Temporal.io, and Kafka is a plus
- Proven ability to mentor engineers, drive operational excellence, and articulate complex technical challenges and solutions to business and IT executive audiences
Pay
Charlotte base salary range: $136,749–$218,798. In addition to a highly competitive base salary, per plan guidelines and vesting requirements, you will also be eligible for an individual annual performance bonus, plus Capital’s annual profitability bonus, plus a retirement plan where Capital contributes 15% of your eligible earnings.