Site Reliability Engineer, Lead
About the role
The Lead Site Reliability Engineer (SRE) role involves ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This includes leading the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments. The role also involves driving root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available, efficient, and scalable services.
Responsibilities
- Ensure the reliability, performance, scalability, and security of critical production systems and platforms.
- Lead the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments.
- Partner closely with DevOps, infrastructure, and security teams to improve system resilience and reduce operational risk.
- Drive root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available, efficient, and scalable services.
- Build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms.
- Work with the development and operation teams to evaluate the health, stability, and reliability of systems and platforms.
- Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems.
- Strengthen the security posture and safeguard national interests.
Requirements
- 8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack.
- 8+ years of experience with Linux systems administration and networking fundamentals within AWS.
- Experience with Python scripting and automation.
- Experience with Infrastructure as Code using Terraform and Terragrunt.
- Knowledge of Kubernetes administration, troubleshooting, and operations.
- TSCI clearance with a polygraph.
Qualifications
- Bachelor’s degree and 8+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering in lieu of a degree.
- Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date.
Skills
- Experience with deploying and managing OpenTelemetry.
- Experience with AWS CloudWatch, AWS EKS, and related AWS services.
- Experience managing Kubernetes environments through Rancher.
- Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management.
- Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence.
- Knowledge of distributed systems, microservices architectures, and containerized workloads.
- Knowledge of NIST 800-53 and NIST-190.
- Master’s degree in a relevant field.
Benefits
- Health, life, disability, financial, and retirement benefits.
- Paid leave.
- Professional development.
- Tuition assistance.
- Work-life programs.
- Dependent care.
- Recognition awards program.
Pay
The projected compensation range for this position is $99,000.00 to $225,000.00 (annualized USD).
Schedule
Full-time and part-time employees working at least 20 hours a week on a regular basis are eligible to participate in Booz Allen’s benefit programs. Individuals that do not meet the threshold are only eligible for select offerings, not inclusive of health benefits.