Senior Site Reliability Engineer
Job Summary
Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team where you grow your career, shape standards, and collaborate with peers - while also serving as the dedicated reliability partner for one of Drata's product engineering teams across the full lifecycle of their work. This is a highly technical role at the intersection of software engineering and systems engineering. The best SREs at Drata are engineers first: they solve problems by building solutions, not by executing manual processes. Automation is a core value, and nowhere is that more visible than in how we approach reliability. Our infrastructure runs on AWS across multiple accounts, defined entirely in Terraform.
What You’ll Do
Reliability Architecture for Your Product Team
Partner with product engineering leads and staff engineers to define SLOs and SLIs for critical services
Eliminate Toil Through Engineering
Central SRE Platform Work
What You'll Bring
6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building and maintaining scalable, resilient services
Robust knowledge of cloud computing technologies: Terraform, Docker, Git, and Linux
Hands-on experience with Datadog for monitoring, alerting, dashboards, SLO tracking, and distributed tracing
Experience building software systems as a software engineer
Experience developing tooling and automation in Python and/or Bash
Experience with CI/CD pipeline automation, specifically GitHub Actions
Experience with disaster recovery practices and incident management
Strong understanding of observability concepts - monitoring, logging, distributed tracing, and metrics - and how to apply them to production systems
Experience with container orchestration and deployment technologies including AWS ECS Fargate and/or Kubernetes
Experience working with relational databases (MySQL proficiency is a plus)
Nice To Have
Experience with AIOps - using AI/ML-based tooling for anomaly detection, predictive alerting, or automated incident triage
Familiarity with the reliability characteristics of AI/ML-backed services (e.g., LLM inference latency, non-determinism, prompt pipeline observability)
Experience with the JavaScript/Node.js ecosystem
Certified Kubernetes Administrator (CKA) certification
Familiarity with compliance frameworks like SOC 2, ISO 27001, or NIST AI
Demonstrated use of AI/AIOps capabilities for reliability tasks - anomaly detection, incident triage, runbook generation, or alert noise reduction
Demonstrated passion for AI through personal projects, contributions, or continuous learning in the context of infrastructure or reliability engineering
Shared Success
Stock equity to ensure that as the company grows, you share directly in that success
Up to 100% employer-paid premiums for medical, dental, and vision coverage for employees and their dependents, along with comprehensive wellness benefits and healthcare concierge services designed to support your needs beyond traditional insurance
A comprehensive suite of financial benefits, including a 401(k) plan, company-paid life and disability insurance, tax-advantaged spending accounts, and a range of discounted voluntary offerings to help you customize and strengthen your overall financial position
Paid Parental Leave policy, after six months of employment. Employees also receive access to Kindbody fertility and family-building benefits and dedicated leave specialists who help guide you through the entire process
Generous annual stipends for both professional and personal development, empowering you to invest in your continued growth. You’ll also have access to a wide range of internal learning opportunities, ensuring you can build new skills, deepen your expertise, and advance your career with confidence
A flexible vacation policy, paid holidays, and other perks to recharge
Compensation
The applicable salary range for this role is: $166,900 - $225,900. A variety of factors are considered when determining someone’s leveling and compensation–including a candidate’s professional background and experience. These ranges may be modified in the future and final offer amounts may vary from the amounts listed above.