Staff Reliability Engineer
About the role
We are seeking an exceptional Staff Site Reliability Engineer to lead critical infrastructure initiatives and drive innovation across our organization. You'll architect scalable solutions, navigate complex technical challenges independently, and deliver results under tight deadlines in a fast-paced environment.
Responsibilities
- Lead enterprise-wide reliability and infrastructure projects across multiple teams with high autonomy
- Navigate ambiguous problem spaces and deliver innovative solutions under tight deadlines
- Architect and deploy solutions for Cloud Prem and SaaS customers at scale
- Drive technical innovation and establish SRE best practices across the organization
- Respond to critical incidents, lead root cause analysis, and implement long-term resolutions
- Develop automation solutions to streamline operations and reduce manual workload
- Participate in on-call rotation and ensure effective incident handoff and documentation
- Lead compliance-focused infrastructure initiatives and partner with Security teams on control implementation across cloud environments
- Provide tier 2/3 technical support to enterprise customers for complex troubleshooting
- Work directly with customer technical teams to resolve deployment, configuration, and integration challenges
- Create customer-facing documentation, troubleshooting guides, and run-books
- Lead customer calls and technical discussions as a trusted advisor
Qualifications
- US Citizenship Required — role requires or may require federal security clearance
- BS degree in Computer Science or related field (or equivalent practical experience)
- 7+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering
- Proven track record leading large-scale, cross-team infrastructure projects from conception to production
- Demonstrated ability to work autonomously on ambiguous projects with tight deadlines
- Hands-on experience with major compliance frameworks (SOC 1/2, ISO 27001, FedRAMP Moderate/High) including infrastructure automation for control implementation and continuous audit evidence collection
- Technical Expertise: 5+ years with AWS (VPC, EC2, RDS, EKS, CloudFormation) and cloud automation
- Expert-level experience with Kubernetes, Helm, Linux, and Terraform
- Strong experience with GitOps model, distributed version control, and CI/CD pipelines
- Proficiency with monitoring tools (Prometheus, Grafana, DataDog)
- Strong programming/scripting skills (Python, Go, Bash) for automation and reading codes
- Deep understanding of distributed systems, microservices, and reliability patterns
- Experience with Bazel and CueLang a plus
- Experience implementing compliance-as-code patterns and automated compliance scanning/remediation
- Leadership & Communication: Exceptional ability to articulate complex technical concepts to diverse audiences
- Track record of driving technical change across organizational boundaries
- Successfully delivered multiple complex projects under tight deadlines
- Strong customer service orientation with patience and empathy
What sets you apart
You don't wait for perfect information—you make calculated decisions and drive progress even when requirements are unclear. You've successfully navigated organizational complexity to deliver critical projects on aggressive timelines while building strong relationships across teams and with customers. You view incidents as learning opportunities and naturally raise the bar for technical excellence wherever you go.
Pay
$148,900 - $260,600, plus equity (when applicable), variable/incentive compensation and benefits.
Schedule
Flexible work arrangements available.