Staff Site Reliability Engineer
Jobgether · United States · 6 days ago
RemoteRemoteEngineering$165k–$230k/yrFull-time
About the role
The Staff Site Reliability Engineer will define the technical vision and reliability strategy for a mission-critical cybersecurity platform operating at global scale. This role combines advanced software engineering, infrastructure architecture, and DevOps leadership to build resilient, secure, and scalable systems.
Responsibilities
- Define and execute the long-term infrastructure strategy supporting secure, repeatable deployments across hosted environments, customer-managed hardware, and restricted environments.
- Architect resilient distributed systems and deployment frameworks that support global-scale operations and diverse customer requirements.
- Lead the evolution of CI/CD platforms, Kubernetes environments, and configuration management practices to improve developer productivity and deployment reliability.
- Design advanced application packaging and infrastructure automation solutions using technologies such as Jsonnet, Grafana Tanka, Kustomize, and GitOps methodologies.
- Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance reliability and feature delivery.
- Develop enterprise observability strategies using monitoring, alerting, and tracing solutions to improve visibility into system health and performance.
- Drive infrastructure security architecture, including container security, zero-trust principles, network segmentation, and automated compliance practices.
- Partner with engineering teams to establish self-service tooling, operational standards, and reliable development workflows.
- Act as a technical advisor across engineering teams, reducing operational complexity and improving engineering effectiveness.
- Lead incident response efforts for high-severity issues, including incident command, root cause analysis, and long-term architectural improvements.
- Conduct blameless post-mortems and implement systemic solutions to prevent recurring failures.
- Mentor engineers and promote a culture of reliability, automation, continuous improvement, and technical excellence.
- Influence technical strategy across the organization by collaborating with engineering and product leadership.
Requirements
- Highly experienced Site Reliability or Platform Engineer with strong software engineering capabilities and a proven ability to lead large-scale infrastructure initiatives.
- 8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related infrastructure roles.
- Proven experience operating at a Staff, Principal, or Lead engineering level with ownership of organization-wide technical initiatives.
- Strong software engineering background with the ability to design and build production-quality systems and internal tooling.
- Advanced experience with Kubernetes in multi-cluster and production environments.
- Extensive experience designing CI/CD pipelines, GitOps workflows, and infrastructure-as-code solutions using tools such as GitHub Actions and ArgoCD.
- Experience architecting complex deployment models across self-hosted, on-premises, VMware-based, and air-gapped environments.
- Strong expertise in observability platforms, monitoring strategies, alerting frameworks, and distributed tracing.
- Experience designing secure infrastructure architectures, including container hardening, vulnerability management, and network security practices.
- Exceptional communication and stakeholder management skills, with the ability to influence technical decisions and align teams around reliability goals.
- Strong mentoring skills and a passion for improving engineering practices across teams.
Benefits
- Competitive base salary range of $165,000-$230,000.
- Eligibility for additional bonuses based on company performance and individual contributions.
- Comprehensive medical, dental, and vision benefits starting on day one.
- Health savings and financial wellness programs.
- 401(k) retirement savings plan with company matching.
- Unlimited vacation policy and dedicated health and wellness days.
- Paid parental leave programs.
- Opportunity to work on impactful cybersecurity technology in a collaborative and growth-focused environment.