Jobs · Engineering · California

Lead Site Reliability Engineer, Platforms

Zoom · San Jose, CA · 1 wk ago
Engineering$124k/yrFull-time

Minimum Salary Range: $124,000 – $271,200. Total direct compensation includes base salary, bonus, and equity value, and may vary by location.

About the role

As a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will act as an uber tech lead, defining projects and guiding work across various teams with a wide scope of responsibility. You will have the opportunity to improve our datacenter Kubernetes infrastructure, cloud infrastructure, security posture, and operation of ZfG environments. Broadly, you will guide all teams toward SRE best practices such as automation, monitoring, and infrastructure as code.

About the team

The DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG). The team is currently working on hardening security posture and building automation and reliability systems that underpin Zoom's global services. This role offers broad visibility, cross-team influence, and the chance to define how infrastructure is built.

Responsibilities

  • Design and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environments
  • Define technical roadmaps and architectural direction for infrastructure automation and security
  • Partner with service teams to understand platform needs and deliver solutions that improve reliability and efficiency
  • Establish and advocate for SRE best practices including infrastructure as code, monitoring, and incident management
  • Mentor team members through design, implementation, and production deployment of complex systems

Requirements

  • 8+ years of SRE or DevOps experience building and operating production infrastructure at scale
  • Proficiency in at least one programming language beyond scripting (e.g., Python, Go, Java)
  • Experience deploying and managing CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrog
  • Experience operating cloud infrastructure on AWS, OCI, or similar platforms using Terraform and Kubernetes
  • Experience implementing observability solutions with logging and monitoring tools such as ELK, Prometheus, or Grafana
  • Ability to communicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditors
  • Participation in on-call rotations and leadership in incident response to maintain system reliability
  • Degree in Computer Science or related field, or equivalent practical experience
  • US citizenship or Green Card status

Preferred Qualifications

  • Experience with security from an SRE perspective
  • Experience with Identity security (e.g., IAM, workload identity, zero trust) and tools (e.g., Teleport, Okta)
  • Experience operating Government environments and understanding their compliance requirements
  • Experience with system design and distributed computing at scale
  • Ability to speak Chinese/Mandarin is a plus

Benefits

As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways.

Schedule

This role follows a structured hybrid approach, centered around our offices and remote work environments.

Similar jobs