Staff DevOps Engineer
Immigration sponsorship is not available for this position.
What You Can Expect
We are hiring a Staff DevOps/Site Reliability Engineer to ensure reliability, scalability, and operational excellence for our real-time communications platform. This platform supports audio/video conferencing, recording, and live-streaming functionalities. The position requires expertise in infrastructure engineering, global team collaboration, and cross-functional partnerships.
About The Team
This team manages essential meeting service operations at Zoom. They handle global, large-scale distributed systems and advance communication technology to connect individuals across physical distances.
Responsibilities
- Ensuring reliability engineering and operations by owning the SLO/SLI framework for real-time services, defining, tracking, and improving latency, availability, jitter, and packet loss.
- Leading incident response for critical outages across the real-time platform, coordinating across time zones and engineering disciplines.
- Promoting a blameless postmortem culture and ensuring action items lead to measurable reliability enhancements.
- Implementing chaos engineering and game day exercises to proactively identify failure modes before user impact occurs.
- Building and evolving observability tools — dashboards, alerting systems, and distributed tracing — tailored for real-time media infrastructure challenges.
- Serving as the architectural authority on deployment patterns, infrastructure design, and operational readiness for real-time services.
- Reviewing and contributing to system design proposals, providing feedback on scalability, fault tolerance, and operational complexity.
- Driving capacity planning, traffic modeling, and cost optimization strategies across globally distributed infrastructure.
- Evaluating and recommending infrastructure tools, platforms, and vendors — including media servers, CDN providers, cloud-native services, and edge networking.
- Ensuring consistent standards for CI/CD pipelines, deployment safety, and progressive rollout strategies across teams.
- Acting as the primary SRE partner for multiple engineering teams building real-time features, attending planning sessions, and providing operational readiness guidance.
- Collaborating closely with network engineering, security, product, and data teams to align on platform-wide reliability requirements.
- Translating infrastructure constraints and reliability trade-offs into actionable recommendations for product leaders and engineering teams.
- Establishing and advocating DevOps best practices — infrastructure-as-code, GitOps, automated testing, and deployment automation — across partner teams.
- Guiding senior engineers on SRE principles, reliability patterns, and operational discipline.
- Serving as a technical liaison between US-based and China/India-based engineering teams, bridging communication gaps and providing technical context.
- Conducting architecture reviews, incident retrospectives, and planning sessions in English and Mandarin as appropriate.
- Maintaining a flexible schedule to ensure meaningful overlap with teams in Beijing, Shanghai, Bangalore, and Hyderabad.
- Building collaborative relationships across cultural and geographic boundaries, adapting communication styles to foster trust and alignment.
- Ensuring engineering documentation, runbooks, and architectural decision records are accessible and understandable for global team members.
Requirements
- 10+ years in DevOps, SRE, or infrastructure engineering roles, with at least 3 years at a staff or principal level scope.
- Proven track record owning reliability for large-scale, distributed, latency-sensitive systems in production.
- Experience in supporting real-time or media-heavy platforms (video conferencing, live streaming, gaming, trading systems, or similar).
- Ability to lead cross-functional technical initiatives without direct authority, driving alignment across engineering, product, and operations.
- Conceptual and architectural understanding of real-time communication protocols: WebRTC, RTP/RTCP, TURN/STUN, SDP, and SFU/MCU topologies.
- Solid expertise in cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Kubernetes, Helm, ArgoCD).
- Proficiency with infrastructure-as-code tooling: Terraform, Pulumi, or equivalent.
- Experience with observability stacks: Prometheus, Grafana, Datadog, Jaeger, OpenTelemetry, or equivalent.
- Understanding of networking fundamentals: BGP, anycast routing, DNS, load balancing, and CDN architecture.
- Utilize CI/CD tools such as GitHub Actions, Jenkins, and Spinnaker to streamline workflows and improve deployment processes.
- Implement deployment safety practices like canary releases, feature flags, and blue/green strategies to ensure reliable software delivery.
- Proficiency in Python, Bash, or Go for automation, tooling, and incident response without requiring advanced software development expertise.
- Occasional weekend work may be required.
- Ability to work across the globe or multiple time zones.
Pay
Minimum Salary Range or On Target Earnings: $124,000.00 - $271,200.00. In addition to the base salary, Zoom has a Total Direct Compensation philosophy that includes base salary, bonus, and equity value. Starting pay is based on qualifications, experience, and location.
Schedule
This role follows a structured hybrid approach, centered around offices and remote work environments. The work style for this position is hybrid.
Benefits
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, including options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Learn more.