Senior Manager, Site Reliability Engineering
H1BConnect · Santa Clara, CA · 6 days ago
Quality Assurance$248k–$397k/yrFull-time
Location: Santa Clara, CA • Full-time
About the role
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead the transformation of IT operations at scale. This role focuses on building AI-powered systems to enhance reliability and employee experience, moving from reactive to predictive operations. The ideal candidate will manage incident, problem, and change management processes, leveraging observability and automation. This position offers an opportunity to lead a high-performing team in a diverse and innovative environment.
Requirements
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
- 5+ years of experience leading and managing global IT operations or service management teams.
- 12+ overall years of experience in Site Reliability Engineering or IT Service Management.
- Proven proficiency in Incident, Problem, and Configuration Management.
- Demonstrated experience applying AI, automation, or advanced analytics to improve operational outcomes.
- Solid understanding of observability and modern reliability practices.
- Strong leadership capability with experience building and scaling engineering-focused teams.
- Ability to deliver executive-level communication and insights.
Responsibilities
- Manage the full lifecycle of Incident, Problem, and Change Management as a 24×7 operational function.
- Transform incident response by implementing AI detection and guided remediation.
- Build and scale intelligent incident workflows that integrate monitoring and telemetry.
- Evolve Problem Management into a data-driven field using AI and analytics.
- Modernize Change Management by introducing risk-aware, data-driven decision-making.
- Drive the adoption of observability to ensure service-level visibility and actionable insights.
- Lead the development of automation and orchestration platforms to reduce manual effort.
- Partner closely with engineering, infrastructure, and business teams to align operations with service reliability goals.
Benefits
- Comprehensive day-one benefits including medical, dental, and vision coverage with HSA support.
- Life and disability insurance, Employee Assistance Program, and 401(k) with auto-enrollment.
- Generous time off and holidays.
- Donation matching up to $10,000.
- Flexible Spending Accounts (FSAs), commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts.
- Optional programs including student-loan and home-purchase support, family care resources, and expert medical services.
Pay
$248,000 - $396,750 per year