Jobs · Management · California

Senior Manager, Production Engineering

CoreWeave · San Francisco, CA · 2 days ago
Management$207k–$275k/yrFull-time

About the role

We're scaling the CoreWeave Cloud Platform and need a Senior Manager of Production Engineering to lead our SRE team and practices. This role involves contributing to the SRE vision, strategy, and roadmap, leading a team, establishing best practices, and collaborating with various teams.

Responsibilities

  • Contribute and execute the SRE vision, strategy, and roadmap for a large-scale, distributed cloud infrastructure.
  • Lead and mentor a high-performing team of SREs, promoting a culture of ownership, collaboration, and continuous learning.
  • Champion automation-first practices, leveraging AI, tools like Terraform, Kubernetes, and Infrastructure-as-Code to minimize toil and manual interventions.
  • Establish and evolve Operational Excellence best practices ensuring the platform is proactive and propagates organizational learning.
  • Drive initiatives for incident management, postmortem culture, root cause analysis, and system hardening.
  • Collaborate with engineering, product, and customer support teams to build scalable, resilient, and self-healing systems.
  • Evolve our on-call strategy and processes to support a 24x7, globally distributed platform with minimal disruptions.

Requirements

  • Bachelor’s degrees in Computer Science, Engineering, or related fields.
  • 10+ years in a leadership or senior management role at a cloud provider, hyperscaler, or high-growth tech company.
  • Experience hiring, developing, and managing geographically distributed 24x7 engineering teams.
  • Experience in designing and implementing incident management processes including on-call rotations, escalation paths, postmortems, and SLO/SLA framework.
  • Strong foundation in systems engineering, with a deep understanding of distributed systems, networking, and storage architecture.
  • Strong cross-functional collaboration skills, with the ability to influence product, platform, hardware, and security teams.

Qualifications

  • Experience building platform tooling and internal developer portals to improve engineering velocity and operational visibility.
  • Familiarity with GPU-accelerated workloads, including resource isolation, scheduling, and performance tuning.
  • Experience working in AI infrastructure environments supporting training or inference at scale.
  • Experience in compliance, reliability risk modeling, or operational maturity assessments (e.g., RTO/RPO, chaos engineering).
  • Prior leadership experience in bare metal infrastructure environments (e.g., custom data centers, edge compute, HPC clusters).
  • Working knowledge of DPUs, service mesh architectures, and multi-tenant security models.

Benefits

  • The base salary range for this role is $207,000 to $275,000.
  • The starting salary will be determined based on job-related knowledge, skills, experience, and market location.
  • We offer a variety of benefits including medical, dental, and vision insurance, company-paid Life Insurance, Voluntary supplemental life insurance, Short and long-term disability insurance, Flexible Spending Account, Health Savings Account, Tuition Reimbursement, Ability to Participate in Employee Stock Purchase Program (ESPP), Mental Wellness Benefits through Spring Health, Family-Forming support provided by Carrot, Paid Parental Leave, Flexible, full-service childcare support with Kinside, 401(k) with a generous employer match, Flexible PTO, Catered lunch each day in our office and data center locations, A casual work environment, A work culture focused on innovative disruption, California Applicants California Consumer Privacy Act, Equal Opportunity & Accommodations, CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace, Export Control Compliance, This position requires access to export controlled information.

Similar jobs