Senior Manager, Production Engineering
CoreWeave · San Francisco, CA · 2 days ago
Management$207k–$275k/yrFull-time
About the role
We're scaling the CoreWeave Cloud Platform and need a Senior Manager of Production Engineering to lead our SRE team and practices. This role involves contributing to the SRE vision, strategy, and roadmap, leading a team, establishing best practices, and collaborating with various teams.
Responsibilities
- Contribute and execute the SRE vision, strategy, and roadmap for a large-scale, distributed cloud infrastructure.
- Lead and mentor a high-performing team of SREs, promoting a culture of ownership, collaboration, and continuous learning.
- Champion automation-first practices, leveraging AI, tools like Terraform, Kubernetes, and Infrastructure-as-Code to minimize toil and manual interventions.
- Establish and evolve Operational Excellence best practices ensuring the platform is proactive and propagates organizational learning.
- Drive initiatives for incident management, postmortem culture, root cause analysis, and system hardening.
- Collaborate with engineering, product, and customer support teams to build scalable, resilient, and self-healing systems.
- Evolve our on-call strategy and processes to support a 24x7, globally distributed platform with minimal disruptions.
Requirements
- Bachelor’s degrees in Computer Science, Engineering, or related fields.
- 10+ years in a leadership or senior management role at a cloud provider, hyperscaler, or high-growth tech company.
- Experience hiring, developing, and managing geographically distributed 24x7 engineering teams.
- Experience in designing and implementing incident management processes including on-call rotations, escalation paths, postmortems, and SLO/SLA framework.
- Strong foundation in systems engineering, with a deep understanding of distributed systems, networking, and storage architecture.
- Strong cross-functional collaboration skills, with the ability to influence product, platform, hardware, and security teams.
Qualifications
- Experience building platform tooling and internal developer portals to improve engineering velocity and operational visibility.
- Familiarity with GPU-accelerated workloads, including resource isolation, scheduling, and performance tuning.
- Experience working in AI infrastructure environments supporting training or inference at scale.
- Experience in compliance, reliability risk modeling, or operational maturity assessments (e.g., RTO/RPO, chaos engineering).
- Prior leadership experience in bare metal infrastructure environments (e.g., custom data centers, edge compute, HPC clusters).
- Working knowledge of DPUs, service mesh architectures, and multi-tenant security models.
Benefits
- The base salary range for this role is $207,000 to $275,000.
- The starting salary will be determined based on job-related knowledge, skills, experience, and market location.
- We offer a variety of benefits including medical, dental, and vision insurance, company-paid Life Insurance, Voluntary supplemental life insurance, Short and long-term disability insurance, Flexible Spending Account, Health Savings Account, Tuition Reimbursement, Ability to Participate in Employee Stock Purchase Program (ESPP), Mental Wellness Benefits through Spring Health, Family-Forming support provided by Carrot, Paid Parental Leave, Flexible, full-service childcare support with Kinside, 401(k) with a generous employer match, Flexible PTO, Catered lunch each day in our office and data center locations, A casual work environment, A work culture focused on innovative disruption, California Applicants California Consumer Privacy Act, Equal Opportunity & Accommodations, CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace, Export Control Compliance, This position requires access to export controlled information.