Staff Software Engineer, Compute Architecture
CoreWeave · Bellevue, WA · 3 wk ago
Engineering$188k–$275k/yrFull-time
About the role
The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.
Responsibilities
- Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.
- Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.
- Translate incidents and hardware failure modes into software improvements that make the platform more resilient.
- Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale.
- Mentor engineers and lead technical projects.
- Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship.
- Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden.
Requirements
- B.S., M.S., or PhD in Computer Science or related field, or equivalent experience.
- 8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environments.
- Expertise in Go and proven experience building REST/gRPC APIs for mission-critical platforms.
- Strong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
- Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams.
- Experience contributing to and collaborating with open source communities.
- Skilled in applying a data-driven approach to reliability, optimization, and continuous improvement.
- Excellent communicator able to work effectively with both technical and non-technical stakeholders.
- Hands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers.
- Track record of leading incident response, postmortems, and driving robust service reliability.
Qualifications
- Working knowledge of Kafka, ClickHouse and CRDB.
- DMTF, RedFish APIs, and GPU servers.
Skills
- Nice to have: Working knowledge of Kafka, ClickHouse and CRDB.
- DMTF, RedFish APIs, and GPU servers.
Benefits
- The base salary range for this role is $188,000 to $275,000.
- In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).
Pay
The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation.
Schedule
Not specified.