Jobs · Engineering · Washington

Staff Software Engineer, Compute Architecture

CoreWeave · Bellevue, WA · 3 wk ago
Engineering$188k–$275k/yrFull-time

About the role

The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.

Responsibilities

  • Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.
  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.
  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient.
  • Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale.
  • Mentor engineers and lead technical projects.
  • Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship.
  • Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden.

Requirements

  • B.S., M.S., or PhD in Computer Science or related field, or equivalent experience.
  • 8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environments.
  • Expertise in Go and proven experience building REST/gRPC APIs for mission-critical platforms.
  • Strong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
  • Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams.
  • Experience contributing to and collaborating with open source communities.
  • Skilled in applying a data-driven approach to reliability, optimization, and continuous improvement.
  • Excellent communicator able to work effectively with both technical and non-technical stakeholders.
  • Hands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers.
  • Track record of leading incident response, postmortems, and driving robust service reliability.

Qualifications

  • Working knowledge of Kafka, ClickHouse and CRDB.
  • DMTF, RedFish APIs, and GPU servers.

Skills

  • Nice to have: Working knowledge of Kafka, ClickHouse and CRDB.
  • DMTF, RedFish APIs, and GPU servers.

Benefits

  • The base salary range for this role is $188,000 to $275,000.
  • In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Pay

The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation.

Schedule

Not specified.

Similar jobs