Senior Software Engineer - Perf and Benchmarking
CoreWeave · Bellevue, WA · 3 wk ago
Engineering$182k–$242k/yrFull-time
About the role
We're looking for a Senior Engineer for CoreWeave’s Benchmarking & Performance team. You will play a crucial role in our planet-scale performance data warehouse, developing and enhancing Kubernetes-native benchmarking services.
Responsibilities
- Develop and enhance Kubernetes-native benchmarking services measuring latency, throughput, jitter, and cost-per-request across CoreWeave’s compute stack.
- Contribute to implementing and maintaining benchmarking workflows for end-to-end MLPerf Training and Inference runs, including workload setup, cluster configuration, and result validation.
- Participate in design discussions and contribute to architecture decisions within the team.
- Break down engineering tasks into clear milestones and deliver reliable, high-quality code.
- Collaborate with teammates to maintain reproducible, well-documented benchmarking processes.
- Provide constructive code reviews and share best practices with peers.
- Mentor junior engineers; review cross-team designs and elevate coding/testing standards.
- Help ensure reproducible, well-documented benchmarking processes.
Requirements
- 3–5 years of experience building distributed systems, high-performance computing components, or cloud services.
- Strong programming skills in Python or Go (C++ a plus) with understanding of networked systems and performance fundamentals.
- Hands-on experience with Kubernetes in production environments plus familiarity with CI/CD and observability tools (e.g., Prometheus, Grafana, OpenTelemetry).
- Exposure to performance-critical GPU systems (CUDA, NCCL, NVLink/PCIe, memory bandwidth) or model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM).
- Effective communicator comfortable working cross-functionally.
Qualifications
- Nice to have: Experience with time-series databases, LSM-based storage engines, or custom data pipelines.
- Familiarity with MLPerf or other large-scale benchmarking frameworks.
- Contributions to OSS projects such as llm-d, vLLM or PyTorch.
- Exposure to benchmarking GPU clusters or multi-region environments.
- Background working with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies.
Skills
- Experience with Kubernetes in production environments.
- Understanding of CI/CD and observability tools.
- Knowledge of performance-critical GPU systems.
- Experience with MLPerf or similar benchmarking frameworks.
- Contributions to open-source projects.
- Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies.
Benefits
- The base salary range for this role is $182,000 to $242,000.
- The starting salary will be determined based on job-related knowledge, skills, experience, and market location.
- We offer a variety of benefits including medical, dental, and vision insurance, company-paid life insurance, short and long-term disability insurance, flexible spending account, health savings account, tuition reimbursement, ability to participate in employee stock purchase program (ESPP), mental wellness benefits through Spring Health, family-forming support provided by Carrot, paid parental leave, flexible, full-service childcare support with Kinside, 401(k) with a generous employer match, flexible PTO, catered lunch each day in our office and data center locations, a casual work environment, and a work culture focused on innovative disruption.
Pay
The base salary range for this role is $182,000 to $242,000.
Schedule
Not specified.