Sr. Software Engineer - Perf and Benchmarking
CoreWeave · Bellevue, WA · 3 wk ago
Engineering$139k–$204k/yrFull-time
About the role
We're looking for a Senior Engineer for CoreWeave’s Benchmarking & Performance team. You will play a crucial role in our planet-scale performance data warehouse, measuring latency, throughput, jitter, and cost-per-request across our global infrastructure.
Responsibilities
- Build and improve Kubernetes-native benchmarking services that measure latency, throughput, jitter, and cost-per-request across CoreWeave’s compute stack.
- Implement and maintain benchmarking workflows for end-to-end MLPerf Training and Inference runs, including workload setup, cluster configuration, runbooks, and result validation.
- Lead design reviews and drive architecture within the team; decompose multi-service work into clear milestones.
- Mentor junior engineers; review cross-team designs and elevate coding/testing standards.
- Help ensure reproducible, well-documented benchmarking processes.
Requirements
- 5+ years of experience building distributed systems, high-performance computing, or cloud services.
- Strong coding in Python or Go (C++ a plus) and deep familiarity with networked systems and performance.
- Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry).
- Experience with performance-critical GPU systems (CUDA, NCCL, RDMA, NVLink/PCIe, memory bandwidth) and model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM).
- Strong communicator comfortable collaborating with cross-functional teams and external partners.
Qualifications
- Nice to have: Experience with time-series databases, LSM-based storage engines, or custom data pipelines.
- Experience running MLPerf submissions or similar large-scale audited benchmarks.
- Contributions to OSS projects such as llm-d, vLLM or PyTorch.
- Exposure to benchmarking large GPU fleets or multi-region clusters.
- Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies.
Skills
- Experience with Kubernetes at production scale.
- Knowledge of networked systems and performance.
- Ability to mentor junior engineers.
- Experience with CI/CD and observability stacks.
- Experience with performance-critical GPU systems.
- Experience with time-series databases, LSM-based storage engines, or custom data pipelines.
- Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies.
Benefits
- The base salary range for this role is $139,000 to $204,000.
- Discretionary bonus.
- Equity awards.
- A comprehensive benefits program (all based on eligibility).
Pay
The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation.
Schedule
Our work schedule is flexible and accommodates your needs.