Jobs · Engineering · Washington

Principal Engineer - Perf and Benchmarking

CoreWeave · Bellevue, WA · 3 wk ago
Engineering$206k–$333k/yrFull-time

About the role

We're looking for a Principal Engineer to lead CoreWeave’s Benchmarking & Performance team. Your responsibilities include defining the benchmarking strategy, leading MLPerf submissions, designing and maintaining benchmarking tools, and collaborating with key partners.

Responsibilities

  • Define the multi-year benchmarking strategy and roadmap; prioritize models/workloads and hardware tiers.
  • Build, lead, and mentor a high-performing team of performance engineers and data analysts.
  • Establish governance for claims: documented methodologies, versioning, reproducibility, and audit trails.
  • Lead end-to-end MLPerf Inference and Training submissions: workload selection, cluster planning, runbooks, audits, and result publication.
  • Coordinate optimization tracks with NVIDIA (CUDA, cuDNN, TensorRT/TensorRT-LLM, Triton, NCCL) to hit competitive results; drive upstream fixes where needed.
  • Design a Kubernetes-native, repeatable benchmarking service that exercises CoreWeave stacks across SUNK, Kueue, and Kubeflow pipelines.
  • Measure and report p50/p95/p99 latency, jitter, tokens/s, time-to-first-token, cold-start/warm-start, and cost-per-token/request across models, precisions, batch sizes, and GPU types.
  • Maintain a corpus of representative scenarios and data sets; automate comparisons across software releases and hardware generations.
  • Build CI/CD pipelines and K8s controllers/operators to schedule benchmarks at scale; integrate with observability stacks and results warehouses.
  • Implement supply-chain integrity for benchmark artifacts (SBOMs, Cosign signatures).
  • Partner with NVIDIA, key ISVs, and OSS projects to co-develop optimizations and upstream improvements.
  • Support Sales/SEs with authoritative numbers for RFPs and competitive evaluations; brief analysts and press with rigorous, defensible data.

Requirements

  • 10+ years building distributed systems or HPC/cloud services, with deep expertise on large-scale ML training or similar high-performance workloads.
  • Proven track record of architecting or building planet-scale data systems (e.g., telemetry platforms, observability stacks, cloud data warehouses, large-scale OLAP engines).
  • Deep understanding of GPU performance (CUDA, NCCL, RDMA, NVLink/PCIe, memory bandwidth), model-server stacks (Triton, vLLM, TensorRT-LLM, TorchServe), and distributed training frameworks (PyTorch FSDP/DeepSpeed/Megatron-LM).
  • Proficient with Kubernetes and ML control planes; familiarity with SUNK, Kueue, and Kubeflow in production environments.
  • Excellent communicator able to interface with executives, customers, auditors, and OSS communities.

Qualifications

  • Nice to have: Experience with time-series databases, log-structured merge trees (LSM), or custom storage engine development.
  • Experience running MLPerf submissions (Inference and/or Training) or equivalent audited benchmarks at scale.
  • Contributions to MLPerf, Triton, vLLM, PyTorch, KServe, or similar OSS projects.
  • Experience benchmarking multi-region fleets and large clusters (thousands of GPUs).
  • Publications/talks on ML performance, latency engineering, or large-scale benchmarking methodology.

Skills

  • Strong analytical and problem-solving skills.
  • Ability to manage multiple projects simultaneously and meet deadlines.
  • Excellent communication and interpersonal skills.
  • Knowledge of industry standards and best practices in benchmarking and performance testing.

Benefits

The base salary range for this role is $206,000 to $333,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

What We Offer

  • Medical, dental, and vision insurance - 100% paid by CoreWeave.
  • Company-paid Life Insurance.
  • Voluntary supplemental life insurance.
  • Short and long-term disability insurance.
  • Flexible Spending Account (FSA) and Health Savings Account (HSA).
  • Tuition Reimbursement.
  • Absolutely free lunch each day in our office and data center locations.
  • A casual work environment.
  • A work culture focused on innovative disruption.
  • California Applicants: California Consumer Privacy Act.
  • Equal Opportunity & Accommodations.
  • Export Control Compliance.

Similar jobs