Jobs · Washington

Staff Software Engineer - Compute

Lambda · Bellevue, WA · 1 wk ago
Hybrid$314k–$465k/yrFull-time

About The Role

As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration. The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.

Responsibilities

  • Designing and implementing a highly available and reliable GPU and CPU "host and instance lifecycle" control plane
  • Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability
  • Guide design of compute platform multi-tenant security model
  • Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy
  • Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities
  • Work with customers on translating vague customer technical requirements into concrete engineering deliverables
  • Set engineering standards and lead design reviews for mission-critical cloud software at scale

Qualifications

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale
  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning
  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems
  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives
  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers)
  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go

Nice to Have

  • Knowledge of Nvidia's AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs), and switches)
  • Knowledge of Nvidia's AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK)
  • Experience with Cloud Service Provider Kubernetes offerings
  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF)

Compensation

Annual salary range: $314K - $465K

Similar jobs