Jobs · Engineering · California

Senior Specialist Field Engineer - Compute Infrastructure

CoreWeave · Sunnyvale, CA · 3 wk ago
Engineering$188k–$275k/yrFull-time

About the role

The Field Engineering organization at CoreWeave is dedicated to ensuring every customer running AI workloads at scale has a seamless, reliable, and high-performance experience. This team supports the infrastructure that powers the AI revolution—working across data centers, hardware systems, and customer workloads to maintain the integrity of our cloud platform. Field Engineering aligns closely with internal and customer engineering teams, offering valuable insights from the field and the chance to shape the CoreWeave product roadmap and development.

Responsibilities

  • Own the technical path from facility and rack design to a validated, production-ready supercomputer—spanning logical design, infrastructure engineering, provisioning, validation, operations, and support.
  • Lead bring-up and acceptance of new large-scale GPU clusters, driving InfiniBand/RoCE fabric validation, HPC performance benchmarking (e.g., NCCL, ib_write_bw), and remediation of fabric, optics, firmware, and node-level issues to meet customer performance targets.
  • Define and operationalize models for managing customer bare-metal fleets at rack-level-and-up—IT service, break-fix, network and firmware management—including Bare Metal as a Service (BMaaS) and customer self-service patterns.
  • Partner with Data Center Operations, Fleet Operations, and Networking teams to align facility, hardware, and fabric readiness with customer go-live timelines and operational SLAs.
  • Review and advise on customer-facing technical contract terms, including service scope, operational responsibilities, SLAs, isolation requirements, and support boundaries.
  • Drive technical leadership and direction during customer meetings, presentations, and workshops, addressing any technical queries or concerns that arise.
  • Offer valuable insights on product features, functionality, and performance, contributing regularly to discussions about product strategy and architecture.
  • Stay informed of the latest developments and trends in Kubernetes, cloud computing and infrastructure, sharing your thought leadership with customers and internal stakeholders.
  • Lead the prototyping and initiation of research and development efforts for emerging products and solutions, delivering prototypes and key insights for internal consumption.
  • Represent CoreWeave at conferences and industry events, with occasional travel as required.

Requirements

B.S. in Computer Science or a related technical discipline, or equivalent experience
7+ years of proven experience as a Solutions Architect, Field Engineer, Infrastructure/Systems Engineer, or Technical Account Manager in Cloud Infrastructure, focusing on building or operating distributed systems or HPC/cloud services, with an expertise focused on bare-metal compute infrastructure and large-scale GPU cluster delivery
Fluency in cloud computing concepts, architecture, and technologies with hands-on experience in designing and implementing cloud solutions
Proven track record with building customer relationships, communicating clearly and the ability to break down complex technical concepts to both technical and non-technical audiences
Deep expertise with modern rack-scale GPU server hardware (e.g., NVIDIA HGX / GB200-class systems), high-speed interconnects (InfiniBand, NVLink), and the firmware/BMC/BIOS layer
Expert-level Linux system administration and command-line troubleshooting, paired with strong networking fundamentals (routing, fabric topologies, TCP/IP)
Hands-on experience bringing up, validating, and operating large GPU clusters—including bare metal node pxe boot, hardware health, fabric validation, and HPC acceptance/performance testing—and integrating bare metal with orchestration layers such as Kubernetes and Slurm
Preferred Experience operating security-sensitive, air-gapped, or otherwise locked-down customer environments
Experience with scripting and automation related to bare-metal provisioning, infrastructure validation, and lifecycle management (Python, Bash, Ansible, or similar)
Experience designing AI supercomputers from MEP designs
Experience delivering bare-metal infrastructure at scale for large strategic customers or AI research labs
Experience with building solutions across multi-cloud or hybrid environment

Qualifications

Fluency in cloud computing concepts, architecture, and technologies with hands-on experience in designing and implementing cloud solutions
Proven track record with building customer relationships, communicating clearly and the ability to break down complex technical concepts to both technical and non-technical audiences
Deep expertise with modern rack-scale GPU server hardware (e.g., NVIDIA HGX / GB200-class systems), high-speed interconnects (InfiniBand, NVLink), and the firmware/BMC/BIOS layer
Expert-level Linux system administration and command-line troubleshooting, paired with strong networking fundamentals (routing, fabric topologies, TCP/IP)
Hands-on experience bringing up, validating, and operating large GPU clusters—including bare metal node pxe boot, hardware health, fabric validation, and HPC acceptance/performance testing—and integrating bare metal with orchestration layers such as Kubernetes and Slurm
Preferred Experience operating security-sensitive, air-gapped, or otherwise locked-down customer environments
Experience with scripting and automation related to bare-metal provisioning, infrastructure validation, and lifecycle management (Python, Bash, Ansible, or similar)
Experience designing AI supercomputers from MEP designs
Experience delivering bare-metal infrastructure at scale for large strategic customers or AI research labs
Experience with building solutions across multi-cloud or hybrid environment

Skills

Fluency in cloud computing concepts, architecture, and technologies with hands-on experience in designing and implementing cloud solutions
Proven track record with building customer relationships, communicating clearly and the ability to break down complex technical concepts to both technical and non-technical audiences
Deep expertise with modern rack-scale GPU server hardware (e.g., NVIDIA HGX / GB200-class systems), high-speed interconnects (InfiniBand, NVLink), and the firmware/BMC/BIOS layer
Expert-level Linux system administration and command-line troubleshooting, paired with strong networking fundamentals (routing, fabric topologies, TCP/IP)
Hands-on experience bringing up, validating, and operating large GPU clusters—including bare metal node pxe boot, hardware health, fabric validation, and HPC acceptance/performance testing—and integrating bare metal with orchestration layers such as Kubernetes and Slurm
Preferred Experience operating security-sensitive, air-gapped, or otherwise locked-down customer environments
Experience with scripting and automation related to bare-metal provisioning, infrastructure validation, and lifecycle management (Python, Bash, Ansible, or similar)
Experience designing AI supercomputers from MEP designs
Experience delivering bare-metal infrastructure at scale for large strategic customers or AI research labs
Experience with building solutions across multi-cloud or hybrid environment

Benefits

In addition to a competitive salary, we offer a variety of benefits to support your needs. These include: Medical, dental, and vision insurance - 100% paid for by CoreWeave, Company-paid Life Insurance, Voluntary supplemental life insurance, Short and long-term disability insurance, Flexible Spending Account, Health Savings Account, Tuition Reimbursement, Ability to Participate in Employee Stock Purchase Program (ESPP), Mental Wellness Benefits through Spring Health, Family-Forming support provided by Carrot, Paid Parental Leave, Flexible, full-service childcare support with Kinside, 401(k) with a generous employer match, Flexible PTO, Catered lunch each day in our office and data center locations, A casual work environment, A work culture focused on innovative disruption, California Applicants California Consumer Privacy Act, Equal Opportunity & Accommodations, CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: careers@coreweave.com. Export Control Compliance This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C.1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.

Similar jobs