Jobs · Information Technology · Texas

Sr. High Performance Computing (HPC) Systems Engineer

SpaceX · Texas, United States · 6 days ago
On-siteInformation TechnologyFull-time

About the role

SpaceX is looking for a Senior High Performance Computing (HPC) Systems Engineer to join the HPC team and support SpaceX personnel and proprietary systems. The ideal candidate will be flexible and flourish in a fast-paced and challenging environment, demonstrating self-motivation and ingenuity.

Responsibilities

  • Administer and manage HPC clusters, storage systems, and high-speed networks
  • Provide application support to SpaceX employees across engineering disciplines
  • Install and integrate Linux-based compute clusters
  • Write instructional documentation and convey highly technical ideas in non-technical terms

Requirements

  • Bachelor's degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree
  • 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
  • Experience with Kubernetes
  • Must be willing to work extended hours and weekends as needed
  • Eligibility for access to classified material up to TS/SCI with Polygraph
  • To conform to U.S. Government export regulations, applicant must be a U.S. citizen or national, U.S. lawful permanent resident (green card holder), Refugee under 8 U.S.C.1157, or Asylee under 8 U.S.C.1158, or be eligible to obtain the required authorizations from the U.S. Department of State

Skills

  • 5+ years of professional experience building, deploying, and troubleshooting Linux systems
  • Experience with a scripting language (Bash, Python) to automate and solve recurring tasks
  • Experience building, deploying, and troubleshooting HPC clusters
  • Familiarity with cluster resource managers (Slurm, PBS, LSF)
  • Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
  • Familiarity with scientific and engineering computing (CFD, FEA)
  • Familiarity with large-scale AI training
  • Familiarity with GPU usage in a compute cluster and CUDA
  • Experience with containers (Docker, Podman, Singularity)
  • Experience deploying and maintaining automated configuration management software (Puppet, Ansible)
  • Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities

Similar jobs