Sr. High Performance Computing (HPC) Systems Engineer
SpaceX · Texas, United States · 6 days ago
On-siteInformation TechnologyFull-time
About the role
SpaceX is looking for a Senior High Performance Computing (HPC) Systems Engineer to join the HPC team and support SpaceX personnel and proprietary systems. The ideal candidate will be flexible and flourish in a fast-paced and challenging environment, demonstrating self-motivation and ingenuity.
Responsibilities
- Administer and manage HPC clusters, storage systems, and high-speed networks
- Provide application support to SpaceX employees across engineering disciplines
- Install and integrate Linux-based compute clusters
- Write instructional documentation and convey highly technical ideas in non-technical terms
Requirements
- Bachelor's degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree
- 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
- Experience with Kubernetes
- Must be willing to work extended hours and weekends as needed
- Eligibility for access to classified material up to TS/SCI with Polygraph
- To conform to U.S. Government export regulations, applicant must be a U.S. citizen or national, U.S. lawful permanent resident (green card holder), Refugee under 8 U.S.C.1157, or Asylee under 8 U.S.C.1158, or be eligible to obtain the required authorizations from the U.S. Department of State
Skills
- 5+ years of professional experience building, deploying, and troubleshooting Linux systems
- Experience with a scripting language (Bash, Python) to automate and solve recurring tasks
- Experience building, deploying, and troubleshooting HPC clusters
- Familiarity with cluster resource managers (Slurm, PBS, LSF)
- Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
- Familiarity with scientific and engineering computing (CFD, FEA)
- Familiarity with large-scale AI training
- Familiarity with GPU usage in a compute cluster and CUDA
- Experience with containers (Docker, Podman, Singularity)
- Experience deploying and maintaining automated configuration management software (Puppet, Ansible)
- Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities