HPC Platform Engineer, Software, Center for Quantum Computing
About the Role
The Models, Quantum, & Silicon (MQS) Center for Quantum Computing (CQC) is a multi-disciplinary team of scientists, engineers, and technicians, on a mission to develop a fault-tolerant quantum computer. We are looking to hire an HPC Platform Engineer to develop, automate, and maintain high-performance computing (HPC) infrastructure on AWS that CQC scientists and engineers use for quantum computing hardware design and simulation. You will work closely with our experimental and theoretical physics teams to enable large-scale HPC workloads with MPI-based parallelism on EC2 instances, manage graphical environments for computer-aided engineering applications, and accelerate the research computing lifecycle through automation and infrastructure-as-code.
The ideal candidate will be able to translate high-level science and simulation requirements into reliable deliverables (including cluster orchestration, job scheduling, CI/CD pipelines, reproducible environments, and artifact management) that are performant, scalable, and secure. This requires someone who (1) has a strong desire to work within a team of scientists and engineers, (2) demonstrates ownership by initiating and driving projects to completion, and (3) is comfortable operating at the intersection of traditional HPC and cloud-native infrastructure.
Responsibilities
- Administer and automate cloud-based HPC environments by deploying and maintaining clusters, managing OS and software stacks, building containers, provisioning users, and securing systems.
- Support computer-aided engineering and computational science workflows (e.g., Palace) by improving HPC environment robustness, performance, and availability across instance types and regions.
- Anticipate and expand CQC computational capacity, leveraging AWS to accelerate the quantum hardware design cycle.
- Design reproducible development environments and own dependency, build, and release management across interdependent research software projects.
- Develop CI/CD pipelines and automate provisioning using infrastructure-as-code (e.g., Docker, AWS CDK).
- Maintain Amazon's high security bar while enabling fast-paced development; implement observability and monitoring for rapid debugging and high uptime.
Basic Qualifications
- Experience in automating, deploying, and supporting large-scale infrastructure
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
- Experience with Linux/Unix
- Experience with CI/CD pipelines build processes
- 2+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- Experience using infrastructure-as-code to design and deploy cloud services
Preferred Qualifications
- Experience with distributed systems at scale
- Experience in an AWS environment, including VPC, EC2, EBS, S3, SQS, CloudFormation and Lambda
- Experience in network fundamentals (DNS, DHCP, TCP/IP, routing, switching, HTTP)
- Experience in Kubernetes, Docker or containers ecosystem
- Experience working with scientists in a research environment
- Experience with high-performance computing (HPC) infrastructure
Pay
USA, CA, San Francisco - 148,700.00 - 201,200.00 USD annually. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location.
Benefits
Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
This response is AI-generated, for reference only.