Site Reliability Engineer (Application Software)
SpaceX · Hawthorne, CA · 1 mo ago
On-siteEngineering$125k–$160k/yrFull-time
Responsibilities
- Deploy, upgrade, operate, maintain, and scale our suite of mission-critical products and services
- Manage our underlying infrastructure as code and use modern observability tools to provide a complete picture of application health
- Closely collaborate with software engineers to design and build highly operable, maintainable, and testable systems
- Engage in and improve the entire software development lifecycle — from inception and design through deployment, operation, and continuous refinement
- Practice sustainable incident response and blameless postmortems
- Provide high-quality end-user support to vehicle software engineers
- Participate in the team’s on-call rotation
- Identify and eliminate performance bottlenecks using measurement and creative engineering
Requirements
- Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree
- 1+ years of experience with Python and Python-based development frameworks
- Experience with Linux operating systems
Basic Qualifications
- Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree
- 1+ years of experience with Python and Python-based development frameworks
- Experience with Linux operating systems
Preferred Skills And Experience
- Experience with build systems (Bazel, Buck, Make, etc.)
- Experience with both container and virtualization technologies (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)
- Experience with databases and data modeling (Postgres, MySQL, ClickHouse, etc.)
- Experience with infrastructure as code (IaC) tools for managing fleets of servers
- Experience with Terraform, Ansible, Puppet, or similar automation frameworks
- Knowledge of the technologies that predate and underpin modern cloud infrastructure, with the ability to translate high-level developer experiences into specific implementations from first principles
- Able to work with mission-critical and sensitive systems with appropriate urgency and care
- Able to communicate effectively with customers, peers, and management in both formal and informal settings
- Experience with full-stack development (the team primarily uses Python, JavaScript, and C#; end users primarily use C#)