Staff Site Reliability Engineer, Core AI Infrastructure
Ladders · United States · 1 wk ago
RemoteRemoteEngineering$218k–$257k/yrFull-time
Responsibilities
- Own reliability, monitoring, and incident response for AI infrastructure services, including on-call AWS support
- Automate operational IT workflows to enhance deployment velocity across CI/CD frameworks and Kubernetes
- Collaborate with the company’s Infrastructure team to extend CI/CD frameworks for IT services and integrate security tooling
- Enhance observability and documentation standards by implementing monitoring solutions and maintaining technical documentation
- Develop full-stack applications for internal AI products and infrastructure using Go or Python
Qualifications
- 8+ years of experience with cloud infrastructure automation (AWS) and network environments
- Hands-on experience with infrastructure-as-code tools (Terraform, Ansible, Chef, Puppet, or Salt)
- Proven ability to manage containerized workloads using Docker and Kubernetes in production environments
- Proficiency in a programming or scripting language (Python, Bash, Ruby, or Go) and git-based CI/CD workflows
- Demonstrated experience leading incident response and conducting root cause analysis in SLA-driven contexts
Benefits
- Medical, dental, and vision insurance
- 401(k) plan with company match
- Equity and bonus eligibility