Jobs · Engineering

Staff Site Reliability Engineer, Core AI Infrastructure

Ladders · United States · 1 wk ago
RemoteRemoteEngineering$218k–$257k/yrFull-time

Responsibilities

  • Own reliability, monitoring, and incident response for AI infrastructure services, including on-call AWS support
  • Automate operational IT workflows to enhance deployment velocity across CI/CD frameworks and Kubernetes
  • Collaborate with the company’s Infrastructure team to extend CI/CD frameworks for IT services and integrate security tooling
  • Enhance observability and documentation standards by implementing monitoring solutions and maintaining technical documentation
  • Develop full-stack applications for internal AI products and infrastructure using Go or Python

Qualifications

  • 8+ years of experience with cloud infrastructure automation (AWS) and network environments
  • Hands-on experience with infrastructure-as-code tools (Terraform, Ansible, Chef, Puppet, or Salt)
  • Proven ability to manage containerized workloads using Docker and Kubernetes in production environments
  • Proficiency in a programming or scripting language (Python, Bash, Ruby, or Go) and git-based CI/CD workflows
  • Demonstrated experience leading incident response and conducting root cause analysis in SLA-driven contexts

Benefits

  • Medical, dental, and vision insurance
  • 401(k) plan with company match
  • Equity and bonus eligibility

Similar jobs