Service Reliability Engineer
Ladders · United States · 1 wk ago
RemoteRemoteEngineering$168k–$334k/yrFull-time
Remote – US based candidates only, no visa sponsorship available.
About the role
This role will lead technical work focused on secure, reliable systems and risk-aware execution. You will work across engineering, product, operations, and business stakeholders to translate complex requirements into practical technology solutions. The position offers the opportunity to influence architecture, execution quality, and the technology capabilities that enable long-term growth within a Telecommunications & Hardware environment.
Responsibilities
- Operate within a 24/7 follow-the-sun support model for global coverage
- Monitor and manage production GPU and Kubernetes environments for high performance
- Analyze logs and metrics to diagnose issues and implement resolutions
- Develop automated routines to prevent production issues before they happen
- Improve automation by integrating insights from incident analysis into solutions
- Perform systems and network administration alongside security monitoring tasks
- Coordinate with service owners to resolve complex system issues efficiently
Requirements
- 8+ years of experience in large-scale production systems, with 3+ years in high-availability settings
- Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management
- Proficient in observability and incident management tools like Grafana and PagerDuty
- Strong Linux system administration skills and automation using Ansible or Python
- Familiarity with GPU hardware and high-performance computing environments
Benefits
- Opportunity to work with cutting-edge technology in AI and computing
- Be part of a diverse and expert team enhancing product reliability
- Flexible work schedule with a 4-day, 10-hour week
- Potential career growth in a leading tech company
Pay
$168,000 – $333,500 annually