Systems Reliability Engineer
The Systems Reliability Engineer will be responsible for maintaining and improving the reliability, performance, and scalability of large-scale production systems. The role requires a balance of software engineering, infrastructure expertise, automation, and operational leadership.
Responsibilities
- Design, build, and maintain reliable distributed systems while improving availability and performance.
- Develop automation tools and solutions using programming languages such as Python, Go, or Java.
- Operate and troubleshoot Linux-based systems at scale, including networking, performance optimization, and system-level issues.
- Manage Kubernetes and containerized workloads in production environments.
- Build, maintain, and improve CI/CD pipelines supporting both applications and infrastructure.
- Implement and enhance observability solutions using tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or similar platforms.
- Maintain system health, define reliability improvements, and contribute to proactive performance management.
- Lead incident response activities, conduct post-incident reviews, and implement preventative improvements.
- Support the development and adoption of reliability practices including automation, service ownership, and operational excellence.
- Collaborate with engineering and cross-functional teams to improve system design and resilience.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- 5+ years of experience in Site Reliability Engineering, DevOps, or production engineering roles supporting large-scale distributed systems.
- Strong programming experience in at least one of the following: Python, Go, or Java.
- Deep hands-on experience managing Linux environments, including networking, troubleshooting, and performance tuning.
- Production experience with Kubernetes and container-based architectures.
- Strong knowledge of observability and monitoring platforms such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or equivalent solutions.
- Experience designing and managing CI/CD pipelines for applications and infrastructure.
- Understanding of distributed system concepts including consistency models, partitioning, and failure handling.
- Proven experience leading incident response and conducting effective post-mortem reviews.
- Excellent communication, collaboration, and technical documentation skills.
- Preferred experience with SLOs, error budgets, chaos engineering practices, and reliability frameworks.
- Preferred familiarity with cloud platforms such as AWS, Azure, or GCP.
- Preferred background in capacity planning, performance engineering, load testing, or service mesh technologies.
Benefits
Competitive annual salary range of $100,000-$150,000.
Fully remote work environment within the United States.
Full-time employment opportunity with career growth potential.
Opportunity to work on large-scale distributed systems and advanced technology projects.
Collaborative environment focused on engineering excellence and continuous improvement.
Exposure to modern cloud, automation, observability, and reliability practices.
Inclusive workplace committed to equal employment opportunities.