Sr Software Engineer - Reliability Engineering
About the role
We're hiring a Senior Software Engineer specializing in Reliability Engineering who can code across the stack and prioritizes reliability. You'll design and build infrastructure, observability tooling, and operational systems, treating resilience and debuggability as first-class concerns. You'll own projects end-to-end—from architecture to code deployment in production. Your work will span infrastructure-as-code, incident response, system improvements, and mentoring, all while maintaining operational excellence for a platform serving millions of dealership transactions daily.
Responsibilities
- SRE Best Practices & System Health Management
- Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
- Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
- Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
- Evolve team's SRE standards and practices.
- Application Monitoring & Observability
- Build observability into systems: logging, metrics, distributed tracing, alert design.
- Improve monitoring frameworks; enable faster incident detection and resolution.
- Design dashboards and alerts that help teams understand system behavior.
- Partner with application teams on service instrumentation.
- AWS Cost Optimization
- Drive significant reductions in cloud spend through architectural improvements and resource utilization.
- Review infrastructure for efficiency; identify and eliminate waste.
- Balance cost, performance, and reliability in design decisions.
- Software Development & Architecture
- Build production systems, APIs, internal tools, and automation with clean, well-tested code.
- Design for maintainability, operational simplicity, and reliability.
- Participate in code review and technical design discussions.
- Mentor junior engineers on code quality and architectural thinking.
- Operations & Incident Response
- Participate in on-call rotations; debug and resolve production incidents.
- Conduct postmortem analysis; drive systemic improvements.
- Develop operational procedures and runbooks.
Qualifications
- 5+ years of software engineering, platform engineering, or infrastructure engineering experience.
- Strong coding skills in Python, Go, Java, or equivalent; writes clean, testable code.
- Hands-on AWS experience: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
- Terraform or equivalent infrastructure-as-code experience.
- Docker and container orchestration (Kubernetes or similar).
- Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
- System design thinking: can architect scalable systems and reason about trade-offs.
- Availability for rotational on-call duties outside of standard business hours may be required.
- Bachelor's degree in a related discipline and 4 years' experience in a related field. The right candidate could also qualify with a different combination, such as a master's degree and 2 years' experience; a Ph.D. and up to 1 year of experience; or 16 years' experience in a related field.
Skills
Highly Valued:
- Experience with observability tools (New Relic, Splunk, Prometheus).
- Incident response experience; familiar with postmortem practices.
- Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
- Cost optimization mindset; has identified and eliminated cloud waste.
- Windows and Linux system troubleshooting and performance analysis.
- Experience with CI/CD pipelines and deployment automation.
Pay
Compensation includes a base salary in the range of $121,800.00 - $203,000.00 per year. The base salary may vary within this range based on factors such as location and the selected candidate's knowledge, skills, and abilities. The position may also be eligible for additional compensation, including an incentive program.
Benefits
- Flexible vacation policy: take as much paid time off as consistent with duties, company needs, and obligations.
- Seven paid holidays throughout the calendar year.
- Up to 160 hours of paid wellness time annually for personal or family wellness.
- Additional paid time off for bereavement, voting, jury duty, volunteering, military leave, and parental leave.