Senior Site Reliability Engineer
ADT · Boca Raton, FL · 2 wk ago
On-siteEngineering$129k–$193k/yrFull-time
Applicants must be authorized to work for any employer in the United States without current or future sponsorship.
Locations and Workstyle
- Blue Bell, PA: Primarily remote; candidates should be within commuting distance of the Blue Bell office and able to work onsite as needed. Option to come onsite more frequently if desired.
- Irving, TX and Boca Raton, FL: Hybrid schedule - onsite a minimum of four days per week, with one remote day. Five days onsite may be required based on business needs.
Responsibilities
- Own the reliability, availability, scalability, and performance of major production platforms and services.
- Work closely with Infrastructure, Development, Product, and Operations teams to keep the ADT platform running and customers protected.
- Drive operational excellence through automation, observability, and proactive problem-solving across large-scale distributed systems.
- Lead reliability efforts across cloud environments, leveraging tools such as Terraform, Ansible, Kubernetes, Dynatrace, and Prometheus.
- Design, build, and evolve infrastructure as code solutions to automate provisioning, configuration management, patching, and repeatable operational processes.
- Own the operational lifecycle of critical services, including production readiness, monitoring strategy, incident response, and continuous improvement.
- Identify reliability gaps, performance bottlenecks, and operational toil, then implement durable solutions that improve stability and efficiency.
- Define and improve observability practices, including dashboards, alerts, runbooks, and detection strategies that reduce time to detect and resolve issues.
- Support software releases and production changes, including validation, rollback planning, post-change verification, and operational readiness reviews.
- Partner with cross-functional teams to improve operational health and establish engineering practices that strengthen platform reliability.
- Mentor junior and mid-level engineers through design reviews, technical guidance, on-call coaching, and operational best practices.
- Participate in an on-call rotation and provide production support, including complex incident response, root cause analysis, remediation efforts, and support for customer-impacting issues during major incidents.
Requirements
- 7+ years of experience in Site Reliability Engineering, Infrastructure Engineering, Systems Engineering, Operations Engineering, DevOps, or related production-focused roles with on-call responsibility.
- Proven experience owning large-scale distributed systems and platforms in production environments.
- Strong Linux and systems engineering fundamentals, including troubleshooting across compute, storage, networking, and application layers.
- Advanced experience with infrastructure as code, including Terraform and Ansible design, implementation, and maintenance.
- Experience operating and optimizing cloud environments, including AWS and GCP.
- Strong experience managing Kubernetes clusters in large-scale production environments.
- Proficiency in Python, Bash, or similar scripting languages, with a focus on automation and operational tooling.
- Strong understanding of software delivery, change management, and safe production operations.
- Experience with monitoring and observability platforms such as Dynatrace, Prometheus, or similar technologies.
- Ability to diagnose and resolve complex production issues while making sound decisions around risk, rollback strategies, and escalation paths.
- Experience with CI/CD pipelines and deployment automation.
- Experience leading or supporting incident response activities and post-incident reviews.
- Strong communication skills and the ability to collaborate effectively across technical and business teams.
- Comfortable operating in complex environments with ambiguity and competing priorities.
- Ability to balance incident-driven operational work with longer-term reliability and automation initiatives.
- Comfortable using AI tools and agents to accelerate investigation, automation, and documentation while maintaining sound engineering judgment.
Preferred Qualifications
- Experience with Java/JVM ecosystems and large-scale customer-facing platforms.
- Experience with Kafka or other distributed messaging platforms.
- Experience leading security remediation efforts at scale, including patch SLAs, CVE response, and operating system upgrades.
- Familiarity with Harness, enterprise Git workflows, and audit-driven change controls.
- Demonstrated success improving reliability metrics such as MTTD, MTTR, availability, or operational efficiency across teams.
Pay
The salary range for this role is $128,800.00 - $193,200.00 and is based on experience and qualifications. Certain roles are eligible for annual bonus and may include equity. These awards are allocated based on company and individual performance.
Benefits
- Healthcare benefits
- 401(k) plan and company match
- Short-term and long-term disability coverage
- Life insurance
- Wellbeing benefits
- Paid time off (employees accrue up to 120 hours in their first year; accrual rate increases after the first year)
- 6 paid holidays