Jobs · Engineering · Florida

Senior Site Reliability Engineer

ADT · Boca Raton, FL · 2 wk ago
On-siteEngineering$129k–$193k/yrFull-time

Applicants must be authorized to work for any employer in the United States without current or future sponsorship.

Locations and Workstyle

  • Blue Bell, PA: Primarily remote; candidates should be within commuting distance of the Blue Bell office and able to work onsite as needed. Option to come onsite more frequently if desired.
  • Irving, TX and Boca Raton, FL: Hybrid schedule - onsite a minimum of four days per week, with one remote day. Five days onsite may be required based on business needs.

Responsibilities

  • Own the reliability, availability, scalability, and performance of major production platforms and services.
  • Work closely with Infrastructure, Development, Product, and Operations teams to keep the ADT platform running and customers protected.
  • Drive operational excellence through automation, observability, and proactive problem-solving across large-scale distributed systems.
  • Lead reliability efforts across cloud environments, leveraging tools such as Terraform, Ansible, Kubernetes, Dynatrace, and Prometheus.
  • Design, build, and evolve infrastructure as code solutions to automate provisioning, configuration management, patching, and repeatable operational processes.
  • Own the operational lifecycle of critical services, including production readiness, monitoring strategy, incident response, and continuous improvement.
  • Identify reliability gaps, performance bottlenecks, and operational toil, then implement durable solutions that improve stability and efficiency.
  • Define and improve observability practices, including dashboards, alerts, runbooks, and detection strategies that reduce time to detect and resolve issues.
  • Support software releases and production changes, including validation, rollback planning, post-change verification, and operational readiness reviews.
  • Partner with cross-functional teams to improve operational health and establish engineering practices that strengthen platform reliability.
  • Mentor junior and mid-level engineers through design reviews, technical guidance, on-call coaching, and operational best practices.
  • Participate in an on-call rotation and provide production support, including complex incident response, root cause analysis, remediation efforts, and support for customer-impacting issues during major incidents.

Requirements

  • 7+ years of experience in Site Reliability Engineering, Infrastructure Engineering, Systems Engineering, Operations Engineering, DevOps, or related production-focused roles with on-call responsibility.
  • Proven experience owning large-scale distributed systems and platforms in production environments.
  • Strong Linux and systems engineering fundamentals, including troubleshooting across compute, storage, networking, and application layers.
  • Advanced experience with infrastructure as code, including Terraform and Ansible design, implementation, and maintenance.
  • Experience operating and optimizing cloud environments, including AWS and GCP.
  • Strong experience managing Kubernetes clusters in large-scale production environments.
  • Proficiency in Python, Bash, or similar scripting languages, with a focus on automation and operational tooling.
  • Strong understanding of software delivery, change management, and safe production operations.
  • Experience with monitoring and observability platforms such as Dynatrace, Prometheus, or similar technologies.
  • Ability to diagnose and resolve complex production issues while making sound decisions around risk, rollback strategies, and escalation paths.
  • Experience with CI/CD pipelines and deployment automation.
  • Experience leading or supporting incident response activities and post-incident reviews.
  • Strong communication skills and the ability to collaborate effectively across technical and business teams.
  • Comfortable operating in complex environments with ambiguity and competing priorities.
  • Ability to balance incident-driven operational work with longer-term reliability and automation initiatives.
  • Comfortable using AI tools and agents to accelerate investigation, automation, and documentation while maintaining sound engineering judgment.

Preferred Qualifications

  • Experience with Java/JVM ecosystems and large-scale customer-facing platforms.
  • Experience with Kafka or other distributed messaging platforms.
  • Experience leading security remediation efforts at scale, including patch SLAs, CVE response, and operating system upgrades.
  • Familiarity with Harness, enterprise Git workflows, and audit-driven change controls.
  • Demonstrated success improving reliability metrics such as MTTD, MTTR, availability, or operational efficiency across teams.

Pay

The salary range for this role is $128,800.00 - $193,200.00 and is based on experience and qualifications. Certain roles are eligible for annual bonus and may include equity. These awards are allocated based on company and individual performance.

Benefits

  • Healthcare benefits
  • 401(k) plan and company match
  • Short-term and long-term disability coverage
  • Life insurance
  • Wellbeing benefits
  • Paid time off (employees accrue up to 120 hours in their first year; accrual rate increases after the first year)
  • 6 paid holidays

Similar jobs