Jobs · Engineering · Texas

Director - Application Site-Reliability Engineering

Caris Life Sciences · Irving, TX · Yesterday
EngineeringFull-time

About the role

Reporting to the Corporate Vice President for Clinical Software Products, the Director owns application reliability and production operations for the clinical software portfolio: production support, incident response, on-call operations, and SLO management for applications under SOX financial controls and FDA regulatory requirements. The Director hires and develops the team, sets standards and selects tooling, establishes production-access governance and segregation-of-duties controls with the information-security, quality, and infrastructure organizations, and participates directly in incident response. The role carries wide latitude to shape how reliability engineering is done here.

Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning incident diagnostics, runbook authoring, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance. The infrastructure organization owns the platform and observability runtime; this role owns application-layer reliability on top of it. The operating model is automation-first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand.

Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.

Responsibilities

  • Own the production-support model for the clinical application portfolio: the on-call rotation, escalation procedures, and incident-response playbooks.
  • Establish and audit production-access and segregation-of-duties controls with engineering, information-security, quality, and infrastructure partners; keep the evidence audit-ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged-action logs.
  • Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational-health metrics such as mean time to detect, mean time to recover, and on-call burden; intervene while an error budget is burning, not after it is spent.
  • Lead incident response for high-severity production events as incident commander or senior technical responder; coordinate cross-functional teams, run post-incident reviews, and drive systemic remediation to closure.
  • Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures.
  • Drive the automation-first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self-service tooling over ticket-driven request work.
  • Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions.
  • Advance automated generation of change and deployment evidence in CI/CD pipelines: deployment records, approval trails, and change documentation as an audit-ready by-product of release.
  • Hire and develop the App-SRE team; set performance expectations, on-call responsibilities, and career growth frameworks.
  • Own application-layer observability alongside the infrastructure and observability platform teams; keep production dashboards, alerting thresholds, and SLO monitors accurate and actionable.
  • Represent App-SRE in engineering leadership forums, operational reviews, and compliance audits; translate operational health and risk into clear executive communication.
  • Run the function AI-first: make AI-assisted practice the team's daily norm, from runbook automation to incident analysis and operational tooling.

Required Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.
  • 10+ years of professional experience in SRE, DevOps, platform engineering, or production operations.
  • 4+ years of direct experience in a people-management or team-lead role within an SRE or production-operations function.
  • Hands-on experience leading incident response for Tier 1 or business-critical production systems, including serving as incident commander or senior technical responder.
  • Experience defining and implementing SLOs, error budgets, and associated alerting and on-call workflows in a production environment.
  • Experience building or significantly maturing a production-support, on-call, or SRE function, including runbook development and on-call-rotation design.
  • Experience operating production systems under a formal regulatory or financial-controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent, including producing documentation that holds up in audit.
  • A record of applying AI-assisted practice to operations or engineering work, personally or through a team.

Preferred Qualifications

  • Domain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software.
  • Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access-control, change-management, and monitoring controls.
  • Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.
  • Experience with deployment pipelines, release-management workflows, and rollback procedures in a continuous-delivery environment.
  • Ability to operate as a player-coach, contributing directly to technical work while building and leading a team.
  • Track record of reducing operational toil through automation programs in an SRE or production-operations organization.
  • Experience presenting operational strategy and risk posture to senior or executive audiences.

Other

This role serves in the senior escalation tier of the production on-call rotation, with after-hours response to high-severity incidents as incident commander or senior escalation point. Periodic travel may be required to support business needs, team on-sites, and leadership reviews.

Similar jobs