Director - Application Site-Reliability Engineering
About the role
Reporting to the Corporate Vice President for Clinical Software Products, the Director owns application reliability and production operations for the clinical software portfolio: production support, incident response, on-call operations, and SLO management for applications under SOX financial controls and FDA regulatory requirements. The Director hires and develops the team, sets standards and selects tooling, establishes production-access governance and segregation-of-duties controls with the information-security, quality, and infrastructure organizations, and participates directly in incident response. The role carries wide latitude to shape how reliability engineering is done here.
Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning incident diagnostics, runbook authoring, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance. The infrastructure organization owns the platform and observability runtime; this role owns application-layer reliability on top of it. The operating model is automation-first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand.
Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.
Responsibilities
- Own the production-support model for the clinical application portfolio: the on-call rotation, escalation procedures, and incident-response playbooks.
- Establish and audit production-access and segregation-of-duties controls with engineering, information-security, quality, and infrastructure partners; keep the evidence audit-ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged-action logs.
- Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational-health metrics such as mean time to detect, mean time to recover, and on-call burden; intervene while an error budget is burning, not after it is spent.
- Lead incident response for high-severity production events as incident commander or senior technical responder; coordinate cross-functional teams, run post-incident reviews, and drive systemic remediation to closure.
- Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures.
- Drive the automation-first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self-service tooling over ticket-driven request work.
- Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions.
- Advance automated generation of change and deployment evidence in CI/CD pipelines: deployment records, approval trails, and change documentation as an audit-ready by-product of release.
- Hire and develop the App-SRE team; set performance expectations, on-call responsibilities, and career growth frameworks.
- Own application-layer observability alongside the infrastructure and observability platform teams; keep production dashboards, alerting thresholds, and SLO monitors accurate and actionable.
- Represent App-SRE in engineering leadership forums, operational reviews, and compliance audits; translate operational health and risk into clear executive communication.
- Run the function AI-first: make AI-assisted practice the team's daily norm, from runbook automation to incident analysis and operational tooling.
Required Qualifications
- Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.
- 10+ years of professional experience in SRE, DevOps, platform engineering, or production operations.
- 4+ years of direct experience in a people-management or team-lead role within an SRE or production-operations function.
- Hands-on experience leading incident response for Tier 1 or business-critical production systems, including serving as incident commander or senior technical responder.
- Experience defining and implementing SLOs, error budgets, and associated alerting and on-call workflows in a production environment.
- Experience building or significantly maturing a production-support, on-call, or SRE function, including runbook development and on-call-rotation design.
- Experience operating production systems under a formal regulatory or financial-controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent, including producing documentation that holds up in audit.
- A record of applying AI-assisted practice to operations or engineering work, personally or through a team.
Preferred Qualifications
- Domain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software.
- Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access-control, change-management, and monitoring controls.
- Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.
- Experience with deployment pipelines, release-management workflows, and rollback procedures in a continuous-delivery environment.
- Ability to operate as a player-coach, contributing directly to technical work while building and leading a team.
- Track record of reducing operational toil through automation programs in an SRE or production-operations organization.
- Experience presenting operational strategy and risk posture to senior or executive audiences.
Other
This role serves in the senior escalation tier of the production on-call rotation, with after-hours response to high-severity incidents as incident commander or senior escalation point. Periodic travel may be required to support business needs, team on-sites, and leadership reviews.