Senior Software Engineer, DevOps
Atria Health and Research Institute · United States · 2 days ago
RemoteRemoteEngineeringFull-time
About Atria: The Atria Health Institute is a membership-based primary and specialty health care practice with a focus on prevention and longevity. We bring together a multidisciplinary team of renowned physicians to provide proactive, preventive, and precision-based care for Atria members and their families. All care, including primary care, advanced screening and diagnostics, urgent care, specialty care, 24/7 home visits, and imaging is included in members' annual fee. Our mission is to make healthspan and lifespan equal for all by translating science into medicine in real-time, all while bringing humanity back into health care. Delivering such robust, personalized, and preventive health care is complex and requires a team-wide dedication to excellence. After successfully opening our flagship Institute in New York in 2022 and expanding to South Florida in 2024, we are now bringing the Atria experience to the West Coast with the launch of our Los Angeles Institute in late spring 2026. Atria Health is seeking a Senior Software Engineer for our DevOps team to help design, build, and operate the infrastructure, deployment pipelines, and observability that the rest of engineering relies on every day. This is an individual contributor role focused on: Owning infrastructure, CI/CD, and automation initiatives end to end; from design through rollout, adoption, and iterationSetting the patterns and standards that keep our systems reliable, observable, and secure as we scaleReducing operational toil and raising the bar on developer productivity across engineering You'll partner with the Tech Lead and other DevOps engineers to shape our infrastructure-as-code, build and release process, and monitoring and security posture — and you'll help define the standards others build on. You'll work closely with the Platform, Data Engineering, and domain product teams (Clinical Experience, Member Experience, and Care Delivery) to anticipate their needs and make their path to production smoother. You'll care deeply about both how we build and what we achieve, while working to accelerate the work of our client teams. Infrastructure & Automation Design, build, and maintain cloud infrastructure on Google Cloud Platform using Terraform, and help define the patterns and standards the team builds onOwn and improve CI/CD pipelines in GitHub Actions to make deployments fast, safe, and repeatableIdentify and eliminate sources of operational toil, building automation and tooling that scales across engineering Reliability & Observability Establish meaningful monitoring, dashboards, and alerting in Datadog, and drive down alert noise across the team's systemsLead incident response within the on-call rotation, and drive postmortems and follow-ups that improve uptime for our applicationsDefine and track Service Level Objectives (SLOs) for our core infrastructure and build systems, and hold the team to them Developer Enablement Partner with product engineering teams on deployment pipelines, environment issues, and build troubleshooting, and proactively remove recurring frictionOwn preview and staging environments, including reliable data sync, masking, and cleanup routines so teams can test against realistic dataImprove developer experience through better tooling, clear runbooks, and documentation Quality & Collaboration Write well-tested, well-reviewed infrastructure code, and lead design reviews and RFCs with a pragmatic, operations-focused perspectiveDesign systems and changes to meet the team's goals, weighing tradeoffs across reliability, performance, and securityGive thoughtful code reviews and mentor other engineers through pairing, knowledge sharing, and documentation Tech Stack Languages: TypeScript, Python, BashInfrastructure: Google Cloud Platform, Terraform, CloudflareCI/CD: GitHub ActionsDatabases: MySQL, PostgreSQL, RedisObservability & Incident Management: Datadog, Sentry, RootlyIntegrations: Athena EMR, wearable platforms, third-party healthcare APIs Requirements Core Experience :5+ years of professional experience in DevOps, SRE, infrastructure, or backend engineering in production environmentsHands-on experience designing and operating infrastructure in at least one cloud provider (ideally Google Cloud Platform)Track record of owning and shipping automation, pipelines, or infrastructure that made a team measurably more productive or reliableAn enthusiasm for developer productivity and making our teams as impactful as possible Technical Skills Deep experience with infrastructure-as-code (ideally Terraform) and building CI/CD pipelines (ideally GitHub Actions)Proficient software engineering ability, and strong command of LinuxStrong instincts for monitoring and observability, and confident debugging across logs, traces, and metricsSolid experience with relational databases (MySQL, PostgreSQL) and containerized workloadsStrong grounding in reliability, performance, and security fundamentals, with the judgment to make sound tradeoffs Nice to Have Experience in healthcare, digital health, or other regulated domains (HIPAA, PHI, SOC 2, etc.)Experience with containers and orchestration (Docker, Kubernetes)Exposure to leading incident response, on-call, and postmortem practicesExperience with database migrations or managing multiple environments at scale