Lead Engineer - Observability Platform
About the role
Join CVS Health Enterprise Technology and put your engineering skills to use shaping the future of developer platforms for software engineers at Fortune 6 scale. The CVS Observability Platform gives engineers at CVS frictionless access to application instrumentation no matter where their applications run. As a Lead Cloud Engineer on the Observability Platform team, you will be architecting and scaling the data pipelines for billions of logs, metrics, and traces produced by thousands of workloads spanning multiple public cloud providers and datacenters. You will leverage open standards, open-source technologies, and in-house custom software to extend and create new platform features. You will work closely with the SRE team and engage with application teams across the company to inform the direction of the platform as it grows.
Responsibilities
- Design, build, and operate core observability platform services using Go, Python, Java (Spring Boot)
- Lead enterprise-wide adoption of OpenTelemetry, including client libraries, semantic conventions, instrumentation patterns, and Collector/agent strategy
- Architect and scale high-throughput, fault-tolerant telemetry pipelines (logs, metrics, traces) with a focus on performance, reliability, and cost efficiency
- Develop self-service observability capabilities that simplify onboarding, troubleshooting, and adoption for application teams
- Implement end-to-end monitoring of the observability platform itself, defining SLOs, health checks, and alerting
- Collaborate with SRE, Platform, and Cloud teams to establish reliability standards, error budgets, and incident response practices
- Participate in on-call rotations and lead incident mitigation, root-cause analysis, and post-incident reviews
- Automate operational workflows and eliminate manual toil through tooling, CI/CD enhancements, and platform automation
- Ensure secure telemetry pipelines through mTLS, secrets management, and zero-trust design patterns
- Produce and maintain high-quality technical documentation, standards, and best practices
- Engage with internal engineering teams to gather requirements, influence roadmap prioritization, and deliver platform improvements
- Provide technical leadership through mentorship, design reviews, architectural guidance, and cross-team collaboration with principal engineers and engineering leadership
Requirements
- 7+ years of experience in Software Engineering
- 5+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management
- 5+ years building production-grade backend services in Go and/or Java
- 5+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns
- 5+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD)
- 5+ years working with public cloud platforms (AWS, GCP, or Azure)
- 3+ years designing and scaling distributed, high-volume data pipelines
- 3+ years of experience with Helm charts, Kustomize, etc.
- 3+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Mimir)
- 3+ years of experience with Infrastructure as Code tools such as Terraform or CloudFormation
- 3+ years with relational databases (PostgreSQL, MySQL)
Preferred Qualifications
- Experience with service meshes and networking technologies such as Envoy and Istio
- Experience integrating or operating commercial observability platforms (Datadog, New Relic, AppDynamics, etc.)
- Experience with On-Call Scheduling tools such as OpsGenie, PagerDuty, GoAlert, etc.
- Experience with streaming and data platforms such as Kafka, Pulsar, or similar technologies
- Familiarity with time-series, NoSQL, or analytical databases (ClickHouse, Bigtable, Cassandra, etc.)
- Experience with cost optimization and capacity planning for large-scale telemetry systems
- Experience with chaos engineering, resiliency testing, or fault injection
- Background in security-aware platform design, including secure service-to-service communication
- Experience mentoring senior engineers and influencing platform standards across organizations
- Strong operational experience supporting 24x7 production systems, including on-call responsibilities
- Strong technical communication and cross-team collaboration skills
- Experience operating in regulated or compliance-heavy environments (e.g., healthcare, finance)
Qualifications
Bachelor’s degree from accredited university or equivalent work experience (HS diploma + 4 years relevant experience)
Pay
The typical pay range for this role is $106,605.00 - $284,280.00. This pay range represents the base annual full-time salary for all positions in the job grade within which this position falls. The actual base salary offer will depend on a variety of factors including experience, education, geography, and other relevant factors. This position is eligible for a CVS Health bonus, commission, or short-term incentive program in addition to the base pay range listed above.
Benefits
This full-time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well-being of colleagues and their families. The benefits for this position include:
- Medical, dental, and vision coverage
- Paid time off
- Retirement savings options
- Wellness programs
- Other resources, based on eligibility
Schedule
- Full time
- 40 anticipated weekly hours