Corporate Vice President - Site Reliability Senior Engineer
New York Life · New York, United States · 2 days ago
Engineering$148k/yrFull-time
Role Overview
New York Life is seeking an experienced Site Reliability Engineer (SRE) to join the Cloud Platform Engineering Team and provide both technical leadership and day-to-day management for multi cloud solution delivery (AWS, GCP, Azure).
What You'll Do
- Support our Cloud Business Office by facilitating Well Architected reviews for new cloud deployments or cloud migration candidates.
- Lead the design and delivery of repeatable, reliable cloud solutions using Terraform, GitHub, and related automation tools.
- Develop and deploy full stack patterns to support ongoing application development, data platform capabilities, and AI services (RAG & Agentic).
- Partner with application and product teams to assess requirements, recommend AWS architecture patterns, and support successful cloud implementations.
- Support the onboarding and lifecycle management of cloud services by producing standard IaC/Terraform Modules.
- Define and mature SRE practices, including SLO/SLI frameworks and error-budget governance.
- Design and implement automation solutions using Java, JavaScript, APIs, SQL, and Terraform.
- Deliver application-level fixes and enhancements through disciplined software engineering.
- Focus on key reliability and performance indicators: uptime, system throughput, system output, and download rate/application load speed.
- Lead the shift from non-standard application platforms to standard software artifacts (Terraform modules, secure base images, YAML templates, Java libraries) integrated into CI/CD pipelines, creating reusable patterns and reducing repetitive configuration and coding tasks.
- Review cloud platform designs, operational issues, and delivery risks; provide practical recommendations and follow-through to resolution.
- Advise teams on cloud cost optimization, resource utilization, governance, and platform standards.
- Maintain detailed records of issues, actions taken, and outcomes to support continuous improvement efforts.
- Collaborate with other IT teams and external vendors to resolve complex issues and implement solutions.
- Identify opportunities to improve support processes and implement best practices to enhance overall efficiency.
- Provide training and guidance to IT staff on network and platform support techniques and best practices.
What You'll Bring
- 8+ years of overall IT, infrastructure, platform engineering, or cloud engineering experience.
- 5+ years of experience with AWS cloud platforms and services.
- 3+ years of experience designing, automating, and operating scalable, secure, and highly available cloud solutions.
- Ability to produce production grade IaC and perform code reviews.
- Demonstrated ability to lead, mentor, and coordinate engineers or technical contributors in a matrixed environment.
- Strong business and technical acumen with the ability to influence decisions, manage tradeoffs, and communicate with leadership.
- Strong facilitation, planning, and execution skills for cross-functional workshops, technical reviews, and delivery initiatives.
- Solid understanding of SDLC, DevOps, change management, incident response, and production operations practices.
- The ability to work directly with business stakeholders, architects, security teams, product owners, and engineering teams to translate needs into practical cloud solutions.
Desired Skills
- Experience with cloud governance, landing zones, identity and access management, networking, compute, storage, and security control patterns.
- Hands-on experience with infrastructure as code practices using Terraform, CloudFormation, AWS CDK, or similar AWS-focused tools.
- Experience with Kubernetes, Amazon EKS, container platforms, CI/CD pipelines, and observability tools.
- Practical understanding of infrastructure technologies, including compute, network, storage, DNS, load balancing, encryption, logging, and monitoring.
- Practical knowledge of scripting or programming languages such as Python, PowerShell, .NET, C#, Java, or similar.
- Strong time management, prioritization, coaching, and problem-solving skills with the ability to balance delivery commitments and operational needs.
- Background with cloud cost management, FinOps practices, compliance requirements, and risk-based decision making.