Service Manager
About The Role
As a Service Manager – AI/LLM Platforms, you will make an impact by leading the operational excellence, reliability, and continuous improvement of AI-powered digital platforms supporting healthcare payer operations. You will oversee hybrid-cloud services leveraging Large Language Models (LLMs), MLOps practices, cloud-native technologies, and modern engineering frameworks to deliver secure, scalable, and compliant solutions that enhance member and provider experiences. You will be a valued member of the Technology & Engineering team, collaborating closely with business stakeholders, product teams, platform engineers, AI specialists, and operations teams to ensure service stability, innovation, and regulatory compliance.
Responsibilities
Own end-to-end service accountability for payer-focused AI, LLM, and ML-enabled platforms, ensuring high availability, performance, and compliance across hybrid environments.
Lead operational governance and continuous improvement initiatives aligned with enterprise service management best practices.
Oversee MLOps processes supporting LLM-powered applications, including model deployment, monitoring, retraining, validation, and rollback strategies.
Capture and implement best practices for Git-based version control, release management, code reviews, and repository governance across application and infrastructure teams.
Implement and govern Infrastructure as Code (IaC) practices using Terraform to provision and manage cloud and on-premises resources consistently and securely.
Support the development, deployment, and operational management of Node.js-based services and APIs that integrate AI and machine learning capabilities.
Drive operational excellence across Linux environments, including security hardening, patch management, performance optimization, and system monitoring.
Establish best practices for Python-based development supporting data pipelines, AI orchestration, automation frameworks, and analytics workloads.
Collaborate with business stakeholders, product owners, and healthcare domain experts to translate complex payer requirements into reliable technology services.
Lead incident management, root cause analysis, problem management, and service restoration activities to minimize business impact.
Maintain platform health through metrics, logs, traces, and observability tools while continuously improving service reliability and resilience.
Foster a culture of knowledge sharing, operational excellence, automation, and continuous improvement across distributed teams.
Requirements
Bachelor's degree in Computer Science, Information Technology, Engineering, Healthcare Informatics, or a related field.
Significant experience managing enterprise platforms and technology services within the Healthcare Payer domain.
Deep understanding of payer operations, including claims processing, member engagement, care management, provider interactions, and regulatory considerations.
Hands-on experience supporting and operating LLM, Generative AI, and MLOps platforms in production environments.
Strong expertise with Kubernetes, Terraform, and Ansible for cloud-native infrastructure management and automation.
Experience developing, supporting, or integrating applications using Node.js and Python.
Strong knowledge of Git-based development workflows, release governance, and software delivery practices.
Proven experience administering and supporting Linux-based environments.
Experience managing incident response, problem management, change management, and service-level objectives (SLOs).
Strong understanding of observability, monitoring, logging, and operational analytics.
Exceptional stakeholder management, communication, and leadership skills.
Qualifications
Experience supporting AI, analytics, or digital transformation initiatives within healthcare payer organizations.
Knowledge of cloud-native AI and machine learning services across major cloud providers.
Experience implementing Platform Engineering, Site Reliability Engineering (SRE), or DevOps best practices.
Familiarity with healthcare regulatory requirements, including HIPAA and data governance frameworks.
Experience managing distributed global teams operating in hybrid work environments.
Certifications in Kubernetes, Cloud Platforms, ITIL, Terraform, or AI/Machine Learning technologies.
Strong background in automation, self-healing architectures, and enterprise-scale service operations.
Skills
Leadership and stakeholder management skills.
Technical proficiency in Kubernetes, Terraform, Ansible, Node.js, Python, Git, Linux, and cloud platforms.
Experience with MLOps, observability, and healthcare regulatory compliance.
Ability to collaborate effectively with cross-functional teams and business stakeholders.
Benefits
Medical/Dental/Vision/Life Insurance.
Paid holidays plus Paid Time Off.
401(k) plan and contributions.
Long-term/Short-term Disability.
Paid Parental Leave.
Employee Stock Purchase Plan.