Director, AI Platform Engineering
NYC Health + Hospitals · New York, NY · 1 wk ago
EngineeringFull-time
About the role
NYC Health + Hospitals is the largest public health care system in the United States, providing essential outpatient, inpatient and home-based services to more than one million New Yorkers every year across the city’s five boroughs. The Director of AI Platform Engineering provides strategic leadership for the cloud, platform, and deployment infrastructure supporting Artificial Intelligence (AI) across the System, ensuring AI systems used in clinical workflows operate safely, reliably, securely, and in compliance with applicable laws and NYC Health + Hospitals rules and regulations.
Responsibilities
- Provides strategic leadership for cloud, platform, and infrastructure engineering, developing and leading multi-year roadmaps, standards, and strategies for the secure and scalable deployment of AI products.
- Oversees architecture and governance of CI/CD pipelines, infrastructure‑as‑code (Terraform), and Kubernetes/Azure Kubernetes Services (AKS) orchestration to support reliable AI deployment.
- Defines and oversees the infrastructure for high-volume, low-latency data pipelines, feature stores, and data access layers required for training and real-time serving of AI models.
- Establishes enterprise-wide reliability and monitoring frameworks to ensure stable, safe operation of AI systems used by clinicians and care teams.
- Implements platform controls and audit trails to monitor and ensure Responsible AI practices, model explainability (XAI), and checks for model drift and bias on an ongoing basis.
- Partners with product management, product development, cybersecurity, Machine Learning Operations (MLOps) engineering, and interoperability teams to ensure AI platform readiness and safe integrations.
- Leads incident management and root‑cause analysis to minimize disruptions to clinical workflows and drive reliability improvements.
- Ensures the AI platform and infrastructure provide necessary controls, logging, and audit capabilities to meet compliance requirements and support AI safety frameworks.
- Develops long‑term platform resilience, disaster recovery, and cost optimization strategies to support System‑wide AI expansion.
- Defines and standardizes the platform's toolchain and API for the Machine Learning lifecycle, including model experimentation tracking (e.g., MLflow, ClearML), model registry, and automated testing/validation frameworks.
- Manages a team of platform and infrastructure engineers.
Requirements
- Master’s Degree from an accredited college or university in Computer Science, Information Systems or Technology, Cybersecurity, Hospital Administration, Health Care Planning, Business Administration, Mathematics, Engineering or Public Administration; and three (3) years of progressively responsible experience in health care information security, multifaced information technology, health and medical service administration, public administration, or a related discipline with an emphasis on systems programming, systems engineering, software developing, or providing technical support as a specialist; two (2) years of which must have been in a related administrative, managerial or supervisory capacity; or
- Bachelor’s Degree from an accredited college or university in disciplines as listed above; and five (5) years of progressively responsible experience in health care information security, multifaced information technology, health and medical service administration, public administration, or a related discipline with an emphasis on systems programming, systems engineering, software developing, or providing technical support as a specialist; two (2) years of which must have been in a related administrative, managerial or supervisory capacity.
Qualifications
- Master’s degree from an accredited college or university in Computer Science, Engineering, Information Systems, or related discipline; and
- Five (5) years of experience in Machine Learning Operations (MLOps), Machine Learning (ML) engineering, Artificial Intelligence (AI) platform engineering, or operating production Machine Learning (ML) / Large Language Model (LLM) systems; or
- Ten (10) years of experience in Software and Data Engineering.
Preferred Skills and Certifications
- Professional certifications in cloud architecture, ML/AI engineering, or DevOps from leading cloud platforms.
- Expertise in Azure architecture, Kubernetes/Azure Kubernetes Services (AKS), Terraform, Continuous Integration and Continuous Deployment (CI/CD), and automation frameworks.
- Experience supporting production AI/ML systems or mission‑critical workloads.
- Knowledge of observability tools, monitoring frameworks, and reliability engineering practices.
- Understanding of security and compliance standards including HIPAA and NIST.
- Demonstrated leadership, cross-functional collaboration, and technical communication skills.
- Strong stakeholder engagement and change‑management skills.
- Experience in healthcare, public sector, or other regulated environments.
- Experience deploying or supporting AI/ML, LLM, or agentic AI systems in production.
- Familiarity with Site Reliability Engineering (SRE) or platform engineering frameworks.
Technical Skills
- Languages: Python, Bash/Shell, HashiCorp Configuration Language (HCL Terraform), YAML/JSON; familiarity with Go preferred for Kubernetes-based platform tooling.
- Platforms: Azure cloud services, Docker, Kubernetes/AKS, Terraform, CI/CD platforms (Azure DevOps, GitHub Actions, Jenkins), monitoring/observability tools (Grafana, Azure Monitor), secrets/IAM security tooling.