AI Platform Engineer
About the role
Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.
The AI Platform Engineer builds and operates the machine learning and generative AI platform used by teams across Abbott Cancer Diagnostics. You'll own the full model lifecycle in production — data and feature pipelines, training and experimentation, evaluation and promotion, serving, and monitoring — along with the platform services, compute and tooling underneath it. This is hands-on infrastructure work backed by solid platform engineering practice: making inference fast and cheap, making the path from experiment to production repeatable and auditable, and shipping interfaces other engineers can build on — in support of software that ultimately reaches patients.
Responsibilities
- Build and maintain data, feature, and training pipelines for ML and LLM workloads — ingestion, transformation, fine-tuning, distributed training, and reproducible experiment execution with lineage tracked from dataset and code to resulting model.
- Implement automated evaluation and promotion gates — performance benchmarks, regression checks, and validation criteria that determine whether a model advances toward production.
- Automate the model lifecycle end to end through CI/CD and GitOps: packaging, promotion across environments, progressive rollout, and rollback.
- Build and operate production model-serving infrastructure for LLMs and predictive models, including inference optimization, autoscaling, and low-latency serving across multiple model formats and runtimes.
- Architect and manage GPU infrastructure — scheduling, autoscaling, resource isolation, and utilization efficiency for training and inference workloads.
- Instrument the platform and the models on it — structured logging, telemetry, drift detection, and cost tracking — and build the triggers and pipelines that close the loop into retraining and revalidation.
- Extend model, dataset, and artifact registries, metadata systems, and versioning so every deployed model has a traceable, auditable history.
- Build the developer-facing surface of the platform: APIs, SDK components, templates, documentation, and runbooks that make it self-service, backed by well-tested code and active participation in design and code review.
- Uphold company mission and values through accountability, innovation, integrity, quality, and teamwork.
- Maintain regular and reliable attendance.
- Act with an inclusion mindset and model these behaviors for the organization.
Minimum Qualifications
- Bachelor's degree in Computer Science, Engineering, AI/ML, or a related field; or equivalent practical experience.
- 3+ years building and operating production software, including significant work on ML or AI infrastructure.
- Strong Python, including software engineering fundamentals — automated testing, code review, and designing code others will read and extend.
- Experience with the ML model lifecycle: pipelines that carry a model from training through evaluation, deployment, monitoring, and retraining.
- Kubernetes experience — deploying, scaling, and debugging containerized workloads on a major cloud provider (AWS preferred).
- Experience with CI/CD, GitOps-based delivery, and infrastructure automation.
Preferred Qualifications
- ML workflow orchestration and experiment tracking (for example Kubeflow, Argo Workflows, Airflow, MLflow, or Weights & Biases).
- GPU infrastructure at scale: scheduling, multi-tenancy and resource isolation, autoscaling, and inference optimization such as quantization, batching, or graph compilation (for example ONNX Runtime or TensorRT).
- Feature stores, data versioning, or dataset lineage tooling.
- Progressive delivery for models — canary, shadow, or A/B deployment and automated rollback.
- Model monitoring and drift detection, including automated retraining triggers.
- Model governance and reproducibility practices: lineage tracking, audit trails, and approval workflows.
- Kubernetes-native serverless and event-driven autoscaling (for example Knative or KEDA).
- Experience building APIs or shared libraries consumed by other engineering teams.
- Production experience in an additional systems language such as Go or Java.
- Fine-tuning foundation models, or building distributed training pipelines.
- Experience delivering software in a regulated environment (HIPAA, CLIA, GxP, SOC 2).
The base pay for this position is $61,300.00 – $122,700.00. In specific locations, the pay range may vary from the range posted.