Jobs · Engineering

Principal AI/ML Engineer

Optum · Eden Prairie, MN · 1 wk ago
Engineering$165k–$282k/yrFull-time

Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives.

About the role

Optum AI is UnitedHealth Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML engineering for health care. We develop AI/ML solutions for the highest impact opportunities across UnitedHealth Group businesses including UnitedHealthcare, Optum Financial, Optum Health, Optum Insight, and Optum Rx. In addition to transforming the health care journey through responsible AI/ML innovation, our charter also includes developing and supporting an enterprise AI/ML development platform.

As a principal engineer within the enterprise AI Platforms team within the Office of AI, you will design, build and operate the platform that healthcare agents run on. This is a software engineering role at its core applied to a domain where the workloads are models, retrieval and agents. You will own systems end to end: analysis and design, coding, code review, testing and debugging, documentation, release and delivery, and the long maintenance tail after launch. Your work is enabling Optum clients with access to an AI Platform that provides secure, PHI-ready model access and favorable inference economics, a capability gateway that makes licensed intelligence callable from any approved agent or application, and the agent harness those agents run on. You will set the engineering standards, patterns and shared utilities other teams build on, and carry meaningful influence over the technical direction of the portfolio. Expect to prototype quickly with design partners, prove or disprove technical bets with evidence, and then harden what survives into platform-grade services that other engineers can easily adopt.

Responsibilities

  • Design, code, test, debug, document and maintain the services that make up the AI platform and agent runtime, taking systems from design review through production operation and long-term support
  • Design the public surface of the platform - APIs, SDKs, client libraries, schemas and versioning strategy - so that internal teams, client engineering teams and external builders can integrate without bespoke work for each consumer
  • Build the Capability Gateway: implement Model Context Protocol (MCP) servers and tool interfaces with typed contracts, authentication and authorization, input and output validation, quota and rate limiting, audit logging, and an automated conformance test suite that new capabilities must pass
  • Build and operate the agent harness and runtime: orchestration and tool-calling execution, state and session management, retrieval paths, concurrency, queuing and retries, timeout and failure handling, idempotency, and graceful degradation under partial outage
  • Establish engineering standards, methods and tooling for the portfolio: coding standards, code review practice, branching and release strategy, CI/CD pipelines, infrastructure as code, environment management, and automated testing at unit, contract, integration and end-to-end levels
  • Engineer the platform for cost and scale, including inference routing across models, caching, batching, connection and resource pooling, capacity planning, and per-tenant cost telemetry that ties spend to completed work
  • Code and deploy the ML and LLM components of the platform into production: model and prompt serving, retrieval and embedding pipelines, adapters and fine-tuned variants, and the deployment mechanics behind them including canary, shadow and rollback paths
  • Implement AI/ML capabilities in support of healthcare workflows, which may include natural language processing and understanding, semantic search, intent classification, information extraction, document AI and computer vision, and automatic speech recognition applied to clinical notes, claims, faxes, referrals, prior authorizations and recorded encounters
  • Build the evaluation and experimentation infrastructure the platform depends on: golden datasets, task-level benchmarks, offline and online evals, human-in-the-loop review workflows, regression gates wired into CI, and experiment tooling with defensible statistical treatment of results
  • Work with large-scale computing frameworks and data analysis systems to build the data foundation behind agent capabilities, including distributed processing, feature and vector stores, streaming ingestion, schema evolution and lineage
  • Engineer for PHI from the first commit: data minimization and de-identification, tenant isolation, encryption in transit and at rest, secrets management, retention and access controls, and enforcement that keeps agent behavior inside approved data and action boundaries
  • Evaluate new tools, techniques and strategies - models, frameworks, orchestration approaches, serving stacks - with enough rigor to make a build, buy or adopt recommendation, and communicate the results and trade-offs to leadership and internal stakeholders
  • Partner across Platform Enablement, product management, design, security, privacy and legal, forward-deployed engineering, operations and external partners such as Azure and Databricks on one integrated backlog with explicit interfaces and SLOs
  • Raise the engineering bar through reference implementations, reusable libraries, design review participation, code review and technical mentorship, and influence thought and leadership on where the platform should go next

Qualifications

  • Bachelor's degree in computer science, engineering or a related technical field, or equivalent experience
  • 10+ years of professional software engineering experience building and operating production systems
  • 5+ years of experience designing distributed services and platform-level APIs consumed by other engineering teams
  • Experience building or operating ML, AI or data-intensive systems in production, including responsibility for them after launch
  • Hands-on experience with LLM-based or agentic systems, including retrieval-augmented generation, tool and function calling, and orchestration
  • Expert-level proficiency in Python and at least one additional production language such as Go, Java, TypeScript, Scala or C++
  • Demonstrated depth in software engineering fundamentals: testing strategy, debugging complex production issues, performance profiling, code review and technical documentation
  • Experience with cloud-native delivery, including containers and Kubernetes, CI/CD, infrastructure as code, and observability tooling
  • Experience with large-scale computing frameworks and distributed data processing such as Spark/Databricks, Ray, Kafka or equivalent
  • Experience building software subject to security, privacy or regulatory constraints
  • Demonstrated ability to explain complex technical results and trade-offs clearly to non-technical stakeholders and senior leaders

Preferred Qualifications

  • Experience building multi-tenant platforms or services, including tenant isolation, quota and entitlement enforcement, and developer experience as an explicit product concern
  • Experience implementing or operating Model Context Protocol (MCP) servers, tool-calling standards, or agent frameworks such as LangGraph or equivalent
  • Experience with model serving and optimization at scale, including vLLM, TensorRT/Triton, batching, KV-cache strategies, quantization and multi-model routing
  • Experience with vector databases, embedding lifecycle management, hybrid retrieval and enterprise knowledge graphs
  • Experience with inference cost management or FinOps for AI workloads
  • Experience with responsible AI practice, including bias and fairness evaluation, model risk management, model documentation, or alignment to frameworks such as the NIST AI Risk Management Framework
  • Experience on a 0-to-1 product or platform, working directly with design partners or customers to shape what gets built
  • Experience with Azure AI services and Databricks in a regulated enterprise environment
  • Experience with healthcare data and standards: claims, EHR data, FHIR, HL7, ICD, CPT, HCPCS, SNOMED CT, LOINC, and working with PHI inside HIPAA/HITRUST boundaries
  • Health care industry experience including provider, payer, medical device, pharmaceutical, or other health services

Benefits

  • Comprehensive benefits package
  • Incentive and recognition programs
  • Equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements)

Pay

The salary for this role will range from $164,600 to $282,200 annually based on full-time employment.

Schedule

You'll enjoy the flexibility to work remotely from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Similar jobs

Principal AI/ML Engineer

Sierra Nevada CorporationLone Tree, CO· 3 mo ago
Engineering$165k–$227k/yrapply on snc.wd1.myworkdayjobs.com