Jobs · Minnesota

Lead SWE, AI Dev/IT Operations - Remote or Hybrid in DC or MN

Optum · Minnetonka, MN · 4 wk ago
$113k–$193k/yrFull-time

Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives.

About the role

We're seeking a skilled Lead Software Engineer to help build intelligent automation, predictive analytics, and agentic AI capabilities for IT Operations. In this role, you will design, build, and help operationalize machine learning models and AI-driven workflows that improve observability, speed up incident response, and support self-healing infrastructure across the enterprise. You'll work closely with senior engineers, architects, and data scientists to bring these capabilities into production. You'll enjoy the flexibility to work remotely from anywhere within the U.S. For hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Responsibilities

  • AI/ML Development: Design, build, and maintain machine learning models and AI-driven automation for IT operations use cases, including predictive analytics, intelligent automation, and support for autonomous remediation
  • Data & Platform Engineering: Build and maintain data pipelines that ingest structured and unstructured data from observability and ITSM platforms (e.g., Splunk, Dynatrace, ServiceNow, logs, metrics, traces)
  • Feature Engineering & Model Support: Develop and maintain feature sets that support anomaly detection, predictive alerting, capacity forecasting, and operational intelligence use cases
  • Model Operationalization: Partner with data scientists and platform teams to deploy and productionize ML models, and help ensure they perform reliably at scale
  • AI-Driven Automation: Integrate ML and AI outputs into orchestration and automation platforms (e.g., Jenkins, Interlink) to support closed-loop, self-healing operational workflows
  • Agentic AI Development: Build components of agentic AI solutions that reason, plan, and execute actions across IT operations, including incident triage, root cause analysis, and resolution
  • MLOps Practices: Follow established MLOps standards, model lifecycle practices, and data governance requirements aligned with enterprise security, privacy, and compliance policies
  • Cloud & Infrastructure: Use cloud native and hybrid architectures to support real-time analytics, scalable model inference, and resilient AI platforms
  • Cross-Functional Collaboration: Work with engineering, operations, security, and business stakeholders to understand requirements, document your work, and contribute reusable patterns for AI-driven operations
  • Responsible AI: Apply ethical AI principles, including fairness, transparency, explainability, and accountability, throughout the development lifecycle

Requirements

  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience
  • 3+ years of experience building and deploying AI/ML solutions in production environments
  • 3+ years of experience in Python and SQL, with hands-on experience using distributed data processing frameworks such as Spark
  • 2+ years of experience applying ML to IT operations, infrastructure, reliability engineering, or observability domains
  • 2+ years of experience with cloud platforms (Azure, AWS, or GCP), including hybrid and cloud native architectures
  • 2+ years of experience with MLOps practices, CI/CD pipelines, model monitoring, and lifecycle management
  • 1+ years of experience with NLP, Large Language Models (LLMs), Generative AI, and modern AI frameworks
  • 1+ years of experience with containerization and orchestration technologies such as Docker and Kubernetes

Preferred Qualifications

  • Hands-on experience contributing to agentic or autonomous AI systems in operational contexts
  • Experience in healthcare, regulated industries, or large-scale enterprise platforms
  • Experience working in agile, product-driven environments delivering scalable, high-impact solutions
  • Working knowledge of IT operations concepts, including incident management, CMDB, monitoring, alerting, and reliability engineering
  • Familiarity with Responsible AI and governance frameworks, including fairness, transparency, and explainability
  • Solid technical communication skills, including the ability to explain AI concepts to non-technical stakeholders

Benefits

In addition to your salary, we offer a comprehensive benefits package, incentive and recognition programs, equity stock purchase, and 401k contribution (all benefits are subject to eligibility requirements).

Pay

The salary for this role will range from $112,700 to $193,200 annually based on full-time employment. Pay is based on several factors including but not limited to local labor markets, education, work experience, and certifications.

Similar jobs