Jobs · Engineering · California

Agentic AI Engineer — Healthcare AI

Deloitte · Costa Mesa, CA · 3 days ago
HybridEngineering$111k–$373k/yrFull-time

About the role

This role is part of Deloitte's AI-first effort to rebuild the decision-making machinery behind American healthcare. You will design, build, and operationalize systems that support clinical reasoning, prior authorization, claims integrity, care navigation, and operational workflows.

Responsibilities

  • Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes.
  • Build stateful workflows using frameworks like LangGraph and LangChain - including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
  • Engineer for long-horizon reliability - multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail.
  • Build the reasoning behind regulated decisions - policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for high-stakes decisions across the industry.
  • Develop end-to-end Retrieval-Augmented Generation (RAG) pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
  • Engineer memory and context management - conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
  • Apply modern context-delivery patterns (e.g., MCP-style tool/context interfaces) so agents access the right information at the right time.
  • Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
  • Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
  • Evaluate agents at the trajectory and task level - multi-step task success, failure-mode and regression analysis, and sandboxed test environments - alongside retrieval- and generation-quality metrics, automated checks, and human review.
  • Engineer healthcare-grade safety - deployment eval gates, human-oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-aware data handling.
  • Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate safely within real business workflows.
  • Deliver production-quality code with strong practices in testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, safety, latency, cost, and model risk.
  • Partner with our modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning - through evaluation-driven feedback and, where it helps, fine-tuned or reasoning-optimized models.
  • Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
  • Demonstrated depth building and shipping production agentic systems - this is your primary craft, not a recent exploration.
  • Strong, hands-on experience building production agent systems with modern orchestration - LangGraph/LangChain or equivalent, including custom orchestration.
  • Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
  • Strong understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
  • Deep, practical understanding of LLM behavior - strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs - and the evaluation methods used to measure them.
  • Experience evaluating and debugging agent behavior - task-success and trajectory analysis, not just output quality.
  • Strong Python engineering skills and modern software practices: testing, CI/CD, version control, and API integration; experience implementing observability, tracing, and debugging for LLM-based systems in production.
  • Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
  • Ability to travel 0-50%, on average, based on the work you do and the clients and industries/sectors you serve.
  • Limited immigration sponsorship may be available.

Qualifications

  • Experience with multi-agent systems and agent collaboration patterns.
  • Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
  • Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
  • Understanding of traditional NLP concepts: tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
  • Experience operating in highly regulated, high-stakes, or operationally complex environments; healthcare exposure - clinical, payer, or life-sciences workflows, or standards such as FHIR - is a plus, not a requirement.
  • Demonstrated habit of staying current with AI research, benchmarks, and emerging engineering patterns.

Skills

  • Strong software/ML fundamentals.
  • Substantial, recent hands-on agentic work.
  • Modern software practices: testing, CI/CD, version control, and API integration.
  • Observability and tracing for LLM-based systems in production.
  • Production tool use and agent capabilities.
  • Model adaptation and fine-tuning techniques.
  • Traditional NLP concepts: tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
  • Highly regulated, high-stakes, or operationally complex environments.
  • Adequate understanding of AI research, benchmarks, and emerging engineering patterns.

Benefits

The role offers a competitive compensation package, including a base salary benchmarked to leading technology companies and a substantial performance-based incentive opportunity designed to grow with the value you help create. The estimated base salary range is $110,700-$372,900 (not adjusted for geographic differential); actual base pay depends on your skills, experience, and level, and you may also be eligible for a discretionary annual incentive based on individual and organizational performance.

Pay

The estimated base salary range is $110,700-$372,900 (not adjusted for geographic differential); actual base pay depends on your skills, experience, and level, and you may also be eligible for a discretionary annual incentive based on individual and organizational performance.

Schedule

Work you'll do Agent architecture & orchestration Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes. Build stateful workflows using frameworks like LangGraph and LangChain - including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns. Engineer for long-horizon reliability - multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail. Build the reasoning behind regulated decisions - policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for high-stakes decisions across the industry, from clinical review and prior authorization to claims integrity and care management. Retrieval, grounding & context engineering Develop end-to-end Retrieval-Augmented Generation (RAG) pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies. Engineer memory and context management - conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection. Apply modern context-delivery patterns (e.g., MCP-style tool/context interfaces) so agents access the right information at the right time. Reliability, evaluation & safety Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior. Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions. Evaluate agents at the trajectory and task level - multi-step task success, failure-mode and regression analysis, and sandboxed test environments - alongside retrieval- and generation-quality metrics, automated checks, and human review. Engineer healthcare-grade safety - deployment eval gates, human-oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-aware data handling. Integration & production craft Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate safely within real business workflows. Deliver production-quality code with strong practices in testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, safety, latency, cost, and model risk. Partner with our modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning - through evaluation-driven feedback and, where it helps, fine-tuned or reasoning-optimized models. Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.

Similar jobs

Agentic AI Engineer

CognizantLouisville, KY· 5 days ago
Engineeringapply on careers.cognizant.com