Agentic AI Engineer — Healthcare AI
About the role
This role is part of a new AI-first initiative at Deloitte aimed at rebuilding the decision-making machinery behind American healthcare. The goal is to make care faster, fairer, and less wasteful by leveraging advanced AI technologies.
Responsibilities
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes.
- Build stateful workflows using frameworks such as LangGraph and LangChain - including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability - multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail.
- Build the reasoning behind regulated decisions - policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for high-stakes decisions across the industry, from clinical review and prior authorization to claims integrity and care management.
- Develop end-to-end Retrieval-Augmented Generation (RAG) pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management - conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply modern context-delivery patterns (e.g., MCP-style tool/context interfaces) so agents access the right information at the right time.
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
- Evaluate agents at the trajectory and task level - multi-step task success, failure-mode and regression analysis, and sandboxed test environments - alongside retrieval- and generation-quality metrics, automated checks, and human review.
- Engineer healthcare-grade safety - deployment eval gates, human-oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-aware data handling.
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate safely within real business workflows.
- Deliver production-quality code with strong practices in testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, safety, latency, cost, and model risk.
- Partner with our modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning - through evaluation-driven feedback and, where it helps, fine-tuned or reasoning-optimized models.
- Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.
Requirements
- Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Demonstrated depth building and shipping production agentic systems - this is your primary craft, not a recent exploration.
- Strong, hands-on experience building production agent systems with modern orchestration - LangGraph/LangChain or equivalent, including custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Strong understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep, practical understanding of LLM behavior - strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs - and the evaluation methods used to measure them.
- Experience evaluating and debugging agent behavior - task-success and trajectory analysis, not just output quality.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, and API integration; experience implementing observability, tracing, and debugging for LLM-based systems in production.
- Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
Qualifications
- Experience with multi-agent systems and agent collaboration patterns.
- Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
- Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
- Understanding of traditional NLP concepts: tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
- Experience operating in highly regulated, high-stakes, or operationally complex environments; healthcare exposure - clinical, payer, or life-sciences workflows, or standards such as FHIR - is a plus, not a requirement.
- Demonstrated habit of staying current with AI research, benchmarks, and emerging engineering patterns.
Skills
- Agentic AI Engineering
- Production agentic systems
- LangGraph/LangChain orchestration
- RAG systems
- Memory and context management
- Observability and tracing
- Safety controls and failure-handling
- Model adaptation and fine-tuning
- Traditional NLP concepts
- Highly regulated environments
- Ai research and benchmarks
Benefits
The role offers a competitive compensation package, including a base salary benchmarked to leading technology companies and a substantial performance-based incentive opportunity. The estimated base salary range is $110,700-$372,900, with actual pay depending on skills, experience, and level. Additionally, there is potential for a discretionary annual incentive based on individual and organizational performance.
Pay
Base salary is benchmarked to leading technology companies rather than traditional consulting scales, and the role carries a substantial performance-based incentive opportunity designed to grow with the value you help create - startup-style upside, with the backing of a committed, well-capitalized platform.
Schedule
Travel is expected to be 0-50% based on the work you do and the clients and industries/sectors you serve.