Staff Machine Learning Engineer, Core Services Eng (GenAI)
Uber · San Francisco, CA · 1 mo ago
Information Technology$232k/yrFull-time
About The Role
Uber's Customer Obsession team builds the platform and AI that powers world-class support across mobile, web, and voice at global scale. We are now hiring a Staff ML Engineer to architect, productionize, and scale an autonomous support agent that resolves customer issues end-to-end. Experience with agentic architectures is a major plus.
What You Will Do
- Own the end-to-end agent architecture: agentic planning and execution loops, long-term memory, persona/voice, knowledge routing, and policy enforcement for compliant, on-brand conversations.
- Advance retrieval & reasoning: Build next-generation retrieval and reasoning pipelines, where the agent can search across different knowledge sources, apply policy-driven tools, and call structured workflows and ensure that responses are consistently grounded.
- Establish evals that matter: offline rubrics, simulated scenarios, safety tests, cost/latency tradeoff suites, and LLM-as-judge (with calibrated human review) wired into CI/CD and experiment platforms.
- Drive automation at scale: partner with Product/Design/Operations on coverage, policy alignment, localization, and rollout strategy to better customer experience and reduce cost per contact.
- Mentor/principle-lead multiple pods; set technical strategy and quality bars; coach senior engineers on agentic patterns, reliability, and experiment velocity.
Basic Qualifications
- 7+ years building production ML/AI systems;
- 2+ years leading complex ML initiatives end-to-end.
- Deep expertise in LLM-driven systems (inference optimization, prompt/program design, fine-tuning, distillation/LoRA, safety/guardrails, evals).
- Track record of shipping customer-facing intelligent experiences with measurable impact (A/B testing, metrics literacy).
- Bachelor's Degree, or above, in Comp Science or related field.
Preferred Qualifications
- Agentic architectures in production (planner/executor, memory, multi-step reasoning) and RAG over complex, policy-heavy knowledge bases.
- Experience building support automation for large consumer platforms (routing, policy codification, internal tooling, co-pilot/auto-resolve).
- Multilingual NLU/NLG (code-switching, low-resource languages), hallucination mitigation, safety red-teaming, and privacy-by-design.
- PRACTICAL expertise balancing speed and reliability at scale: experiment frameworks, feature flags, canary/guarded rollouts, and clear kill-switches.