Software Engineer, Machine Learning Infrastructure - Generative AI
About The Team
DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI.
About The Role
You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, with a primary focus on our evals and LLM observability platform: the systems that let teams evaluate, trace, and continuously improve the quality of LLM and agent products. You’ll work across evaluation frameworks and SDKs, OpenTelemetry-based trace/score ingestion, LLM-as-judge and offline/online eval pipelines, agent simulations, data pipelines, backend services, and observability. This role is ideal for an engineer who enjoys building reliable measurement and quality primitives in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and evaluation methodologies are evolving quickly.
Requirements
- B.S., M.S., or PhD. in Computer Science or equivalent
- 3+ years of industry experience in software engineering
- Strong backend engineering fundamentals, especially in Python and distributed systems
- Experience building production services, APIs, data pipelines, or ML infrastructure at scale
- Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization
- Hands-on experience with evaluation, LLM observability, or measurement systems for ML/LLM products in production — eval pipelines, tracing/scoring, offline/online quality metrics, or experimentation
- Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software
Qualifications
- Excited About This Opportunity
- Depth in evaluation methodology — LLM-as-judge design and calibration, judge/eval drift detection, human-in-the-loop labeling, or eval harness design for agents and multi-step systems
- Experience with LLM observability and tracing (e.g., OpenTelemetry, trace/score ingestion) and building instrumentation SDKs
- Experience building and deploying AI agents or MCP servers in production, including agent evaluation or simulation
- Experience with data pipelines, streaming ingestion, and analytical stores (e.g., SQL, columnar/OLAP) for high-volume telemetry
- Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
- Experience building developer platforms, internal platforms, or self-serve infrastructure
- Experience with Kubernetes, cloud infrastructure (AWS/GCP), or high-throughput batch systems
- Experience with RAG, search, vector databases, or open-weights LLM inference and fine-tuning
Benefits
We offer a comprehensive benefits package to all regular employees, including a 401(k) plan with employer matching, 16 weeks of paid parental leave, wellness benefits, commuter benefits match, paid time off and paid sick leave in compliance with applicable laws (e.g. Colorado Healthy Families and Workplaces Act). We also provide medical, dental, and vision benefits, 11 paid holidays, disability and basic life insurance, family-forming assistance, and a mental health program, among others.