Principal Engineer, Model Development Platform
About the role
The Principal Engineer for the Model Development Platform at Wayve will own the end-to-end architecture behind Wayve's AI model lifecycle, from data ingestion and training to experiment scheduling and on-road testing. This role involves working at the intersection of AI research, large-scale distributed systems, and robotic operations, ensuring the platform is reliable, scalable, and coherent.
Responsibilities
Design and evolve the platform's overall architecture for reliability, observability, and scalability.
Set performance, latency, and availability targets, and drive the engineering standards to meet them.
Unify the platform across disciplines, from front-end UIs and distributed training to Spark data pipelines and optimization-based experiment scheduling, ensuring systems interoperate cleanly.
Dive into the hardest challenges across subteams, lead architectural reviews, and propose pragmatic solutions that balance innovation with operational simplicity.
Build systems that optimize how models are tested in simulation and on-road, using techniques like linear programming and heuristic optimization to balance hardware, safety, and research priorities while improving throughput and turnaround.
Architect pipelines that ingest, transform, and enrich petabytes of fleet sensor data, and drive efficient compute use across GPU, CPU, cloud, and edge for both prototyping and large-scale training.
Partner with Product, Research, and Operations to align architecture with user needs and co-own the platform's long-term roadmap.
Requirements
10+ years of experience designing and building large-scale distributed systems, ML/AI infrastructure, full stack web application, or developer platforms, including at least 3 years as a staff or principal-level engineer.
Proven ability to design systems spanning web platforms, ML pipelines, and large-scale compute orchestration (e.g., Spark, Ray, Kubernetes, Airflow, MLflow).
Experience driving platform reliability improvements, defining SLAs/SLOs, and building self-healing and observable systems that operate at “four nines” availability or better.
Deep understanding of distributed computing, workflow orchestration, data modeling, and API design, with the ability to write and review production-quality code.
Excellent communication and cross-functional collaboration skills; ability to guide engineers, managers, and researchers toward unified technical direction.
Demonstrated success in mentoring engineers across levels and cultivating a culture of engineering excellence.
Desirable
Experience applying algorithmic or mathematical optimization (e.g., linear programming, graph algorithms) to operational or scheduling problems.
Familiarity with end-to-end model lifecycle tooling, from data ingestion and training CI to model artifact tracking and evaluation workflows.
Prior exposure to autonomous systems, robotics, or other safety-critical domains.
Full-Stack Fluency in modern web frameworks (e.g., React, Flask, FastAPI) and how they integrate into backend systems.
Understanding of data privacy, compliance, and secure handling practices for large-scale sensor data.
What You'll Own
System architecture & reliability
Cross-domain technical leadership
Hands-on problem solving
Experimentation & scheduling systems
Data & compute infrastructure
Strategic collaboration
Essential
Technical Leadership at Scale – 10+ years of experience designing and building large-scale distributed systems, ML/AI infrastructure, full stack web application, or developer platforms, including at least 3 years as a staff or principal-level engineer.
Architectural Depth & Breadth – Proven ability to design systems spanning web platforms, ML pipelines, and large-scale compute orchestration (e.g., Spark, Ray, Kubernetes, Airflow, MLflow).
Reliability and performance – Experience driving platform reliability improvements, defining SLAs/SLOs, and building self-healing and observable systems that operate at “four nines” availability or better.
Hands-On Systems Design – Deep understanding of distributed computing, workflow orchestration, data modeling, and API design, with the ability to write and review production-quality code.
Collaborative Influence – Excellent communication and cross-functional collaboration skills; ability to guide engineers, managers, and researchers toward unified technical direction.
Mentorship & Culture – Demonstrated success in mentoring engineers across levels and cultivating a culture of engineering excellence.
Desirable
Optimization & Scheduling Expertise – Experience applying algorithmic or mathematical optimization (e.g., linear programming, graph algorithms) to operational or scheduling problems.
ML Ops & Experimentation Systems – Familiarity with end-to-end model lifecycle tooling, from data ingestion and training CI to model artifact tracking and evaluation workflows.
Domain Experience – Prior exposure to autonomous systems, robotics, or other safety-critical domains.
Full-Stack Fluency – Experience with modern web frameworks (e.g., React, Flask, FastAPI) and how they integrate into backend systems.
Data Governance – Understanding of data privacy, compliance, and secure handling practices for large-scale sensor data.