Full Stack Software Engineer, Evaluation Tools
Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. We build intelligent, mapless, and hardware-agnostic AI products for automakers, accelerating the transition from assisted to automated driving.
About the role
The Evaluation Tools group builds the internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behavior. Model developers, researchers, and QA engineers across Wayve depend on our tools to understand driving performance and scale evaluation to millions of scenarios. Every major model release runs through them.
You'll join the Search & Agents squad in Sunnyvale, building search, scenario mining, and agentic tooling that lets anyone find driving scenarios and turn them into tests without writing SQL. This work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat. You'll work across the stack, own features end-to-end, and shape new projects as user needs evolve.
Why this role matters
- Make natural language scenario search reliable and self-service, removing dependency on scarce SQL and mining expertise
- Scale scenario mining from proof-of-concept to production, unlocking certification-grade test creation for teams like Validation
- Extend the evaluation MCP so agents can run search-to-test workflows end-to-end
- Build trust in search and mining results, enabling fast, confident decisions grounded in clear evidence
Responsibilities
- Build full-stack features across search, scenario mining, agentic tooling, and web applications for scenario curation and test suite creation
- Work directly with users to understand pain points, define success metrics (adoption, time saved, reliability), and measure impact
- Ship high-quality, reliable, performant, and maintainable software
- Contribute to architectural decisions that scale across the team’s products
- Use AI tools to improve development speed, team workflows, and product quality
Requirements
- Strong development skills in Python, TypeScript, and JavaScript, with experience using React or similar front-end frameworks
- Strong SQL skills with a solid understanding of database design and optimization
- Experience designing and building production-level full-stack systems and reliable, high-performance APIs
- Track record of building robust, maintainable systems and applying engineering best practices
- Excellent communication and collaboration skills, including working effectively across time zones
- Comfortable working independently in a fast-paced, high-context, ambiguous environment
- Power user of AI development tools with a mindset for building AI-augmented workflows
Qualifications
- Experience with search, embeddings, or retrieval systems
- Experience with the Databricks platform and large-scale data processing with Spark or similar
- Experience with job orchestration frameworks such as Flyte
- Experience building agentic or LLM-powered developer tools (e.g., MCP servers)
- Experience building tools for technical users (ML, robotics, infrastructure)
- Experience with scientific data visualization or complex data interfaces
Pay
Reasonably estimated salary for this role ranges from $209,000 to $266,000, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.
Schedule
This is a full-time, hybrid role based in Sunnyvale, CA. We operate core working hours, allowing flexibility to determine your schedule while fostering collaboration.