Full Stack Software Engineer, Evaluation Tools
Wayve
Sunnyvale · Onsite · Full Time
Posted
Job description
About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role The Evaluation Tools group builds the internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behaviour. Model developers, researchers and QA engineers across Wayve depend on our tools to understand driving performance and scale evaluation to millions of scenarios, and every major model release runs through them. You'll join the Search & Agents squad in Sunnyvale. We build the search, scenario mining and agentic tooling that lets anyone find the right driving scenarios and turn them into tests without writing SQL. Our work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat. You'll work across the stack and own features end-to-end: talking to users, shaping ideas, building robust software, and validating impact. Evaluation is an evolving challenge, so you'll also have plenty of opportunity to define new projects as user needs emerge. Why this role matters Make natural language scenario search reliable and self-service, so finding test data no longer depends on scarce SQL and mining expertise Take scenario mining from proof-of-concept to production scale, unlocking certification-grade test creation for teams like Validation Extend the evaluation MCP so agents can run search-to-test workflows end to end Build…