Machine Learning Engineer, Performance Tooling

Wayve

London; Sunnyvale · Onsite · Full Time

Posted

Job description

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role Wayve is building autonomous driving technology that runs on real vehicles. Getting our models onto embedded hardware — correctly, quickly, and reproducibly — is one of the hardest problems between research and product. As a ML Compiler Engineer, you will own the compilation pipeline that makes that possible. You will build and extend Wayve's ML compiler end-to-end: designing passes, integrating with vendor toolchains like NVIDIA TensorRT and Qualcomm QNN, and delivering deployable bundles that meet our accuracy and latency requirements on every target platform. Each stage in the pipeline — capture, decomposition, precision assignment, legalisation, partitioning — can affect accuracy, latency, or whether a vendor backend accepts the graph. Your work spans the full lowering stack, building compiler passes and infrastructure that scale across architectures and target platforms. Key responsibilities Own the ML compilation pipeline end-to-end — from checkpoint to deployable bundle on NVIDIA (TensorRT) and Qualcomm (QNN) targets. Design and implement compiler passes with accuracy and latency gates, so bad compiles are caught before they reach hardware. Build compilation infrastructure that scales across platforms, model architectures, and SoCs — without re-engineering for each new target. Partner with model and training teams on compilability; build regression and benchmarking to…

Apply for this job