Staff ML Performance Engineer (Inference Optimisation)
WayveJob Description
📋 Description Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime Implement and validate optimisations in compilers, runtimes, and/or kernels (e.g. operator fusion Build robust benchmarking and regression testing to ensure performance improvements hold across Optimise for multiple targets (e.g. NVIDIA Orin/Thor, Qualcomm) and work with teams to support Collaborate with model developers to influence architecture and training/deployment decisions that Contribute to technical roadmaps and tooling and help raise the standard of performance engineering 🎯 Requirements Proven experience improving performance in production systems with tight constraints (latency Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN Comfort operating at multiple levels of abstraction — from high-level model behaviour down to Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code). Clear communicator and collaborative teammate; able to align multiple stakeholders on performance 🎁 Benefits Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and Experience with NVIDIA and/or Qualcomm SoCs and performance tooling. Python and C++ proficiency. Experience mentoring others and/or driving technical direction in a small, fast-moving team.