Staff ML Performance Engineer (Compiler)
WayveRemotely
pythoncudatensorrtonnxtritonmliropenclqualcomm qnn
Job Description
📋 Description Role focuses on optimizing ML inference for edge accelerators and GPUs. Work on production-ready systems to run large transformer models on low-power edge devices. Collaborate across ML systems, compilers, runtimes, kernels, and embedded deployment. 🎯 Requirements Proven experience improving performance in constrained production systems. Strong proficiency with TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX. Comfort operating from high-level model behavior to low-level kernel/runtime execution. Solid software engineering fundamentals: debugging, profiling, testing, maintainable code. Clear communicator and collaborative teammate able to align stakeholders on trade-offs. 🎁 Benefits Not specified in the listing.
Back to all jobs