Job Description
📋 Description Design and implement kernels for linear algebra and tensor ops (GEMM, batched GEMM, convolutions Own performance and correctness - add microbenchmarks, regression tests, numerics validation Profile and optimise kernel launches for next-gen AI hardware Profile architectures, memory layout, threading, and cache locality Integrate with native code and reading/extending kernels for ML frameworks Mentor teammates and share knowledge within the ML Kernels & Runtime team 🎯 Requirements Excellent programming and scripting skills using C++ and Python Hands-on with BLAS/DNN stacks and kernel development Understanding of processor architectures, Linux profiling, and performance tooling Experience with PyTorch or similar custom ops/extensions is a plus Strong communication, collaboration, and problem-solving abilities 🎁 Benefits Competitive salary and annual leave policy Medical and dental health plans, gym card, pension (matched up to 4%) Inclusive, flexible interview process and reasonable adjustment options