Reinforcement Learning Engineer - Manipulation
HumanoidJob Description
📋 Description Train language-vision conditioned manipulation policies via RL in simulation and real world. Construct diverse manipulation task suites in simulation. Partner with teleoperations to collect trajectories for behavior cloning. Partner with testing and operations to establish real-world RL pipelines. Experiment with sim-to-real transfer of policies. 🎯 Requirements 3+ years building deep-learning systems with shipped models or artifacts. Hands-on with LLMs, VLMs, or image/video generative models. Experience solving real problems using RL with deep neural networks. Strong Python + PyTorch/JAX; profiling numerics and maintainable code. Self-driven, proactive, good communicator; document experiments and trade-offs. 🎁 Benefits Competitive equity: stock options with upside as we scale. 30+ paid days off, including 23 days of annual leave and UK holidays. Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer). Free daily breakfast, catered lunch, and snacks in-office.