Remotely
pythondatareinforcement learninglarge language modelspytorchtensorflowrlhfdpogrpo
Job Description
📋 Description
- Design and implement end-to-end RL systems that combine model-based RL, offline RL, simulation, and
- Build latent world models and marketplace state representations that capture supply-demand
- Develop systems that optimize across multiple marketplace levers simultaneously—pricing
- Create policy evaluation frameworks and establish monitoring systems for safe deployment of new
- Fine-tune open-source large language models on domain-specific data for prediction, reasoning, and
- Design and implement training strategies for language models, including supervised fine-tuning
🎯 Requirements
- PhD in Computer Science, Operations Research, Applied Mathematics, or related field with at least
- Proficient in RL fundamentals (Markov Decision Processes, stochastic control, reward design)
- Experience building production ML/RL systems with online learning or simulation-based optimization
- Knowledge in world models and sequential modelling (RNNs, transformers, state-space models)
- Hands-on experience fine-tuning large language models in production (SFT, RLHF, DPO, or GRPO; LoRA
- Ability to design evaluation frameworks for generative models (factual accuracy, reasoning quality)
🎁 Benefits
- Term Life Insurance
- Comprehensive Medical Insurance
- GrabFlex: create a benefits package that suits your needs and aspirations
- Parental and Birthday leave
- Love-all-Serve-all (LASA) volunteering leave
- Confidential Grabber Assistance Programme
Back to all jobs