Job Description
📋 Description Build the ML infrastructure platform for researchers and engineers. Automate resource provisioning and scalable GPU/CPU node management. Advance high-performance workload scheduling for large-scale training. Design robust data pipelines for ML data to ML-ready formats. Implement feature caching and storage for low-latency access. Contribute to a unified platform that abstracts cloud infra. 🎯 Requirements 3+ years in ML infrastructure, backend platform, or distributed systems. Terraform, Pulumi or Crossplane for IaC. Experience with Kubernetes, Ray, Slurm, or similar schedulers. Distributed data processing with Spark or Beam. Experience with feature stores and caching (Feast, Redis). Strong systems design, networking, storage knowledge. 🎁 Benefits Competitive base pay: $160,360 - $240,540 Annual bonus, equity, comprehensive benefits Hybrid/remote-friendly policies where applicable