Also available on
Job Description
📋 Description Provide frontline technical support for AI training, fine-tuning, and inference workloads. Serve as product expert, guiding customers through complex GPU cluster challenges. Collaborate with Customer Experience, Product, and Sales to improve offerings. Be flexible with weekend coverage and on-call hours as needed. 🎯 Requirements 6+ years in a customer-facing technical role (SRE/DevOps/infrastructure) with AI service support Experience with Kubernetes; strong fault isolation and incident handling Proficient in Python, TypeScript and/or JavaScript; REST, curl, Postman Observability tooling knowledge (Prometheus, Grafana) at scale Experience with LLM inference frameworks, LoRA fine-tuning, and HPC storage/compute Cloud experience (AWS, GCP, Azure) and IaC tooling (Ansible, infra-as-code) 🎁 Benefits Remote-friendly with competitive compensation, equity, and benefits Startup environment with opportunities to influence roadmap Flexible remote work options and holiday/weekend coverage as needed