Job Description
📋 Description Design and maintain inference infrastructure for generative audio models. Implement and manage high-performance inference engines. Orchestrate deployments with Kubernetes and autoscaling. Develop and automate CI/CD pipelines for model artifacts. Monitor production systems and track latency, resources, performance. Collaborate with research teams to optimize serving paths. 🎯 Requirements Deep understanding of audio model architectures (TTS/ASR). Hands-on Kubernetes (K8S) and autoscaling for production workloads. Strong MLOps background with CI/CD and scalable cloud infra. Proficiency in Python or Go for tooling and backend services. Experience with GPU-accelerated inference and profiling. Familiarity with high-performance inference engines (e.g., Triton, vLLM-Omni). 🎁 Benefits Competitive salary and equity Medical, dental, and vision insurance 42 days PTO: 15 PTO, 10 sick, 15 holidays, 2 floating Parental leave and fertility support 401(k) retirement plan Lifestyle spending account: $500/month