North AmericaFull TimeEngineering
Remotely
kubernetespytorchjaxvllmtensorflowtensorrtonnxray
Job Description
📋 Description
- Design and build scalable ML services for data extraction, classification, enrichment, and
- Develop end-to-end ML pipelines: data prep, training, eval, deployment, monitoring.
- Deploy/optimize models using ONNX, vLLM, TensorRT for low latency, high throughput.
- Collaborate with software engineers and product teams on data needs and feature engineering.
- Build observability and evaluation systems to ensure model quality and reliability.
- Apply distributed systems and cloud infra to ML services at scale.
🎯 Requirements
- 5+ years in production ML systems.
- Experience with Kubernetes and Ray; ONNX, vLLM, TensorRT or similar.
- Proficiency with PyTorch, TensorFlow, or JAX.
- Designing and scaling ML services for large data with low latency.
- Full ML lifecycle expertise: data prep, feature engineering, training, eval, deploy, monitoring.
- Strong software engineering skills; distributed systems and cloud infra.
🎁 Benefits
- Competitive base salary $150,000 to $215,000, plus equity.
- Health, dental, vision insurance.
- Remote-friendly with access to WeWork locations.
- Unlimited PTO, holidays, year-end closure.
- 401(k) with employer match; stipends for lifestyle/wellbeing.
- Military reserve pay top-up; paid parental leave.
Back to all jobs