North AmericaFull TimeEngineering
Remotely
pythonjavadistributed systemsobservabilitytrainingscalaml
Job Description
📋 Description
- Design, build, and operate platform subsystems for observability, evaluation, tooling, and
- Prove platform capabilities against production ops, incl. anomaly detection and automated workflows.
- Build observability infra for model behavior, training health, serving latency, and data quality.
- Identify cost optimization across ML training/serving with automated tooling.
- Architect reliability improvements across AI/ML stack for production excellence.
- Contribute to modernized ML platforms and manage tech dependencies with partner teams.
🎯 Requirements
- Significant experience designing, building, operating production AI/ML systems at scale (training
- Hands-on infra for advanced agentic/complex model architectures (memory, trace, evaluation, replay
- Strong software fundamentals with deep Python and working proficiency in Java or Scala.
- Proven ability to improve reliability, scalability, and cost efficiency of AI/ML infra.
- Background building observability/monitoring for ML workloads across training, serving, data
- Deep distributed-systems expertise (batch and real-time serving).
🎁 Benefits
- Annual salary range of $600,000–$1,066,000, with compensation based on
- Compensation structure with salary and stock options, adjustable annually.
- Comprehensive health plans and mental health support.
- 401(k) with employer matching.
- Stock option program.
- Health Savings Accounts and Flexible Spending Accounts.
Back to all jobs