San FranciscoFull TimeEngineering
Remotely
pythonmachine learningdistributed systemsobservabilityml infrastructurefeature store
Job Description
📋 Description
- Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to
- Own the end-to-end feature store lifecycle—from ingestion and transformation to production serving
- Design and implement observability, monitoring, and validation frameworks to detect performance
- Collaborate with cross-functional partners to translate product requirements into robust, scalable
- Automate model deployment and reliability testing to improve developer velocity and ensure system
- Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability
🎯 Requirements
- Deep experience building, deploying, and maintaining production-grade ML infrastructure at scale
- Strong background in distributed systems and backend engineering, with the ability to write robust
- Systematic approach to debugging complex, high-throughput systems and performance bottlenecks.
- Energy to build from 0 to 1 infrastructure that stands the test of time and provides a reliable
- Strong communication skills and ability to create clear documentation for system architectures and
- Growth mindset and attention to detail in code reviews to improve developer velocity.
🎁 Benefits
- Competitive salary with equity opportunities.
- Hybrid work model with in-office days in San Francisco/New York.
- Flexible time off, healthcare, and 401k with matching as part of Patreon’s benefits package.
- Supportive, diverse culture focused on creators and collaboration.
Back to all jobs