Remotely
gokubernetessparkdistributed systemskafkabigqueryclickhouseflink
Job Description
📋 Description
- Lead reliability initiatives across critical advertising domains, including ad serving, auctions
- Partner with engineering leadership to establish and execute roadmaps focused on reliability
- Design and build scalable platforms, tooling, automation, and infrastructure capabilities that
- Lead architecture reviews and influence technical decisions for high-traffic, revenue-critical
- Establish and monitor reliability metrics and SLOs around critical advertiser and platform
- Participate in on-call rotations, lead complex incident investigations, and coordinate
🎯 Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related
- Proven experience evolving and supporting high-traffic, user-facing production environments with
- Deep expertise in distributed systems, scalability engineering, cloud-native architectures, and
- Strong software engineering capabilities, ideally with experience in backend programming languages
- Extensive knowledge of observability practices and technologies, including metrics, logging
- Demonstrated experience improving reliability through SLOs, automation, incident management
🎁 Benefits
- Base salary range of $217,000–$303,900 USD, with final compensation determined by factors such as
- Eligibility for equity in the form of restricted stock units.
- Comprehensive health benefits, including medical, dental, and vision coverage.
- 401(k) program with employer matching.
- Workspace benefits and support for a home office.
- Personal and professional development funds.
Back to all jobs