Remotely
gokubernetescloud infrastructuredistributed systemsincident managementobservabilityhigh availabilityslos
Job Description
📋 Description
- Lead reliability initiatives across multiple advertising technology domains, including ad serving
- Partner with engineering leadership to define and execute a roadmap focused on reliability
- Design and build platforms, tooling, automation, and engineering solutions that improve system
- Lead architecture reviews and influence technical decisions affecting critical, revenue-generating
- Participate in on-call rotations, lead complex production investigations, and coordinate
- Identify systemic reliability risks and drive durable engineering solutions that strengthen
🎯 Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related
- Strong experience evolving and supporting high-traffic, user-facing production environments.
- Deep expertise in distributed systems, scale engineering, cloud-native architectures, and highly
- Strong software engineering capabilities in a general-purpose backend language such as Go.
- Solid understanding of observability practices and technologies, including metrics, logging
- Proven experience improving reliability through SLOs, automation, incident management, performance
🎁 Benefits
- Comprehensive health benefits, including medical, dental, and vision coverage.
- 401(k) program with employer matching.
- Equity compensation in the form of restricted stock units, subject to the position offered.
- Flexible vacation policy and global company days off.
- 4+ months of paid parental leave.
- Family planning support.
Back to all jobs