Senior Site Reliability Engineer
JobgetherNorth AmericaFull TimeEngineering
Remotely
pythonaiprometheusbgpgrafanaopentelemetryipv6loki
Job Description
📋 Description
- Develop and scale robust Python-based tooling, infrastructure-as-code utilities, and automation
- Build automated workflows and API integrations across corporate ticketing systems to accelerate
- Apply modern AI and LLM-based development tools to improve technical execution, automate scripting
- Work with advanced private cloud and compute technologies to improve availability, latency
- Design and implement telemetry pipelines, Prometheus and Grafana dashboards, and AI-driven anomaly
- Define operational KPIs, monitoring standards, telemetry baselines, alerting thresholds, and
🎯 Requirements
- 5+ years of relevant Site Reliability Engineering, infrastructure engineering, systems engineering
- Exceptional proficiency in Python and experience developing scalable operational tooling, API
- Hands-on experience with modern observability technologies such as Prometheus, Grafana
- Strong understanding of advanced networking concepts, including high-bandwidth routing and
- Experience designing and launching new services with clear operational readiness requirements
- Extensive experience creating technical runbooks, leading complex incident response processes, and
🎁 Benefits
- Comprehensive benefits designed to support employee health, well-being, financial security, and
- Flexible working options that allow employees to work from home, in an office, or through a
- Opportunity to work on cutting-edge AI hardware, private cloud, distributed infrastructure, and
- Exposure to large-scale, business-critical systems serving global digital experiences.
- Collaborative environment with opportunities to work alongside experienced infrastructure
- Opportunities to develop expertise in SRE, observability, automation, AI-assisted engineering
Back to all jobs