Senior Site Reliability Engineer
JobgetherNorth AmericaFull TimeEngineering
Remotely
awskubernetesterraformprometheusgrafanaeksnew reliccloudwatch
Job Description
📋 Description
- Define, measure, and continuously improve service reliability through SLIs, SLOs, error budgets
- Build and maintain observability across infrastructure and applications with metrics, logs, traces
- Participate in production incident response, troubleshooting, escalation, root-cause analysis
- Improve operational excellence through production readiness reviews, runbooks, documentation
- Optimize infrastructure for reliability and performance while managing cloud consumption and FinOps
- Analyze system performance, resource utilization, latency, throughput, and growth trends to
🎯 Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud
- Strong hands-on experience designing and operating production workloads in AWS, including
- Advanced Terraform expertise, including reusable modules, remote state, dependency management
- Strong experience with observability and production telemetry, using Prometheus, Grafana
- Practical understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning
- Deep Kubernetes knowledge covering EKS, Helm, workload scheduling, networking, storage
🎁 Benefits
- Flexible remote or hybrid work options.
- Comprehensive health and wellness benefits, including medical, dental, and vision coverage with
- Generous paid time off and paid holidays.
- MyShare Employee Ownership Program.
- 401(k) plan with up to a 4% employer match, with full vesting from day one.
- Opportunities to work alongside industry leaders and experienced engineering professionals.
Back to all jobs