Job Description
📋 Description Own reliability, scalability, and observability of financial SaaS apps and infra Design, implement, and maintain SLOs/SLIs across critical systems Lead observability strategy with monitoring, logging, and tracing Build/run incident procedures and post-incident reviews; mentor team Architect/deploy cloud infra on AWS or Azure; implement IaC for high availability Develop automation and AIOps to reduce toil and enable self-healing 🎯 Requirements 7+ years in SRE/DevOps/platform engineering with production responsibility Expert-level AWS/Azure experience; scalable infra knowledge Observability expertise with Prometheus, Grafana, ELK, Datadog, etc. Strong SLO/SLA knowledge; handling error budgets Distributed systems knowledge; resilient, highly available design Python/PowerShell/Bash tooling for automation 🎁 Benefits Salary: $104,148 – $177,600 per year Remote work option; US-based or global remote considered Comprehensive benefits package (health, vision, dental)