Senior Software Engineer, SRE / Observability Tooling
GliaJob Description
📋 Description Develop dashboards, alerts, and monitors as code. Partner with development teams to ensure production readiness. Build tooling to define, measure, and report on SLOs/SLIs. Automate observability and operational workflows to reduce toil. Improve incident response tooling and workflows. 🎯 Requirements Expert AWS and Kubernetes (EKS) for observability, networking, auto-scaling. Experience with observability platforms (DataDog, Prometheus) and metrics/logs/traces. Deep understanding of SRE principles: SLOs, error budgets, toil reduction. Proven experience analyzing and troubleshooting large-scale distributed systems. Strong Python or Go skills for building operational tools. CI/CD pipelines for microservices (ArgoCD, Github Actions, Helm). Data-driven problem-solving and root cause analysis. Thoughtful use of AI tools with ownership of output. 🎁 Benefits Remote-first with teams across Canada, Portugal, Poland, Estonia. Optional offices in Tallinn and Tartu, Estonia. Flexible remote collaboration with biannual Estonia gatherings. Equal-opportunity employer and inclusive culture. Collaborative, cross-country engineering community.