Senior Site Reliability Engineer
LodgifyRemotely
pythondatadogkubernetespostgresqlprometheusgrafana
Job Description
📋 Description
- Define meaningful SLIs, SLOs, and reliability targets for the platform.
- Collaborate with software teams to adopt observability best practices.
- Strengthen production readiness with runbooks, ownership, and rollback paths.
- Improve reliability, scalability, and performance of cloud, Kubernetes, and infra.
- Build observability using metrics, logs, traces, and golden signals (Datadog, Prometheus, Grafana).
- Automate repetitive ops tasks with Python or similar languages.
🎯 Requirements
- 7+ years of production experience operating Kubernetes-based platforms and cloud infra.
- Strong SRE practices: SLIs, SLOs, error budgets, incident response, post-incident learning.
- Design observability and alerting for critical systems, plus troubleshooting of distributed systems.
- Ability to write maintainable automation to reduce manual intervention.
- Experience with stateful production systems (databases, caches, queues, streaming platforms).
- Balance reliability, performance, cost, and delivery speed; comfortable in transitional
🎁 Benefits
- Remote flexibility: work from home any day.
- Alan health insurance: premium coverage.
- Meal perk: €150/month allowance and related benefits.
- Tax-free savings: flexible remuneration options up to limits.
- Home office gear provided (table, chair, monitor).
- Language learning: Free Spanish classes.
Back to all jobs