Senior Principal Site Reliability Engineer
BybitJob Description
📋 Description Design and build an enterprise-grade chaos platform across multi-cluster (K8s + EC2) and Develop core chaos capabilities: fault injection at pod/node/AZ level Support latency, packet loss, partitions, and dependency timeout injection Production safety: blast radius control, one-click Kill Switch, auto rollback Integrate fault scenarios with monitoring and SLOs for automated feedback loop Enable traffic and fault isolation for experiments 🎯 Requirements 8+ years backend/infrastructure experience; 3+ yrs chaos/stability engineering Hands-on with large-scale production fault injection under safety constraints Expertise in Kubernetes fault injection (Chaos Mesh / Litmus / custom) Go or another backend language; system architecture design Deep understanding of distributed failure modes Observability stack: Prometheus, Grafana, Thanos, OpenTelemetry 🎁 Benefits Study Growth Fund for professional development Internal events and team-building activities Global collaboration with international colleagues Career advancement opportunities Internal mobility for long-term growth