Job Description
📋 Description Define customer-meaningful SLOs and set error budgets for broker flows (order execution and market Develop reliability standards, including SLO methods, error-budget policy, observability guide, and Contribute reliability patterns (circuit breakers, retries, bulkheads, load-shedding) into Ruby Extend observability stack (Prometheus, Honeycomb, OpenTelemetry) as workloads scale on Nomad. Design and run tabletop exercises and fault-injection tests to stress platforms against real-world Mentor engineers to build a lasting site reliability culture across teams. 🎯 Requirements Production-quality coding in Ruby and/or Java, plus Python for automation. Experience embedding SRE practices (SLOs, error budgets, burn-rate alerting). Hands-on OpenTelemetry, Prometheus, Grafana instrumentation. Strong Linux internals and networking (TCP/IP, UDP/multicast, packet capture). On-call production experience and blameless post-incident reviews. Experience influencing standards; HashiCorp Nomad/Consul/Vault is a strong plus. 🎁 Benefits Performance bonuses Stock purchase options Medical/Vision/Dental benefits 401k plan Paid vacation and sick days Gym membership reimbursement