Remotely
Listed on 3 job boards, which is a stronger signal than one board alone.
awsdatadogmysqlkubernetespostgresqlprometheusgrafanards
Also available on
Job Description
📋 Description
- Senior SRE focusing on scalable infrastructure, reliability, and AI-enabled tooling.
- Own architecture, upgrades, and design of distributed systems and microservices.
- Lead capacity planning, stress testing, and capability improvements for growth.
- Define SLAs, alerts, and improve observability and CI/CD across teams.
- Mentor engineers and translate business goals into technical roadmaps.
- Collaborate cross-functionally to improve system reliability at scale.
🎯 Requirements
- 6+ years of software or infrastructure engineering with SRE focus.
- Strong Golang skills and experience with REST APIs.
- Expert in SQL-based RDBMS (MySQL, PostgreSQL) and query optimization.
- Proficient with observability tools (Prometheus, Grafana, Datadog, New Relic).
- Experience with distributed systems design patterns and AI tooling (Claude/LLMs).
- Bachelor’s degree in Computer Science or equivalent.
🎁 Benefits
- Remote-friendly with flexible collaboration hours (EST overlap).
- Competitive compensation and leadership opportunities.
- Open source culture and emphasis on automation and AI enablement.
Back to all jobs