New YorkFull TimeEngineering
Remotely
goawskubernetesgcpincident managementci/cdprometheusgrafana
Job Description
📋 Description
- Role focused on development, deployment, and monitoring of services that manage bare-metal
- Collaborate with cross-functional teams, vendors, and stakeholders to deliver reliable
- Own incident response, RCA, PIRs, and post-incident improvements to prevent degradation.
- Develop and improve incident response playbooks for a wide range of failure scenarios.
- Improve production services performance, robustness, and observability.
- Contribute to automation, CI/CD pipelines, and scalable tooling for reliability.
🎯 Requirements
- 7+ years in cloud operations, SRE, or related roles.
- Experience with Kubernetes, AWS, GCP; cloud infrastructure knowledge.
- Familiarity with incident management frameworks (ITIL, SRE).
- Proficiency in Go; experience with Prometheus and Grafana.
- Experience deploying containerized apps with Kubernetes; strong documentation and problem-solving
- On-call production support experience.
🎁 Benefits
- Comprehensive benefits package and discretionary bonus/equity.
- 401(k) with employer match; flexible PTO; various health and wellness programs.
- Employee stock purchase program (ESPP) and parental leave.
- Casual work environment; flexible, full-service childcare support; Career growth opportunities.