San FranciscoFull TimeEngineering
Remotely
awsdatadogkubernetesazureansibleterraformci/cdcloudwatch
Job Description
📋 Description
- Manage Forge's Site Reliability Engineering team responsible for keeping Forge systems highly
- Drive strong incident management practices in partnership with engineering teams, including
- Build, improve, and manage observability infrastructure in partnership with Platform Engineering
- Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support
- Champion reliability best practices across engineering, including service ownership, operational
- Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
🎯 Requirements
- 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar
- 10+ years of total software engineering, infrastructure, platform, cloud, or production operations
- Bachelor's degree in Computer Science, Engineering, or a closely related field, or equivalent
- Experience building, operating, and maintaining large-scale cloud infrastructure and distributed
- Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting
- Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
Back to all jobs