Site Reliability Engineer II - AI Infrastructure (f/m/d)
FocusedRemotely
postgresqlazurelinuxterraformnetworkinggithub actionsgitlab ciopentofu
Job Description
📋 Description
- Own the reliability and day-to-day operation of multiple internal applications, deployment
- Lead the investigation and resolution of incidents and deployment failures across cloud services
- Drive automation improvements across build, release, deployment, monitoring, alerting, backup
- Implement safe, well-tested code, configuration, infrastructure, and database fixes to restore or
- Develop proposals for cloud, hosting, database, and runtime migrations, as well as horizontal
- Build and maintain clear operational documentation, including runbooks, recovery procedures
🎯 Requirements
- Strong experience with Linux, networking, structured troubleshooting, and cloud hosting concepts.
- Experience operating internal applications and deployment platforms in Azure or a comparable cloud
- Practical knowledge of relational databases, particularly PostgreSQL or an equivalent platform
- Experience with infrastructure as code and CI/CD tools such as OpenTofu, Terraform, GitLab CI
- Knowledge of monitoring, alerting, backup, recovery, rollback automation, incident response, and
- Experience with horizontal scaling, stateless system design, migrations, and reversible change
🎁 Benefits
- Focused Energy is an equal opportunity employer and committed to creating an inclusive environment.
- Compensation details will be determined by location, level, and experience; incentives and equity
Back to all jobs