Site Reliability Engineer
JobgetherRemotely
awsdatadogdockerkubernetesterraformpulumislosli
Job Description
📋 Description
- Define and implement reliability strategy across the platform (SLOs/SLIs, incident practices
- Drive architectural decisions to keep systems scalable, resilient, and observable as AI workloads
- Design and own event-driven messaging, transitioning from synchronous to durable asynchronous
- Manage AWS infrastructure with IaC for provisioning, configuration, deployment, and operations.
- Ensure Kubernetes and containerized workloads scale reliably with increasing load.
- Build and maintain observability with monitoring, dashboards, APM, and distributed tracing.
🎯 Requirements
- Extensive SRE/Platform Engineering/DevOps experience with production-scale systems ownership.
- Deep expertise in event-driven architectures and messaging (Kafka, NATS, RabbitMQ) and related
- Strong AWS expertise (EC2, VPC, IAM, S3, RDS) plus solid networking fundamentals.
- Hands-on IaC experience with Terraform, Pulumi, or similar tools with Git workflows.
- Strong Kubernetes and Docker production experience (lifecycle, resource limits, health checks
- Observability expertise with Datadog or equivalent (dashboards, monitoring, APM, tracing, alerting).
🎁 Benefits
- Competitive compensation.
- Fully remote with flexible locations.
- One-time home office allowance.
- Company-provided equipment.
- Stock options.
- Health plan available worldwide.
Back to all jobs