Site Reliability Engineer
YunoRemotely
awsdatadogdockerkubernetespostgresqlterraformkafkapulumi
Also available on
Job Description
📋 Description
- Remote, full-time, individual contributor, +7 years experience.
- Own reliability strategy, architecture, and standards across infra.
- Lead event-driven messaging and scalable platform design.
- Own IaC provisioning and cloud deployment at scale.
- Build observability, dashboards, alerts for 24/7 ops.
- Mentor senior/mid engineers and lead incident reviews.
🎯 Requirements
- Event-driven architecture with Kafka/NATS/RabbitMQ; at-least-once.
- Advanced AWS: EC2, VPC, IAM, S3, RDS; internal VPC networking.
- Infrastructure as Code: Terraform or Pulumi.
- Kubernetes and Docker in production; lifecycle and health checks.
- Observability and SLOs; Datadog dashboards/monitors; tracing.
- Distributed systems debugging; Go or Python; leadership & English.
🎁 Benefits
- Remote work from anywhere
- Home Office Bonus
- Work equipment provided
- Stock options
- Health plan worldwide
- Flexible days off