Job Description
📋 Description Lead complex troubleshooting and incident resolution from investigation to permanent solutions. Own production issues across distributed systems, infra, and app layers. Design and improve monitoring processes, playbooks, runbooks, and knowledge resources. Leverage AI-powered tools and automation to speed up detection and resolution. Collaborate with engineering, customer experience, and clients to resolve challenges. Develop and maintain automation scripts to reduce manual work. 🎯 Requirements 3–5 years of experience in technical support, SRE, Forward Deployed Engineering, or full-stack Strong experience troubleshooting distributed systems and production environments. Expertise with observability and monitoring platforms such as DataDog, ELK Stack, Prometheus Knowledge of IaC, Docker/Kubernetes, cloud-native architectures, and CI/CD pipelines. Experience with AWS, databases, and system monitoring. Strong English communication and collaboration skills with global teams. 🎁 Benefits Competitive compensation based on experience, skills, and expertise. Total compensation around USD 72K–90K plus equity. Equity participation to share in company growth. Unlimited paid time off (PTO) with flexibility. Annual learning and development stipend for growth. Opportunity to work remotely from Brazil.