Job Description
📋 Description Acknowledge and triage incidents; apply runbooks to restore service quickly and consistently Design and implement metrics and monitors for our distributed app; integrate with the SRE observability framework Develop and improve runbooks for FCM processes; ensure clear, repeatable procedures Understand and maintain the DAG for all FCM processes; ensure correct sequencing and reliable execution Implement robust monitoring and logging with OpenTelemetry Update technical docs for FCM critical paths; set SLA/SLO/SLI targets 🎯 Requirements BS or MS in CS/SE or related field, or equivalent practical experience 2+ years of experience with distributed systems in production Comfort using AI tooling to improve workflows and surface insights Strong understanding of security best practices for backend services Excellent verbal communication with a collaborative, team-first approach 🎁 Benefits Generous PTO 7 Paid Holidays Annually + 5 Conditional Holidays Annually 1 Service Day Annually 401k with 3.5% Company Match Paid Parental Bonding Leave Health, Vision, Dental Coverage