Lead Site Reliability Engineer
ZetaJob Description
📋 Description Establish a SRE site and help build an effective, inclusive SRE team. Provide technical leadership for the local team and partner team leads. Guide on availability and performance; build automation to prevent recurrence and automate responses. Manage execution of project priorities, deadlines, and deliverables. Lead Incident Management during incidents. Drive MTTR per the Incident SLA. 🎯 Requirements Experience designing, analyzing, and troubleshooting large-scale distributed systems. Experience with MySQL or PostgreSQL. Hands-on with Kubernetes and cloud platforms. Excellent communication, ownership, and systematic problem-solving. Bachelor’s/Master’s degree in engineering (computer science, information systems). 6-10 years in distributed systems, storage, databases, algorithms, or Unix/Linux internals.