North AmericaFull TimeEngineering
Remotely
pythonawskubernetesazuregcpsplunkdynatraceelk stack
Job Description
📋 Description
- Drive operational excellence, reliability, scalability, and performance of critical production
- Apply SRE principles to enterprise-level systems; improve resilience and availability.
- Lead incident response, troubleshoot production issues, and prevent recurrence.
- Design and implement automation to reduce toil and improve efficiency.
- Develop and maintain internal tools and workflows for scalable, reliable infra.
- Work across AWS, GCP, Azure; manage microservices and Kubernetes.
🎯 Requirements
- 8+ years of IT experience, with senior SRE, infra or systems engineering background.
- Deep SRE principles in enterprise-scale environments.
- Strong experience with AWS, GCP, and/or Azure.
- Expertise in microservices and Kubernetes.
- Production-quality coding, Python or similar languages.
- Automation design to reduce manual work and enable self-healing systems.
🎁 Benefits
- Fully remote position within the United States.
- Full-time, direct W-2 employment.
- Opportunity to work on mission-critical enterprise systems and cloud tech.
- Technical ownership and influence over reliability, automation, and ops.
- Collaborate with experienced teams and senior stakeholders.
- Career growth within a technology-focused environment.