Remotely
awsdatadogkubernetessreobservabilityappdynamicsdynatraceopentelemetry
Job Description
📋 Description
- Lead observability, reliability, and operational excellence across cloud-native environments.
- Define SRE standards, governance, and monitoring strategies for distributed systems.
- Partner with product and engineering to improve system reliability and performance.
- Design and implement monitoring, alerting, dashboards, and reporting solutions.
- Collaborate across distributed teams in a multilingual and global context.
- Drive technical initiatives, prioritization, and timely delivery within budgets.
🎯 Requirements
- Significant experience in Site Reliability Engineering (SRE) in large-scale production.
- Strong expertise in observability platforms, SLIs/SLOs, and alerting governance.
- Hands-on experience with Dynatrace, OpenTelemetry, and distributed tracing.
- Proficient with Dynatrace, Datadog, New Relic, AppDynamics, and related APM/tools.
- Experience with cloud and container tech (AWS, Kubernetes), and monitoring tooling (RUM, APM).
- Strong collaboration, stakeholder management, and communication in English; French/English
🎁 Benefits
- Competitive salary: CAD 100,000 - 150,000 CAD annually.
- Comprehensive benefits after three months: health insurance, disability coverage, EAP, and mental
- Personal tech reimbursement, DPSP, flexible vacation, and retirement matching.
- Distributed team with remote-friendly setup and a borderless global framework.
Back to all jobs