Remotely
pythonkubernetesterraformgrafanatrinoapache iceberghive metastorenessie
Job Description
📋 Description
- Own platform health and reliability for Trino, Lightdash, Coder, Jupyter, Redash, Hive Metastore
- Define SLOs for availability, performance, and workspace provisioning.
- Manage end-to-end Trino cluster operations and optimizations.
- Lead platform upgrades, patch cadence, and security fixes.
- Develop runbooks, on-call processes, and incident handling.
- Ensure security, access controls, and audit logging across services.
🎯 Requirements
- Experience integrating AI into workflows or decision making.
- 12+ years in software/platform engineering; 5+ years in management with production platforms; or
- Proven open-source data infra leadership in production.
- Hands-on with distributed query engines and clustering.
- Strong Kubernetes/container deployment and debugging skills.
- Establishing SLOs, observability, and incident response maturity.
🎁 Benefits
- Open-source stack with architectural influence.
- Culture of craftsmanship, quality, and innovation.
- AI/automation tools to enhance engineering excellence.
- Equity where applicable, health plans, 401(k), and family leave programs.
Back to all jobs