New York, SunnyvaleFull TimeEngineering
Remotely
gorestkubernetesci/cdprometheusgrafanagrpcpromql
Job Description
📋 Description
- Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data
- Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and
- Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health
- Improve observability, alerting, and operational tooling for quick issue detection and resolution.
- Translate incidents and hardware failure modes into software improvements for increased resilience.
- Collaborate with hardware-adjacent, infrastructure, operations, and software teams to design
🎯 Requirements
- B.S., M.S., or PhD in Computer Science or related field, or equivalent experience.
- 8+ years of software engineering experience in infrastructure, cloud, and distributed systems.
- Expertise in Go and building REST/gRPC APIs for mission-critical platforms.
- Strong background in cloud-native Kubernetes infrastructure and distributed services.
- Experience mentoring engineers and leading technical projects.
- Proven ability to work with open source communities and apply data-driven reliability improvements.
🎁 Benefits
- Medical, dental, and vision insurance (US-based offering for full-time employees).
- 401(k) with generous employer match and ESPP.
- Flexible PTO and parental leave options.
- Tuition reimbursement and comprehensive wellness benefits.
- On-site facilities and data center perks; a casual, innovative work culture.
- Career development opportunities and exposure to cutting-edge AI infrastructure.