Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
JobgetherRemotely
kuberneteslinuxterraformdrivercudagpunvidiagpu cluster
Job Description
📋 Description
- Deploy and operate GPU cloud infra in datacenters
- Coordinate hardware, networking, storage, and platform software
- Validate BIOS, firmware, NICs, DPUs, GPUs, NVMe, and PCIe topology
- Troubleshoot across hardware and software layers
- Establish deployment standards and runbooks
- Collaborate with engineering, network, and service teams
🎯 Requirements
- Datacenter infra experience in GPU/AI cloud envs
- Bare-metal bring-up and hardware validation
- GPU servers, drivers, NVLink/NVSwitch familiarity
- Linux troubleshooting and kernel/driver experience
- Networking: VLANs, BGP, EVPN, OVS/OVN
- Observability tools: Prometheus, Grafana, Zabbix
🎁 Benefits
- Competitive compensation
- Full-time or contract depending on arrangement
- Europe-based remote work environment
- Hands-on GPU platform exposure and career growth
- International, diverse team
Back to all jobs