Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
JobgetherRemotely
dockerkuberneteslinuxansibleterraformcudagpunvidia
Job Description
📋 Description
- Deploy, commission, and operate GPU cloud infrastructure across regional/core datacenters.
- Transform validated architectures and BOMs into production-ready hardware, networking, storage
- Own intersection of datacenter operations, GPU infra, network engineering, and cloud platform ops.
- Troubleshoot complex issues across physical and software layers; drive incidents to resolution.
- Establish deployment standards, validation procedures, and operational documentation.
🎯 Requirements
- Datacenter infrastructure hands-on experience in GPU, HPC, AI cloud, or high-density compute.
- Bare-metal deployment from physical install to production readiness.
- Experience with NVIDIA GPU servers, drivers, PCIe topology, and high-performance compute.
- Familiarity with NVL/NVLink, in-rack networking, and OEM/NVIDIA validation needs.
- Linux troubleshooting, virtualization and containers (KVM/QEMU, Docker, Kubernetes, VFIO).
- Networking knowledge: VLANs/VRFs, BGP/ECMP, OVS/OVN, OOB management, high-speed datacenter nets.
🎁 Benefits
- Attractive compensation; full-time or contract engagement.
- Europe-based remote working environment with flexibility.
- Opportunity to work on cutting-edge GPU cloud and AI infra at scale.
- Hands-on exposure to NVIDIA GPUs, RoCE/RDMA, Kubernetes, storage, and observability.
Back to all jobs