Build and scale ML-optimized GPU/TPU superclusters for AI training.
Design and maintain large-scale GPU and TPU clusters using Kubernetes, Python, and Go to support AI model training at production scale. Work with low-level networking technologies like RDMA and NCCL to optimize performance across distributed systems running frameworks such as PyTorch, TensorFlow, and JAX. This senior-level role is based in Canada and available as a permanent, full-time remote position where you'll operate at the infrastructure layer serving machine learning workloads.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 22 July 2026 · InsideJobs links you to the original posting.