Staff engineer building AI research infrastructure for large-scale training and inference workloads.
Design and operate distributed systems that manage GPU-intensive ML experiments across thousands of machines, from cluster scheduling to job orchestration and monitoring. Partner with researchers and ML engineers to turn experimental workloads into production pipelines while building tooling for rapid iteration. This role requires deep expertise in large-scale backend services, cluster management systems, and modern ML training workflows, with a focus on translating research needs into reliable infrastructure.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 21 August 2026 · InsideJobs links you to the original posting.