Optimize large AI models for high-volume, low-latency production inference.
You'll optimize large AI models to run efficiently in high-volume, low-latency production environments using PyTorch, CUDA, and distributed systems technologies like NCCL, MPI, and InfiniBand across GPU and Azure infrastructure. This senior-level permanent role requires full-time onsite work in San Francisco and demands fluency in English. The position focuses on model inference performance at scale, working with advanced GPU acceleration and networking technologies.
Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.
Found at a specialist agency · listed 27 July 2026 · InsideJobs links you to the original posting.