← All roles

Software Engineer, Model Inference

Permanentonsite5+ yrs📍 San Francisco, 🇺🇸 United States🗣 English

Optimize large AI models for high-volume, low-latency production inference.

You'll optimize large AI models to run efficiently in high-volume, low-latency production environments using PyTorch, CUDA, and distributed systems technologies like NCCL, MPI, and InfiniBand across GPU and Azure infrastructure. This senior-level permanent role requires full-time onsite work in San Francisco and demands fluency in English. The position focuses on model inference performance at scale, working with advanced GPU acceleration and networking technologies.

Tech stack
PyTorchCUDANCCLAzureGPUInfiniBandMPINVLink
Rate
Not stated by the agency
via

Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.

Found at a specialist agency · listed 27 July 2026 · InsideJobs links you to the original posting.