← All roles

Staff Software Engineer- Foundation Model Inference

Permanentonsite5+ yrs📍 San Francisco, California, 🇺🇸 United States🗣 English

Build LLM inference infrastructure powering enterprise-scale generative AI workloads.

You'll build and optimize LLM inference infrastructure that handles enterprise-scale generative AI workloads, working with PyTorch, MLflow, Ray, vLLM, and SGLang across GPU systems. This is a permanent, full-time senior role based in San Francisco requiring onsite work and English proficiency.

Tech stack
PyTorchMLflowRayvLLMSGLangGPU
Rate
Not stated by the agency
via

Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.

Found at a specialist agency · listed 14 August 2026 · InsideJobs links you to the original posting.