← All roles

Software Engineer, Hardware Health

Permanentonsite5+ yrs📍 San Francisco, 🇺🇸 United States🗣 English

Build infrastructure to monitor and manage OpenAI's global GPU compute fleet health.

Build infrastructure to monitor and manage GPU compute fleet health across global operations. You'll work with Python, SQL, PromQL, and Linux to develop monitoring and management systems for GPU and InfiniBand hardware. This is a senior-level, full-time onsite position in San Francisco requiring fluency in English.

Tech stack
PythonSQLPromQLLinuxGPUInfiniBand
Rate
Not stated by the agency
via

Membership is €29/month, cancel anytime: every rate, every original listing link, and a daily alert for roles matching your filters.

Found at a specialist agency · listed 27 July 2026 · InsideJobs links you to the original posting.