Software Engineer, AI Workload Scheduling, AI Infra Supercomputing - Palo Alto [269362]

$140,000 - $252,000 yearly

Job Description

As a Software Engineer within the Supercomputing AI Infrastructure team, you will work on scaling and optimizing our training compute clusters at the core of Robotaxi and the Optimus development. Tesla Supercomputing maintains, on premise, some of the largest supercomputers in the world. We’re currently scaling 100K+ GPU clusters, which are central to developing our autonomy capabilities. Robustly managing such a large system of AI training and HPC clusters requires software design through the stack: from the operating system to workload scheduling to ultimately the training loop.

We are building and improving the workload scheduling systems that govern how hundreds of thousands of GPUs are allocated across training jobs, experiments, and data pipelines. In this role, you will own the Slurm-based scheduler at the heart of our AI HPC compute infrastructure, ensuring that cluster resources are utilized efficiently and that engineers can iterate quickly without friction

The Role

  • Own and evolve the Slurm scheduler configuration, plugins, and policies that govern job scheduling across our GPU supercomputers
  • Design and implement scheduling policies that maximize cluster utilization, minimize job queue times, and enforce fair-share allocation across teams and priorities
  • Build tooling and automation around job submission, preemption, backfill, and resource reservation to support the dynamic needs of AI engineering
  • Develop monitoring and observability infrastructure to provide real-time visibility into cluster utilization, job throughput, and scheduling efficiency
  • Debug and root cause scheduling failures and resource contention issues across thousands of nodes and implement fixes to prevent recurrence
  • Coordinate with the operations team managing the training cluster to maintain high availability and job throughput
  • Work closely with the ML team to understand workload patterns and evolving resource requirements

Requirements

Anthony Antonucci

Qualifications

Relevant experience in Engineering & Information Technology; strong problem-solving skills; commitment to Tesla mission.

Immediate Fill

No