$118,000 - $390,000 yearly
Tesla
Palo Alto, CA
As a Software Engineer within the Autopilot AI Infrastructure team, you will work on reinforcing, optimizing, and scaling our infrastructure components supporting AI research activities for Autopilot and the Optimus.
At the core of our autonomy capabilities are neural networks that the research team is designing to train on very large amounts of data, across large-scale GPU clusters. Robustly training these models at scale and in the shortest amount of time is critical to our mission.
The Role
Reduce wall clock time to convergence of our training jobs by identifying bottlenecks in the ML stack, from data-loading up to the GPU
Integrate efficient, low-level code with the overall high-level training framework
Profile our workloads and implement solutions to increase training efficiency
Optimize workloads for efficient hardware utilization (e.g. CPU and GPU compute, data throughput, networking)
Requirements
Members of the Autopilot AI Infrastructure...