Senior AI Infrastructure Engineer - Model Training
Kodiak
<div class="content-intro"><p>Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.</p></div><p><span style="font-size: 12pt;">Kodiak's AI is only as good as the speed at which we can train it. Every improvement to our models – from GigaFusionNet to large-scale world models – depends on infrastructure that turns thousands of hours of multimodal driving data into training throughput. We are looking for engineers who make model training fast: streaming massive camera, LiDAR, and radar datasets without stalling a single GPU, sharding data and models efficiently across nodes, and extracting every FLOP from the latest hardware. If you measure your impact in tokens per second and GPU utilization, this role is for you.</span><br><br><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>In this role, you will:</strong></span></p> <ul data-list-tree="true" data-indent="0" data-border="0"> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Design high-throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturated</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Build and optimize distributed training infrastructure across multi-node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single device</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Maximize utilization of modern accelerators such as NVIDIA B200s through mixed-precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlap</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Profile end-to-end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU compute</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Develop scalable dataset construction pipelines that convert petabytes of raw driving logs into training-ready, streamable formats</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Partner with ML teams to scale new architectures from prototype to full-cluster training runs efficiently and reliably</span></li> </ul> <div><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>What you’ll bring:</strong></span></div> <ul data-list-tree="true" data-indent="0" data-border="0"> <li style="font-size: 12pt;"><span style="font-size: 12pt;">BS, MS, or PhD in Computer Science or a related field, and at least 2-3 years of industry experience in ML systems or infrastructure</span></li> <li style="font-size: 12pt;"><span style="font-size: 1