Training / AI Infrastructure — Genesis AI Genesis AI, a London-based company, is hiring a Training/AI Infrastructure engineer to optimize foundation model training, focusing on reducing wall-clock time to convergence and improving GPU utilization. The role requires 8+ years of experience in distributed systems, ML infrastructure, or high-performance computing, with expertise in Python, CUDA, and PyTorch. Training / AI Infrastructure - Salary - Not published - Location - London - Work type - Hybrid - Posted - today Apply on company site opens in new tab https://jobs.ashbyhq.com/genesis/2196f682-bd01-4f83-992d-367bbb3c8e8b/application What You’ll Do - Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels - Design, build, and optimize distributed training systems PyTorch for multi-node GPU clusters, ensuring scalability, robustness, and high utilization - Implement efficient low-level code CUDA, cuDNN, Triton, custom kernels and integrate it seamlessly into high-level training frameworks - Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking - Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures What You’ll Bring - Deep experience in distributed systems, ML infrastructure, or high-performance computing 8+ years - Production-grade expertise in Python - Low-level performance mastery: CUDA/cuDNN/Triton, CPU–GPU interactions, data movement, and kernel optimization - Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism - System-level mindset with a track record of tuning hardware–software interactions for maximum utilization