About this role
Phone numbers and emails in this ad are masked until you log in.
auto_translated_note
What You’ll DoDrive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernelsDesign, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilizationImplement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworksOptimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networkingDevelop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failuresWhat You’ll BringDeep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)Production-grade expertise in PythonLow-level performance mastery: CUDA/cuDNN/Triton, CPU - GPU interactions, data movement, and kernel optimizationScaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelismSystem-level mindset with a track record of tuning hardware - software interactions for maximum utilizationFind more English Speaking Jobs in United Kingdom on Arbeitnow
Community Q&A
Anyone worked here? Ask before you apply.
No threads yet for this job or company.