Description
A frontier model training run may use thousands of accelerators for long periods of time, which means small inefficiencies become enormous compute costs and small reliability problems can destroy valuable training progress.
Magic's Pre-training Systems team is responsible for preventing those failures while pushing training throughput as high as possible.
The work spans data, tensor, and pipeline parallelism; communication and gradient synchronization; checkpoint design; fault recovery; orchestration; storage; networking; and profiling across multi-node GPU environments. Long-context models add additional pressure because large sequences increase memory consumption and make efficient sequence packing and communication even more important.
Rather than simply operating an existing training platform, the engineer is expected to improve it. That could mean finding bottlenecks across compute and networking, increasing accelerator utilization, reducing checkpoint overhead, improving reproducibility, or redesigning parts of distributed execution to support new model architectures.
Candidates should understand distributed systems as well as large-model training. Practical experience with multi-node GPU workloads and the trade-offs among different parallelism strategies is especially relevant. The ability to diagnose problems that cross ML frameworks, communication libraries, hardware, and infrastructure will matter frequently.
The role works closely with Magic's research and kernel engineers. Changes to the model may require systems work, while systems constraints may inform how the model itself is designed.
Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.
https://magic.dev/careers/f1d3988f-f93c-42b7-ad1a-f9fb3d07ff26