Job

Kernel Engineer — High-Performance AI Training & Inference

Trainety Curated Opportunities

Location
Remote / Worldwide
Industry
Technology & Internet
Organization size
Individual
Updated
September 9, 2026

Description

For Magic's long-context models, some of the most important performance improvements happen below the model framework itself.


The Kernels team focuses on the low-level operations that determine how efficiently training and inference use modern AI accelerators. Memory movement, kernel fusion, communication, sequence length, numerical correctness, and hardware utilization can all materially affect whether a new model architecture is practical at scale.

This engineer will design, implement, benchmark, and maintain high-performance kernels used across Magic's training, inference, and reinforcement-learning systems. Performance matters, but so do robustness and correctness: optimized kernels eventually become production infrastructure and need extensive testing rather than existing only as research benchmarks.

The technical environment may involve NVIDIA architectures such as Blackwell as well as alternative accelerators. Relevant frameworks and libraries include Triton, CUTLASS, CuTeDSL, NCCL, MSCCL++, FlashAttention-style implementations, Pallas/Mosaic, Mojo, and other low-level kernel or accelerator programming systems.

Magic favors deep expertise here. Someone who understands GPU architecture, memory hierarchies, compiler behavior, synchronization, communication primitives, and low-level optimization can be a better fit than an ML generalist with shallow accelerator knowledge.

The engineer will also co-design solutions with training, inference, and RL teams, making this an opportunity to influence the interaction between model architecture and the hardware-level implementation beneath it.

Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.

https://magic.dev/careers/42010571-7943-43d5-b182-c77d8342b088

Expertise

  • GPU Kernels
  • CUDA
  • Triton
  • CUTLASS
  • FlashAttention
  • AI Accelerators
  • Performance Engineering
  • Computer Architecture

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.