Job

Research Engineer — Large-Scale Inference & RL Systems

Trainety Curated Opportunities

Location
United States
Category
Reinforcement Learning
Industry
Technology & Internet
Organization size
Individual
Updated
September 9, 2026

Description

Serving an ultra-long-context model efficiently creates problems that ordinary model APIs rarely encounter. KV-cache growth, sequence length, GPU memory pressure, batching decisions, rollout duration, networking overhead, and fault recovery all become part of the model's practical performance.


Magic's Inference & RL Systems team works directly on that execution layer.

The engineer will design and scale high-performance inference systems while also supporting the infrastructure used for reinforcement-learning and post-training workflows. That includes optimizing KV-cache management and scheduling, improving throughput and latency, maintaining distributed rollout systems, and making evaluation and reward pipelines resilient enough for repeated large-scale experimentation.

This position is particularly systems-heavy. Debugging may cross GPU execution, memory behavior, networking, storage, schedulers, and distributed services rather than staying within a single ML framework. Engineers are expected to reason explicitly about trade-offs among latency, throughput, reliability, and compute cost.

A strong background could come from large-scale inference, distributed ML training, GPU systems, model-serving infrastructure, or other performance-critical distributed platforms. Deep familiarity with RL algorithms is useful, but the ability to make the systems underneath those algorithms reliable and fast is central to the job.

The team works closely with Magic's kernel and research groups, so infrastructure decisions may directly influence how model architectures and post-training techniques are designed.Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.

https://magic.dev/careers/427ffdee-d4d1-4a39-a730-4a96435daa67

Expertise

  • LLM Inference
  • Distributed Systems
  • GPU Systems
  • KV Cache
  • Post-Training
  • Performance Optimization
  • Model Serving
  • Reinforcement Learning

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.