Job

Research Engineer — RL Environments, Evals & Post-Training Data

Trainety Curated Opportunities

Location
United States
Category
Reinforcement Learning
Industry
Technology & Internet
Organization size
Individual
Updated
September 9, 2026

Description

Once pre-training ends, improving a model becomes a much more targeted process: identify a weakness, create an environment that exposes it, produce useful training signals, run an intervention, and determine whether behavior actually improved.


That loop is the center of Magic's RL Research & Environments position.

Engineers on this team will build datasets and environments for post-training using approaches such as synthetic generation, targeted collection, and self-play. They will design filtering and scoring systems, define reward signals, maintain evaluation frameworks, and run ablations to understand which data mixtures or training strategies lead to measurable capability gains.

Magic's focus on ultra-long context creates unusually difficult evaluation problems. A model may need to stay coherent through extended trajectories, use information buried far back in context, call tools correctly, and maintain useful behavior across long-running tasks. Traditional short benchmarks may miss many of these failures.

The position therefore combines experimentation with infrastructure. Candidates need enough research judgment to design meaningful tests, but also enough engineering depth to make data and evaluation pipelines reliable and repeatable at scale.

Relevant experience can come from RL, model evaluation, large-scale ML systems, data infrastructure, synthetic-data generation, reward design, or post-training research. Strong attention to data quality is especially important because the feedback loop is only as useful as the signals used to train and measure the model.

Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.

https://magic.dev/careers/b8d6a107-3f98-4341-b114-311c7070e59f

Expertise

  • Reinforcement Learning
  • Model Evaluation
  • Synthetic Data
  • Reward Modeling
  • Post-Training
  • Long-Context Reasoning
  • Self-Play
  • Experiment Design

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.