Job

Machine Learning Research Scientist — Frontier Model Evaluation & Failure Analysis

Trainety Curated Opportunities

Location
United States
Category
Model Evaluation/RLHF
Industry
Technology & Internet
Organization size
Individual
Updated
September 19, 2026

Description

When a frontier model performs poorly, knowing that the score dropped is not enough. Someone has to determine what failed, why it failed, whether the problem comes from the model, data, reward signal, or evaluation itself, and which intervention might actually improve behavior.


That diagnostic work defines this position.


Scale’s evaluation team works with leading AI labs on post-training and model assessment. The Research Scientist will design benchmarks and diagnostic methods that expose capability gaps, reasoning failures, robustness problems, and alignment issues across large language models and agents.


The work covers both text and multimodal models. That means evaluation can involve reasoning, instruction following, agent behavior, multimodal understanding, or other capabilities where a simple accuracy metric may hide important failure patterns.


Post-training knowledge is especially important because evaluation is closely connected to improvement. After identifying a weakness, the scientist may use expertise in supervised fine-tuning, RLHF, reward modeling, preference modeling, or instruction tuning to understand which training intervention is likely to address it.


Root-cause analysis is a major theme. Instead of producing a benchmark number and moving on, the role asks researchers to characterize failure modes deeply enough that model builders can use the findings as technical input for the next generation of systems.


The scientist will also collaborate with external foundation-model teams. This makes communication important: evaluation findings need to be rigorous enough for research use while also being clear enough to influence technical strategy.


Scale is looking for graduate-level ML or AI expertise, strong understanding of deep learning and reinforcement learning, and previous work with model fine-tuning or evaluation. Publication experience at major conferences is particularly relevant.


For researchers interested in measuring frontier models rather than only training them, this offers direct exposure to how evaluation drives real post-training decisions.


Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.


https://scale.com/careers/4728014005

Expertise

  • Model Evaluation
  • RLHF
  • Reward Modeling
  • Benchmarking
  • Multimodal Evaluation
  • Failure Analysis
  • Deep Learning
  • Curated Opportunity

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.