Description
Building stronger autonomous agents is partly a model problem and partly an environment-and-data problem.
Scale’s Agent Capabilities & Environments team focuses heavily on the second half of that equation: what experiences should an agent train on, how should success be measured, what kinds of tasks expose weaknesses, and which interaction data actually teaches useful behavior.
The environments can be diverse. Agents may interact with code repositories, browsers, graphical interfaces, databases, or other software systems rather than producing a single text response. Each environment introduces different notions of action, reward, state, failure, and completion.
The scientist will study the data required to advance these systems and help shape Scale’s broader agent-data strategy. That can involve browser agents, software-engineering agents, GUI agents, tool-use problems, planning, reward signals, and evaluation frameworks.
Research is expected to translate into prototypes quickly. Familiarity with PyTorch, JAX, or TensorFlow is useful not simply for running existing models but for implementing ideas from recent literature and testing whether they improve agent behavior.
Agent reasoning methods and orchestration frameworks can also be relevant. Scale mentions areas such as tool use, text-to-SQL, browser interaction, coding, GUI environments, LangGraph-style systems, and planning approaches.
Evaluation is tightly connected to training. Researchers need to understand whether an agent is failing because of planning, data, reward design, model capability, environment design, or the evaluation itself.
The role includes publishing research and working with external researchers while also collaborating with engineering teams that turn promising methods into scalable systems.
Scale is looking for several years of sophisticated ML research or product-development experience, evidence of published work, and practical exposure to LLM agents. Fine-tuning open models and cloud ML infrastructure are useful additions.
For someone interested in how autonomous systems learn to act—not merely how language models generate text—this is one of the more research-heavy agent opportunities in the current batch.
Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.
https://scale.com/careers/4488520005