AutoEnvScaling: Automating the Data Flywheel with Terminal Agents

Paper · Project page · Code

A terminal agent can write code, run experiments, and debug a failing program. Those same abilities can help it build the environments used to train the next version of itself.

AutoEnvScaling treats environment design as a terminal task. A proposer agent studies the solver’s recent attempts, builds new environments, and tests them. The solver then trains on the accepted environments, producing fresh experience that guides the next round of design.

Put environment design inside the loop

A fixed training pool eventually becomes familiar. Researchers then need to inspect failures, identify missing skills, and build new tasks. AutoEnvScaling gives that workflow to an agent with a terminal, a workspace, tools, and access to the solver’s rollouts.

The loop has four steps:

  1. Inspect experience. The proposer reads the solver’s rollouts to identify what it still struggles with.
  2. Build and test. It writes complete training environments and uses execution feedback to repair them.
  3. Train the solver. Validated environments provide tasks and verifier rewards for reinforcement learning.
  4. Refresh the curriculum. The solver’s new rollouts become evidence for the next generation round.

Executable checks matter here. A plausible task description is only the beginning: the environment also needs to run and provide a useful learning signal. Giving the proposer a terminal lets it inspect and repair the artifact it is building.

Better environments improve the solver

In the paper’s environment-generation experiments, AutoEnvScaling increases the valid-environment rate by up to 25× over prompt-only generation, with up to 75% lower cost per valid environment. These are the largest gains across the tested models; costs decrease for six of the eight proposer models.

The environments also help downstream training. With GPT-5.6-sol as the proposer, Qwen3.5-35B-A3B reaches 38.9 on Terminal-Bench 2.1, compared with 34.6 for the matched static Tmax baseline. Removing solver feedback reduces that advantage, showing why the proposer needs to see the learner’s actual experience.

One model can play both roles

Environment design and task solving both use a terminal. This gives us a shared task interface for recursive self-improvement: the same model can propose environments, learn to solve them, and use the updated checkpoint in the next round.

We test this with Qwen3.6-35B-A3B in both roles. It first receives a cold start from 600 high-quality proposer trajectories; both AutoEnvScaling and the static baseline begin from that same checkpoint.

Evaluation Cold-start checkpoint AutoEnvScaling
Terminal-Bench 2.1 (avg@5) 41.1 53.3
Terminal-Bench 4.0 1.6 3.2

The Terminal-Bench 2.1 gain is 12.2 points over the cold-start checkpoint and 6.1 points over static Tmax training. Terminal-Bench 4.0 remains difficult, but performance also improves there. Training the proposer matters: when its training is removed, environment-validation failures grow and the run becomes unstable.

What can the agent build next?

The pipeline generates environments for software engineering, computer use, professional work, and GPU kernel programming. The terminal supplies a common interface for building and testing them, while solver feedback guides which environments to create next.

This is a concrete step toward agents maintaining their own training curriculum. The design of that curriculum becomes something the system can improve through experience, alongside its ability to solve the resulting tasks.

See the full paper for the validation pipeline, training setup, comparisons, and ablations.