SCOPE: Self-Play with Co-Evolving User Simulator
An agent can get better at talking to its training partner while getting worse at talking to everyone else. That is the problem we study in One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL, accepted at EMNLP 2026.
When the simulator becomes the shortcut
Multi-turn reinforcement learning often uses another language model to play the user. A frozen simulator tends to repeat a narrow set of behaviors. The agent learns to exploit those habits: training reward rises, but performance against unfamiliar simulators falls. Its responses also become less diverse.
We call this simulator collapse. Better performance against one training partner does not necessarily mean better interaction with real users.
Let the training partner evolve
SCOPE addresses this through population co-training. The agent and user simulator both learn from their conversations, while a pool of saved simulator checkpoints supplies different training partners. This keeps the agent from specializing to a single, unchanging behavior pattern.
The released implementation samples a simulator checkpoint before each rollout. Simulator rewards depend on the task: in Persuasion for Good, the simulated persuadee is rewarded for keeping donations low; in τ²-bench, the curriculum reward favors tasks on which the agent sometimes succeeds and sometimes fails. The aim is to maintain useful challenges as the agent improves.
When updating the simulator is impractical, Verbalized Sampling provides an inference-time alternative. The frozen model proposes several possible replies, and the interaction continues with one sampled reply. This broadens the behaviors the agent encounters without training a second model.
Does it transfer?
The paper studies Persuasion for Good, τ²-bench, and CooperBench. It reports held-out gains of up to 9% with Verbalized Sampling and 14% with Co-Training over single-simulator RL, along with improvements in a human study. The methods also preserve more policy diversity.
These are complementary interventions, with different task coverage: population co-training is implemented for the dialogue environments, while the released CooperBench experiments use a frozen coding partner.
The training partner is part of the learning problem. For agents that must work with unfamiliar people, maintaining a diverse, evolving set of partners matters alongside improving the agent itself.