A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control

Octi Zhang1,2*Mateo Guaman Castro1*Patrick Yin1*Ignacio Dagnino1Abhishek Gupta1Rosario Scalise1†Byron Boots1†

1University of Washington2NVIDIA*Equal contribution†Equal advising

CoRL 2026Paper ↗Code (coming soon)

Highlights

All videos at 1×

UR5e, NutHardware

UR5eHardware

UR5eSimulation

FrankaSimulation

ANYmal CSimulation

ANYmal DSimulation

Summary

Success-Guided Sampling (SGS) trains on the task configurations that are neither too easy nor too hard for the current policy. With it, reinforcement learning scales much more effectively, to over a million simulated robots at once, in legged locomotion and contact-rich manipulation.

Overview

Locomotion

A single policy, trained from scratch, takes ANYmal C and D across every terrain, from stepping stones to floating islands.

Manipulation

UR5e and Franka learn per-task contact-rich assembly from the NIST task board, with zero demonstrations and the same reward function across all tasks.

Real world

Trained only in simulation, the UR5e assembles parts on the real task board from RGB images, zero-shot.

Scale

SGS keeps improving as parallel environments grow to a million, where uniform sampling and other methods fall behind.

Method

01The change to PPO

SGS is standard PPO with a small outer loop. It changes only which task configuration an environment resets to.

PPO
  1. 1initialize policy π
  2. 2for each iteration:
  3. 3for each environment, in parallel:
  4. 4if its episode ended:
  5. 5reset to a task configuration drawn uniformly
  6. 6step π, store the transition
  7. 7update π with PPO
PPO with SGS
  1. 1initialize policy π
  2. 2fix N task configurations, each with an empty history buffer
  3. 3for each iteration:
  4. 4for each environment, in parallel:
  5. 5if its episode ended:
  6. 6record its success in that task configuration’s history buffer
  7. 7reset to a task configuration drawn weighted by success rate (SGS)
  8. 8step π, store the transition
  9. 9update π with PPO

added changed

02Task configurations

A task configuration is a triple (s0, g, e): the robot’s initial state s0, its goal g, and the environment e, such as the terrain type or the layout of the scene. SGS samples a large, fixed set of task configurations before training.

For assembly, the task configurations come in equal parts from three reset strategies.

Franka, nut on boltSimulation

The nut starts anywhere in the workspace, next to the gripper but not always in it, so the robot may have to fetch it first.

03Sampling during training

We follow one fixed task configuration of each UR5e task through the course of training. SGS gives it small weight while the policy never solves it, large weight when it’s getting better but not perfect, and small weight again once the policy has mastered it.

Iteration 2,8000 of 64 successes

SGS weight above 0.9× peak0%50%100%Success rate0%02,0004,000Training iteration0Peak0%50%100%Success rateSGS weight0.14× peak
Nut on bolt: one task configuration, 64 trials per checkpoint, with the 95% Wilson interval and the SGS weight relative to its peak
IterationSuccesses95% intervalSGS weight
00 of 640%–6%0.14
1,0000 of 640%–6%0.14
1,8000 of 640%–6%0.14
2,8000 of 640%–6%0.14
3,8009 of 648%–25%0.83
4,00027 of 6431%–54%0.99
4,20051 of 6468%–88%0.90
4,40052 of 6470%–89%0.88
4,60057 of 6479%–95%0.79
5,40059 of 6483%–97%0.73

01Scaling

With SGS, success keeps rising as the number of parallel environments grows to a million. The baselines stall or collapse.

SGSPLRUniformLinear curriculum

LocomotionANYmal D, all terrains
00.514K32K256K1MParallel environmentsSGS 0.73PLR 0.54Uniform, Linear 0.00

Success rate

Locomotion, success rate, mean of three seeds with 95% confidence interval
Method4K32K256K1M
SGS0.46 ± 0.020.49 ± 0.110.72 ± 0.050.73 ± 0.01
PLR0.00 ± 0.000.13 ± 0.010.55 ± 0.140.54 ± 0.13
Uniform0.00 ± 0.000.42 ± 0.040.00 ± 0.000.00 ± 0.00
Linear0.00 ± 0.000.32 ± 0.030.00 ± 0.000.00 ± 0.00
ManipulationFranka, nut-and-bolt assembly
00.514K32K256K1MParallel environmentsSGS 0.70Uniform 0.06, PLR 0.05

Success rate

Manipulation, success rate, mean of three seeds with 95% confidence interval
Method4K32K256K1M
SGS0.00 ± 0.010.06 ± 0.010.10 ± 0.140.70 ± 0.19
PLR0.00 ± 0.000.00 ± 0.000.02 ± 0.040.05 ± 0.11
Uniform0.00 ± 0.000.00 ± 0.000.01 ± 0.040.06 ± 0.13

02Real world

Trained only in simulation, the UR5e runs on the real task board from RGB images, zero-shot.

Nut on bolt

Rod in hole

03Continuous runs

Nut on bolt, Franka, simulation

All terrains, ANYmal C, simulation

Gear mesh, UR5e, real world

Clips

UR5eReal world

Nut on bolt

Rod in hole

Gear mesh

UR5eSimulation

Nut on bolt
Rod in hole
Gear mesh
BNC connector
Waterproof connector
Rectangular peg

ANYmal CSimulation

All terrains
Climbing box
Stepping stones
Gap
Balancing beam
Contour
Floating island
Inverted slope
Jump box
Parallel boxes
Pit
Radiating beam
Stairs

ANYmal DSimulation

Climbing box
Stepping stones
Gap
Balancing beam
Contour
Floating island
Inverted slope
Jump box
Maze
Parallel boxes
Pit
Radiating beam
Stairs

FrankaSimulation

Nut on bolt

BibTeX

@inproceedings{zhang2026balanced,
  title     = {A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control},
  author    = {Zhang, Octi and Guaman Castro, Mateo and Yin, Patrick and
               Dagnino, Ignacio and Gupta, Abhishek and Scalise, Rosario and
               Boots, Byron},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}

If you want to adapt this website or use its videos, see the Terms of Use.