Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning

Minho Park1* Kinam Kim1* Donghu Kim2 Byungkun Lee1 Dongyoon Hwang1 Yongjae Shin1 Junha Hyung1 Hojoon Lee2† Jaegul Choo1†

1KAIST AI 2Holiday Robotics

*Equal contribution †Corresponding authors

Robot data is expensive to collect

Robot demonstrations are often collected by hand, one episode at a time. Even the 20 short demonstrations behind H20 add up—and that is before resetting the scene.

2× speed
0.0 sclip 1 of 6

20 episode durations; six shown in video


for 20 demonstrations, excluding resets.

3,000 demonstrations at this pace: of continuous teleoperation.

Teleoperation alone cannot grow robot data to the scale of image and text datasets. Circle area is proportional to the number of samples.

Simulation scales from a small set of demonstrations

Starting from 20 human demonstrations per task, DiscoDemo generates thousands more in parallel without an operator for each rollout. The resulting data can train policies that transfer to a real robot.

36 DiscoDemo rollouts: nine per task, from the same starting scene within each task.

PnP-BananaStackCubeFMB-RoundFMB-SqCircle
3,000successful demonstrations per task
≈15 GPU-haverage per task, generator training through rendering (one GPU process)
20human seed demonstrations per task

Compute excludes human data collection, scene setup, and downstream policy fine-tuning.

Reconstructing the real workcell

We recreate four real-robot tasks in Isaac Sim: pick-and-place, cube stacking, and two peg insertions. Below, compare a recorded pick-and-place episode with its simulation replay, frame for frame.

PnP-Banana · Episode 10
RealSimulation

Drag to reveal · Use ← / → keys

0.0 s

The same recorded episode and its physics-based replay. Real and simulation share all 302 frames at 20 fps. Drag to compare either the external or wrist camera.

What makes simulated demonstrations useful?

1. How do we generate successful trajectories?

Motion planningDesign paths with rules

Motion planning connects task poses with task-specific logic.

MimicGenAdapt a human demonstration

MimicGen adapts human demonstrations to new scenes.

Reinforcement learningLearn actions from rewards

Reinforcement learning learns from task feedback, but can converge on similar motions.

PnP-Banana rollouts from motion planning, MimicGen, and P-RFCL.

2. How do we discover different solutions?

More demonstrations can still repeat the same narrow behavior, leaving an imitation policy unprepared for unfamiliar states. DiscoDemo discovers different successful solutions to broaden the states covered by its data.

Augmentation expands around one solution; DiscoDemo discovers distinct solutions.

How DiscoDemo works

Train a diverse generator, filter its demonstrations, then learn a policy from images.
  1. Learn diverse solutions

    A skill-conditioned RL policy learns from simulator state, with a demonstration-guided reverse curriculum for competence and a success-gated diversity reward for distinct solutions.

  2. Generate and filter demonstrations

    Each rollout samples a new skill; successful trajectories that pass safety checks are kept, 3,000 per task.

  3. Imitate from pixels

    Shoulder- and wrist-camera images with their actions fine-tune a DROID-pretrained π0.5, which acts without simulator state.

What the generated demonstrations look like

Each line traces a successful gripper path from the same starting scene. DiscoDemo keeps the direct motions of RL while discovering a wider range of approaches.

Up to 24 successful paths per source, from one fixed placement, at 3× speed.

Demonstration samples

Five successful rollouts per source from one fixed placement. H20 shows human demonstrations replayed in simulation, each from its original placement.

0.0 s

Analysis

Wider coverage, at a lower cost

DiscoDemo scores highest on our average coverage and diversity metrics, with near-95% generation yield before safety filtering. Episodes stay short, and generation compute—including generator training, rollouts, and rendering—remains among the lowest.

Balancing diversity and task success

More diversity helps up to a point: average imitation success peaks at a moderate diversity weight. We use α = 0.3 for all four tasks.

Efficient generator training

Parallel, state-based RL reaches high success within a few GPU-hours, with no systematic slowdown from the diversity objective. DiscoDemo and P-RFCL times are estimated; PLD and ExpertGen times are recorded.

Downstream imitation results

With the same imitation learner and training recipe, DiscoDemo reaches 62% success in simulation and 78% on the real robot, averaged across four tasks; the strongest baselines reach 40% and 39%. Generated datasets contain 3,000 successful demonstrations per task; H20 uses 20 human demonstrations.

Downstream policies in simulation

Columns are the eight data sources, rows five scenes per task; every policy starts from the same state within a scene. Column heads show the success rate on that task.

0.0 s

Real-robot policies trained on eight data sources

Every source on the same three real-robot scenes. Policies trained on generated data transfer from simulation without real-world fine-tuning.

0.0 s

Complementary to corrective data collection

DiscoDemo remains stronger than P-RFCL after the same DAgger or DART corrections, especially on FMB-SqCircle. Diverse successful demonstrations and corrective data offer complementary benefits.

Takeaways

  1. Discover, don't just multiply. A skill-conditioned RL generator turns 20 human demonstrations into distinct successful solutions in simulation.
  2. Better imitation. With the same π0.5 recipe, policies reach 62% in simulation and 78% on a real robot, vs. 40% and 39% for the strongest baselines.
  3. Wider coverage, lower cost. Generated data covers more of the task space with short episodes and among the lowest generation compute.

Citation

@article{park2026discodemo,
  title   = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
  author  = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
  journal = {arXiv preprint},
  year    = {2026}
}