Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning
1KAIST AI 2Holiday Robotics
*Equal contribution †Corresponding authors
Robot data is expensive to collect
Robot demonstrations are often collected by hand, one episode at a time. Even the 20 short demonstrations behind H20 add up—and that is before resetting the scene.
20 episode durations; six shown in video
for 20 demonstrations, excluding resets.
3,000 demonstrations at this pace: of continuous teleoperation.
Simulation scales from a small set of demonstrations
Starting from 20 human demonstrations per task, DiscoDemo generates thousands more in parallel without an operator for each rollout. The resulting data can train policies that transfer to a real robot.
36 DiscoDemo rollouts: nine per task, from the same starting scene within each task.
Compute excludes human data collection, scene setup, and downstream policy fine-tuning.
Reconstructing the real workcell
We recreate four real-robot tasks in Isaac Sim: pick-and-place, cube stacking, and two peg insertions. Below, compare a recorded pick-and-place episode with its simulation replay, frame for frame.
Drag to reveal · Use ← / → keys
The same recorded episode and its physics-based replay. Real and simulation share all 302 frames at 20 fps. Drag to compare either the external or wrist camera.
What makes simulated demonstrations useful?
1. How do we generate successful trajectories?
Motion planningDesign paths with rules
MimicGenAdapt a human demonstration
Reinforcement learningLearn actions from rewards
PnP-Banana rollouts from motion planning, MimicGen, and P-RFCL.
2. How do we discover different solutions?
More demonstrations can still repeat the same narrow behavior, leaving an imitation policy unprepared for unfamiliar states. DiscoDemo discovers different successful solutions to broaden the states covered by its data.
How DiscoDemo works
-
Learn diverse solutions
A skill-conditioned RL policy learns from simulator state, with a demonstration-guided reverse curriculum for competence and a success-gated diversity reward for distinct solutions.
-
Generate and filter demonstrations
Each rollout samples a new skill; successful trajectories that pass safety checks are kept, 3,000 per task.
-
Imitate from pixels
Shoulder- and wrist-camera images with their actions fine-tune a DROID-pretrained π0.5, which acts without simulator state.
What the generated demonstrations look like
Each line traces a successful gripper path from the same starting scene. DiscoDemo keeps the direct motions of RL while discovering a wider range of approaches.
Up to 24 successful paths per source, from one fixed placement, at 3× speed.
Demonstration samples
Five successful rollouts per source from one fixed placement. H20 shows human demonstrations replayed in simulation, each from its original placement.
Analysis
Wider coverage, at a lower cost
DiscoDemo scores highest on our average coverage and diversity metrics, with near-95% generation yield before safety filtering. Episodes stay short, and generation compute—including generator training, rollouts, and rendering—remains among the lowest.
Balancing diversity and task success
More diversity helps up to a point: average imitation success peaks at a moderate diversity weight. We use α = 0.3 for all four tasks.
Efficient generator training
Parallel, state-based RL reaches high success within a few GPU-hours, with no systematic slowdown from the diversity objective. DiscoDemo and P-RFCL times are estimated; PLD and ExpertGen times are recorded.
Downstream imitation results
With the same imitation learner and training recipe, DiscoDemo reaches 62% success in simulation and 78% on the real robot, averaged across four tasks; the strongest baselines reach 40% and 39%. Generated datasets contain 3,000 successful demonstrations per task; H20 uses 20 human demonstrations.
Downstream policies in simulation
Columns are the eight data sources, rows five scenes per task; every policy starts from the same state within a scene. Column heads show the success rate on that task.
Real-robot policies trained on eight data sources
Every source on the same three real-robot scenes. Policies trained on generated data transfer from simulation without real-world fine-tuning.
Complementary to corrective data collection
DiscoDemo remains stronger than P-RFCL after the same DAgger or DART corrections, especially on FMB-SqCircle. Diverse successful demonstrations and corrective data offer complementary benefits.
Takeaways
- Discover, don't just multiply. A skill-conditioned RL generator turns 20 human demonstrations into distinct successful solutions in simulation.
- Better imitation. With the same π0.5 recipe, policies reach 62% in simulation and 78% on a real robot, vs. 40% and 39% for the strongest baselines.
- Wider coverage, lower cost. Generated data covers more of the task space with short episodes and among the lowest generation compute.
Citation
@article{park2026discodemo,
title = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
author = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
journal = {arXiv preprint},
year = {2026}
}