CoRL 2026

SkillWeaver Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

He Zhu1, Lusen Zhao2†, Kwan Man Cheng1, Su Li1, Katerina Fragkiadaki1

1Carnegie Mellon University 2University of Illinois Urbana-Champaign

†Work done during an internship at Carnegie Mellon University

31 100 mcts select retrieve expand execute verify backprop accept distill search tree MEMORY VLM PLANNER ROBOT SKILL VLM VERIFIER input SIM SCENE INSTRUCTION retrieve accept propose add lesson strategy verified demos VISUOMOTOR POLICY train

Video

What SkillWeaver does

Dozens of reconstructed simulation assets — cereal boxes, jars, bottles, a teapot, a woven basket, plates — arranged on a table.

Asset generation

10,192
Sim-ready assets across indoor categories

VLMs and 3D vision tools take care of object category proposal, image generation, 3D reconstruction, and attribute annotation.

A generated scene: a salad dressing bottle, a cucumber, a spatula and two oranges on a wooden table below a robot gripper, in a living-room setting.

Scene generation

14K+
Generated scenes

Scenes are composed procedurally from the asset library, or reconstructed from a seed observation as an interactive digital twin.

Rows of robot arms train in parallel in simulation, each at its own table; in the foreground one grasps an orange bottle.

Neural Interaction skills

4
action types

Pick,Place,OPEN, and CLOSE skills are trained with PPO in IsaacLab and split into sub-policies by object geometry, pose, and articulation type.

An MCTS planner exchanges past experience with memory and feedback with a verifier, and selects pick, place, open and close skills that execute in a simulated environment.

Data collection

39K+
Demonstrations · 11.0M frames

A VLM agent selects and grounds skills in a verifier-guided search tree, and experience memory transfers what worked or failed across episodes.

Overall Pipeline

Examples of Collected Data

“Grasp the pickle jar.”(round object)
“Take hold of the lemon.”(round object)
“Pick up the banana.”(elongated lying object)
“Lift the green cup.”(concave object)
“Take hold of the curvy bottle.”(elongated upright object)
“Grasp the long neck beer bottle.”(elongated upright object)
“Lift the yellow milk carton.”(elongated upright object)
“Pick up the bowl.”(concave object)

Simulation Results

SkillWeaver data improves OOD generalization

We reconstruct LIBERO environments in IsaacLab, apply extensive augmentations, and collect a dataset about 10× larger than the original one. We then train π0.5 on mixtures of the original LIBERO data and our generated trajectories, and evaluate across the four LIBERO-PRO suites. Significant improvements on OOD generalization are observed across all four LIBERO-PRO suites under Position, Task, and Environment perturbations.

Success rate under single perturbations, by LIBERO-PRO suite Four panels, one per LIBERO-PRO suite, each comparing training on ground-truth data alone against adding SkillWeaver data under position, task and environment perturbations. Adding our data wins all twelve comparisons. GT only SkillWeaver Data + GT LIBERO-Object 0 25 50 75 100 Success rate (%) 0.0 77.0 Position 10.0 43.2 Task 48.8 82.6 Environment LIBERO-Spatial 0 25 50 75 100 Success rate (%) 46.6 71.6 Position 49.4 72.4 Task 52.4 72.4 Environment LIBERO-Goal 0 25 50 75 100 Success rate (%) 11.3 60.3 Position 18.7 50.7 Task 74.7 78.7 Environment LIBERO-10 0 25 50 75 100 Success rate (%) 4.3 43.4 Position 4.6 33.4 Task 20.3 43.4 Environment

A clear data-scaling trend is observed: OOD success rate on the LIBERO-PRO-Object suite improves with data scaling, averaged over the Position, Task and Environment perturbations.

OOD success rate against amount of generated data Average success rate on the LIBERO-PRO Object suite over the position, task and environment perturbations, rising from 13.1 at 3.125% of our data to 67.6 at 100%, against a 19.6 ground-truth-only baseline. 0 20 40 60 80 Average Success Rate (%) Data Fraction 3.125% 6.25% 12.5% 25% 50% 100% 21 44 89 180 384 708 Avg. Demos per Task GT only: 19.6% 13.1 21.2 35.7 55.8 61.9 67.6

Training data visualization and simulation rollouts:

Instruction to VLA

“pick up the milk and place it in the basket”

Layout change: the milk and the cream cheese have swapped places

Instruction to VLA

“pick up the bbq sauce and place it in the basket”

Layout change: the bbq sauce and the chocolate pudding have swapped places

Instruction to VLA

“pick up the orange juice and place it in the basket”

Layout change: the orange juice and the bbq sauce have swapped places

Instruction to VLA

“pick up the tomato sauce and place it in the basket”

Layout change: the tomato sauce and the milk have swapped places

We further stress-test compositional generalization by applying two perturbations at a time, covering all 10 pairwise combinations. Even under this challenging setting, SkillWeaver data significantly improves success on 9 of the 10 perturbation combinations that include at least one non-illusory dimension, i.e., Position, Task, or Environment, with gains of up to 82.6 percentage points.

Success rate under pairwise augmentation Grouped bars over ten augmentation pairs, five per row. Training on ground-truth data alone collapses to zero on three of them and stays under 21 percent on five more; adding SkillWeaver data reaches 34 to 83 percent. GT only SkillWeaver Data + GT 0 25 50 75 100 Success rate (%) 97.8 89.4 Language + Object 0.0 78.4 Language + Position 10.0 40.8 Language + Task 39.8 77.1 Language + Environment 0.0 82.6 Object + Position 0 25 50 75 100 Success rate (%) 10.0 34.0 Object + Task 40.2 57.3 Object + Environment 0.0 73.0 Position + Task 21.3 70.6 Position + Environment 0.4 50.6 Task + Environment

Simulation rollouts:

Instruction to VLA

“pick up the cream cheese and place it in the basket”

Position + Environment: the alphabet soup and the cream cheese swap starting places, and the background and textures change

Instruction to VLA

“pick up the ketchup and place it in the basket”

Position + Environment: the cream cheese and the ketchup swap starting places, and the background and textures change

Instruction to VLA

“pick up the salad dressing and place it in the basket”

Position + Environment: the milk and the salad dressing swap starting places, and the background and textures change

Instruction to VLA

“pick up the chocolate pudding and place it in the basket”

Position + Environment: the chocolate pudding and the ketchup swap starting places, and the background and textures change

SkillWeaver data enables robust sim-to-sim transfer

π0.5 trained with SkillWeaver data collected in IsaacLab can achieve zero/few-shot transfer to the LIBERO-Object and SIMPLER-WidowX benchmarks.

Zero- and few-shot sim-to-sim transfer to LIBERO-Object Success rate on LIBERO-Object after training on SkillWeaver data with 0, 1 or 5 ground-truth demonstrations per task added, against a 50-demonstration ground-truth-only baseline. 0 25 50 75 100 Success rate (%) 50 GT demos only: 95.0% 66.0% 0 63.4% 1 82.2% 5 GT demos per task added to SW data
Zero-/few-shot sim-to-sim transfer to LIBERO-Object.
Zero-shot transfer to SIMPLER-WidowX Zero-shot success rate on four SIMPLER-WidowX tasks after training on SkillWeaver data alone. 0 25 50 75 100 Success rate (%) 91.7% Eggplant 75.0% Carrot 62.5% Spoon 79.2% Stack cube SIMPLER-WidowX task
Zero-shot transfer to SIMPLER-WidowX.

Zero-shot simulation rollouts:

“pick up the bbq sauce and place it in the basket”
“pick up the milk and place it in the basket”
“pick up the ketchup and place it in the basket”
“pick up the salad dressing and place it in the basket”

Zero-shot Sim-to-Real Transfer All training data collected with SkillWeaver in simulation

The real-world pipeline has two stages. Large-scale simulation training (LST) uses demonstrations that SkillWeaver generates across diverse in-the-wild scenes in broad exploration mode. Post-training (PT) then adds task-specific trajectories generated in targeted exploration mode, inside a digital twin reconstructed from the target real-world scene. Both stages train only on SkillWeaver-generated simulation data, with no teleoperated demonstrations. We evaluate object, color, and spatial grounding under the original setup and two distribution shifts — swapped object positions and a changed tablecloth.

Quantitative Results

Task Original PT Original LST+PT Swap PT Swap LST+PT Bg. PT Bg. LST+PT
Object pick up the apple 40 82+42 0 86+86 32 50+18
pick up the banana 82 66−16 10 80+70 78 44−34
place the apple in the bowl 4 94+90 0 60+60 0 78+78
place the lemon in the bowl 0 50+50 0 48+48 0 76+76
Spatial pick up the fruit on the left 36 80+44 0 72+72 16 44+28
pick up the fruit on the right 72 58−14 0 58+58 18 20+2
place the fruit at the front in the bowl 0 26+26 0 6+6 0 26+26
place the fruit at the back in the bowl 0 18+18 0 30+30 0 30+30
Colour pick up the red fruit 50 90+40 0 82+82 34 58+24
pick up the yellow fruit 84 52−32 6 78+72 16 18+2
place the red fruit in the bowl 0 76+76 0 50+50 0 50+50
place the yellow fruit in the bowl 2 58+56 0 36+36 0 6+6
Average 30.8 62.5+31.7 1.3 57.2+55.8 16.2 41.7+25.5

Success rate (%) per real-world task, 50 trials each. PT = post-training on the reconstructed scene only; LST+PT = large-scale simulation training followed by the same post-training.

Real-world success rate by scenario Average success rate over twelve real-world tasks in each of three scenarios, comparing post-training alone against large-scale simulation training followed by the same post-training. The second wins every scenario, by most under swapped objects, where post-training alone reaches 1.3 per cent. PT LST+PT 0 25 50 75 100 Success rate (%) 30.8 62.5 Original 1.3 57.2 Swap Objects 16.2 41.7 Change Background Real-world scenario
Averaged over the twelve tasks above, LST+PT wins every scenario. The gap is widest under swapped objects, where post-training alone all but fails.

Qualitative Results

“pick up the apple”
“pick up the banana”
“place the apple in the bowl”
“place the lemon in the bowl”

Ablations

MCTSfailed branch dropped · verified prefix kept
START 4 EXPANSIONS, 8 ROLLOUTS, 62.5% NODE-LEVEL SUCCESS
Linearchain grows on failure · restarts at max depth
MAX DEPTH START RESTART MAX DEPTH 11 EXPANSIONS, 11 ROLLOUTS, 63.6% NODE-LEVEL SUCCESS
verifier passverifier failaccepted trajectory
Search Linear MCTS (ours)
Breadth12
Max Depth204
Simulation Budget2020
Expansions / Success Traj.6.743.60
Time / Success Traj. (s)891645

Tested on five long-horizon tasks under the same simulation budget. MCTS needs fewer expansions and less wall-clock time per successful trajectory, by reusing successful skill prefixes across branches.

BibTeX

@inproceedings{zhu2026skillweaver,
  title     = {SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation},
  author    = {Zhu, He and Zhao, Lusen and Cheng, Kwan Man and Fragkiadaki, Katerina},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}