Asset generation
VLMs and 3D vision tools take care of object category proposal, image generation, 3D reconstruction, and attribute annotation.
CoRL 2026
1Carnegie Mellon University 2University of Illinois Urbana-Champaign
†Work done during an internship at Carnegie Mellon University
VLMs and 3D vision tools take care of object category proposal, image generation, 3D reconstruction, and attribute annotation.
Scenes are composed procedurally from the asset library, or reconstructed from a seed observation as an interactive digital twin.
Pick,Place,OPEN, and CLOSE skills are trained with PPO in IsaacLab and split into sub-policies by object geometry, pose, and articulation type.
A VLM agent selects and grounds skills in a verifier-guided search tree, and experience memory transfers what worked or failed across episodes.
Overall Pipeline. Simulated scenes and task instructions go into the SkillWeaver agent, which searches over neural interaction skills, verifies each outcome and distills what it learns into memory. The successful trajectories it collects are distilled into visuomotor policies.
Nothing in “straighten it up” says what upright means. Both the planner and verifier have to work that out for themselves — that the object’s own +z should end up in line with the table’s — the planner to aim for it, the verifier to judge whether it was reached.
What makes "insert" hard is orientation reasoning. For example, a long spoon carried in flat will never enter a holder however well it is centred: its long axis has to be brought in line with the container’s +z first. The instruction says none of that — the planner has to identify the challenge and reason about the correct orientation by itself.
What makes "pour" hard: the planner has to decompose pouring into a sequence of primitives—“pick,” “move above the container,” and “rotate”—while correctly reasoning about the source object’s target orientation required for tilting and pouring.
Pour · clip 2
Pour · clip 3
Pour · clip 4
Open · clip 1
Open · clip 2
Open · clip 3
Close · clip 1
Close · clip 2
Close · clip 3
These tasks are challenging not just because they are long-horizon—- often requiring four or five skills in sequence—- but also because the scene may violate the instruction's implicit assumptions. If a container starts on its side, “put it in” first requires standing it upright, even though that step is never stated. The planner has to discover such missing steps through search, with memory using past failures to guide the next attempt.
Long-horizon · clip 3
Long-horizon · clip 4
We reconstruct LIBERO environments in IsaacLab, apply extensive augmentations, and collect a dataset about 10× larger than the original one. We then train π0.5 on mixtures of the original LIBERO data and our generated trajectories, and evaluate across the four LIBERO-PRO suites. Significant improvements on OOD generalization are observed across all four LIBERO-PRO suites under Position, Task, and Environment perturbations.
A clear data-scaling trend is observed: OOD success rate on the LIBERO-PRO-Object suite improves with data scaling, averaged over the Position, Task and Environment perturbations.
Training data visualization and simulation rollouts:
Every generated scene applies all augmentations at once — object sizes, layouts, swapped and replaced objects, backgrounds, lighting, camera poses and rewritten instructions — instead of one perturbation at a time. Each clip below therefore comes from its own scene, with its own objects and its own instruction.
Instruction to VLA
“pick up the milk and place it in the basket”
Layout change: the milk and the cream cheese have swapped places
Instruction to VLA
“pick up the bbq sauce and place it in the basket”
Layout change: the bbq sauce and the chocolate pudding have swapped places
Instruction to VLA
“pick up the orange juice and place it in the basket”
Layout change: the orange juice and the bbq sauce have swapped places
Instruction to VLA
“pick up the tomato sauce and place it in the basket”
Layout change: the tomato sauce and the milk have swapped places
Instruction to VLA
“pick the salad dressing and place it in the basket”
original task: pick up the chocolate pudding and place it in the basket
Instruction to VLA
“pick the cream cheese and place it in the basket”
original task: pick up the alphabet soup and place it in the basket
Instruction to VLA
“pick the butter and place it in the basket”
original task: pick up the milk and place it in the basket
Instruction to VLA
“pick the chocolate pudding and place it in the basket”
original task: pick up the orange juice and place it in the basket
Every generated scene applies all augmentations at once — object sizes, layouts, swapped and replaced objects, backgrounds, lighting, camera poses and rewritten instructions — instead of one perturbation at a time. Each clip below therefore comes from its own scene, with its own objects and its own instruction.
Instruction to VLA
“pick up the black bowl next to the ramekin and place it on the plate”
Layout change: the black bowl and the plate have swapped places
Instruction to VLA
“pick up the black bowl on the stove and place it on the plate”
Layout change: the plate and the ramekin have swapped places
Instruction to VLA
“pick up the black bowl on the wooden cabinet and place it on the plate”
Layout change: the plate and the ramekin have swapped places
Instruction to VLA
“pick up the black bowl on the cookie box and place it on the plate”
Layout change: the plate and the wooden cabinet have swapped places
Instruction to VLA
“pick the akita black bowl on the top of the wooden cabinet and place it on the plate”
original task: pick up the black bowl in the top drawer of the wooden cabinet and place it on the plate
Instruction to VLA
“pick the akita black bowl next to the plate and place it on the plate”
original task: pick up the black bowl from table center and place it on the plate
Instruction to VLA
“pick the akita black bowl on the top of the cabinet and place it on the plate”
original task: pick up the black bowl on the cookie box and place it on the plate
Instruction to VLA
“pick the akita black bowl next to the ramekin and place it on the plate”
original task: pick up the black bowl next to the plate and place it on the plate
Every generated scene applies all augmentations at once — object sizes, layouts, swapped and replaced objects, backgrounds, lighting, camera poses and rewritten instructions — instead of one perturbation at a time. Each clip below therefore comes from its own scene, with its own objects and its own instruction.
Instruction to VLA
“put the bowl on top of the cabinet”
Layout change: the objects on the table are shuffled into each other’s places
Instruction to VLA
“put the cream cheese in the bowl”
Layout change: the objects on the table are shuffled into each other’s places
Instruction to VLA
“put the wine bottle on the rack”
Layout change: the black bowl and the wine bottle have swapped places, and so have the wine rack and the wooden cabinet
Instruction to VLA
“put the bowl on the stove”
Layout change: the objects on the table are shuffled into each other’s places
Instruction to VLA
“put the wine bottle in the bowl”
original task: put the wine bottle on top of the cabinet
Instruction to VLA
“put the plate on the stove”
original task: put the bowl on the stove
Instruction to VLA
“put the plate on the top of the drawer”
original task: put the bowl on top of the cabinet
Instruction to VLA
“put the wine bottle in the bowl”
original task: put the cream cheese in the bowl
Every generated scene applies all augmentations at once — object sizes, layouts, swapped and replaced objects, backgrounds, lighting, camera poses and rewritten instructions — instead of one perturbation at a time. Each clip below therefore comes from its own scene, with its own objects and its own instruction.
Instruction to VLA
“put the white mug on the left plate and put the yellow and white mug on the right plate”
Layout change: the two mugs have swapped starting places
Instruction to VLA
“put both the cream cheese box and the butter in the basket”
Layout change: the cream cheese and the orange juice swap places, and so do the butter and the tomato sauce
Instruction to VLA
“put both the alphabet soup and the tomato sauce in the basket”
Layout change: the alphabet soup and the ketchup swap places, and so do the tomato sauce and the orange juice
Instruction to VLA
“put both the alphabet soup and the cream cheese box in the basket”
Layout change: the alphabet soup and the ketchup have swapped places, and so have the cream cheese and the tomato sauce
Instruction to VLA
“put the yellow and white mug on the left plate and put the white mug on the right plate”
original task: put the white mug on the left plate and put the yellow and white mug on the right plate
Instruction to VLA
“put both the cream cheese and the tomato sauce in the basket”
original task: put both the alphabet soup and the tomato sauce in the basket
Instruction to VLA
“put both the ketchup and the cream cheese box in the basket”
original task: put both the alphabet soup and the cream cheese box in the basket
Instruction to VLA
“put both the alphabet soup and the butter in the basket”
original task: put both the cream cheese box and the butter in the basket
We further stress-test compositional generalization by applying two perturbations at a time, covering all 10 pairwise combinations. Even under this challenging setting, SkillWeaver data significantly improves success on 9 of the 10 perturbation combinations that include at least one non-illusory dimension, i.e., Position, Task, or Environment, with gains of up to 82.6 percentage points.
Simulation rollouts:
Instruction to VLA
“pick up the cream cheese and place it in the basket”
Position + Environment: the alphabet soup and the cream cheese swap starting places, and the background and textures change
Instruction to VLA
“pick up the ketchup and place it in the basket”
Position + Environment: the cream cheese and the ketchup swap starting places, and the background and textures change
Instruction to VLA
“pick up the salad dressing and place it in the basket”
Position + Environment: the milk and the salad dressing swap starting places, and the background and textures change
Instruction to VLA
“pick up the chocolate pudding and place it in the basket”
Position + Environment: the chocolate pudding and the ketchup swap starting places, and the background and textures change
Instruction to VLA
“pick the cream cheese and place it in the basket”
Position + Task: the butter and the cream cheese swap starting places, and the goal object changes
original task: pick up the alphabet soup and place it in the basket
Instruction to VLA
“pick the tomato sauce and place it in the basket”
Position + Task: the ketchup and the tomato sauce swap starting places, and the goal object changes
original task: pick up the salad dressing and place it in the basket
Instruction to VLA
“pick the ketchup and place it in the basket”
Position + Task: the ketchup and the salad dressing swap starting places, and the goal object changes
original task: pick up the bbq sauce and place it in the basket
Instruction to VLA
“pick the bbq sauce and place it in the basket”
Position + Task: the bbq sauce and the milk swap starting places, and the goal object changes
original task: pick up the tomato sauce and place it in the basket
Instruction to VLA
“pick the cream cheese and place it in the basket”
Task + Environment: the goal object changes, and the background and textures change
original task: pick up the alphabet soup and place it in the basket
Instruction to VLA
“pick the bbq sauce and place it in the basket”
Task + Environment: the goal object changes, and the background and textures change
original task: pick up the tomato sauce and place it in the basket
Instruction to VLA
“pick the salad dressing and place it in the basket”
Task + Environment: the goal object changes, and the background and textures change
original task: pick up the chocolate pudding and place it in the basket
Instruction to VLA
“pick the ketchup and place it in the basket”
Task + Environment: the goal object changes, and the background and textures change
original task: pick up the bbq sauce and place it in the basket
Instruction to VLA
“grab alphabet soup and put it into basket”
Position + Language: the alphabet soup and the salad dressing swap starting places, and the instruction is rephrased
Instruction to VLA
“grab cream cheese and put into basket”
Position + Language: the alphabet soup and the cream cheese swap starting places, and the instruction is rephrased
Instruction to VLA
“grab salad dressing and put into basket”
Position + Language: the milk and the salad dressing swap starting places, and the instruction is rephrased
Instruction to VLA
“pick up bbq sauce and put it into basket”
Position + Language: the bbq sauce and the chocolate pudding swap starting places, and the instruction is rephrased
Instruction to VLA
“grab butter and put it into basket”
Task + Language: the goal object changes, and the instruction is rephrased
original task: pick up the milk and place it in the basket
Instruction to VLA
“grab salad dressing and put into basket”
Task + Language: the goal object changes, and the instruction is rephrased
original task: pick up the chocolate pudding and place it in the basket
Instruction to VLA
“grab cream cheese and put into basket”
Task + Language: the goal object changes, and the instruction is rephrased
original task: pick up the alphabet soup and place it in the basket
Instruction to VLA
“pick orange juice and put into basket”
Task + Language: the goal object changes, and the instruction is rephrased
original task: pick up the butter and place it in the basket
Instruction to VLA
“grab butter and put it into basket”
Environment + Language: the background and textures change, and the instruction is rephrased
Instruction to VLA
“grab alphabet soup and put it into basket”
Environment + Language: the background and textures change, and the instruction is rephrased
Instruction to VLA
“grab milk and put into basket”
Environment + Language: the background and textures change, and the instruction is rephrased
Instruction to VLA
“lift chocolate pudding and put it in basket”
Environment + Language: the background and textures change, and the instruction is rephrased
Instruction to VLA
“pick up the alphabet soup and place it in the basket”
Position + Object: the red alphabet soup and the salad dressing swap starting places, and objects are replaced
Instruction to VLA
“pick up the cream cheese and place it in the basket”
Position + Object: the alphabet soup and the red cream cheese swap starting places, and objects are replaced
Instruction to VLA
“pick up the salad dressing and place it in the basket”
Position + Object: the milk and the red salad dressing swap starting places, and objects are replaced
Instruction to VLA
“pick up the bbq sauce and place it in the basket”
Position + Object: the chocolate pudding and the green bbq sauce swap starting places, and objects are replaced
Instruction to VLA
“pick the butter and place it in the basket”
Task + Object: the goal object changes, and objects are replaced
original task: pick up the milk and place it in the basket
Instruction to VLA
“pick the salad dressing and place it in the basket”
Task + Object: the goal object changes, and objects are replaced
original task: pick up the chocolate pudding and place it in the basket
Instruction to VLA
“pick the alphabet soup and place it in the basket”
Task + Object: the goal object changes, and objects are replaced
original task: pick up the cream cheese and place it in the basket
Instruction to VLA
“pick the chocolate pudding and place it in the basket”
Task + Object: the goal object changes, and objects are replaced
original task: pick up the orange juice and place it in the basket
Instruction to VLA
“pick up the cream cheese and place it in the basket”
Environment + Object: the background and textures change, and objects are replaced
Instruction to VLA
“pick up the milk and place it in the basket”
Environment + Object: the background and textures change, and objects are replaced
Instruction to VLA
“pick up the chocolate pudding and place it in the basket”
Environment + Object: the background and textures change, and objects are replaced
Instruction to VLA
“pick up the orange juice and place it in the basket”
Environment + Object: the background and textures change, and objects are replaced
π0.5 trained with SkillWeaver data collected in IsaacLab can achieve zero/few-shot transfer to the LIBERO-Object and SIMPLER-WidowX benchmarks.
Zero-shot simulation rollouts:
The real-world pipeline has two stages. Large-scale simulation training (LST) uses demonstrations that SkillWeaver generates across diverse in-the-wild scenes in broad exploration mode. Post-training (PT) then adds task-specific trajectories generated in targeted exploration mode, inside a digital twin reconstructed from the target real-world scene. Both stages train only on SkillWeaver-generated simulation data, with no teleoperated demonstrations. We evaluate object, color, and spatial grounding under the original setup and two distribution shifts — swapped object positions and a changed tablecloth.
| Task | Original PT | Original LST+PT | Swap PT | Swap LST+PT | Bg. PT | Bg. LST+PT | |
|---|---|---|---|---|---|---|---|
| Object | pick up the apple | 40 | 82+42 | 0 | 86+86 | 32 | 50+18 |
| pick up the banana | 82 | 66−16 | 10 | 80+70 | 78 | 44−34 | |
| place the apple in the bowl | 4 | 94+90 | 0 | 60+60 | 0 | 78+78 | |
| place the lemon in the bowl | 0 | 50+50 | 0 | 48+48 | 0 | 76+76 | |
| Spatial | pick up the fruit on the left | 36 | 80+44 | 0 | 72+72 | 16 | 44+28 |
| pick up the fruit on the right | 72 | 58−14 | 0 | 58+58 | 18 | 20+2 | |
| place the fruit at the front in the bowl | 0 | 26+26 | 0 | 6+6 | 0 | 26+26 | |
| place the fruit at the back in the bowl | 0 | 18+18 | 0 | 30+30 | 0 | 30+30 | |
| Colour | pick up the red fruit | 50 | 90+40 | 0 | 82+82 | 34 | 58+24 |
| pick up the yellow fruit | 84 | 52−32 | 6 | 78+72 | 16 | 18+2 | |
| place the red fruit in the bowl | 0 | 76+76 | 0 | 50+50 | 0 | 50+50 | |
| place the yellow fruit in the bowl | 2 | 58+56 | 0 | 36+36 | 0 | 6+6 | |
| Average | 30.8 | 62.5+31.7 | 1.3 | 57.2+55.8 | 16.2 | 41.7+25.5 | |
Success rate (%) per real-world task, 50 trials each. PT = post-training on the reconstructed scene only; LST+PT = large-scale simulation training followed by the same post-training.
Articulation · clip 1
Articulation · clip 2
Articulation · clip 3
Articulation · clip 4
| Search | Linear | MCTS (ours) |
|---|---|---|
| Breadth | 1 | 2 |
| Max Depth | 20 | 4 |
| Simulation Budget | 20 | 20 |
| Expansions / Success Traj. | 6.74 | 3.60 |
| Time / Success Traj. (s) | 891 | 645 |
Tested on five long-horizon tasks under the same simulation budget. MCTS needs fewer expansions and less wall-clock time per successful trajectory, by reusing successful skill prefixes across branches.
MCTS success rate (%) over ten searches per task at maximum depth K = 6, on tasks that hinge on orientation adjustment, spatial reasoning, or an implicit prerequisite. Memory lifts average search success from 37.5% to 85.0%.
@inproceedings{zhu2026skillweaver,
title = {SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation},
author = {Zhu, He and Zhao, Lusen and Cheng, Kwan Man and Fragkiadaki, Katerina},
booktitle = {Conference on Robot Learning (CoRL)},
year = {2026}
}