CoRL 2026Conference on Robot Learning

Where Should
Action Generation
Begin?

A Learnable Source Prior
for Generative Robot Policies

Meipo Dai1,*Qiyuan Zhuang1,*He-Yang Xu1,*Ying-Jie Shuai1Yijun Wang1Qi Dou2Xiu-Shen Wei1,†

1 Southeast University
2 The Chinese University of Hong Kong

* Equal contribution † Corresponding author

81.6%

Average success
15 RoboTwin tasks

+25.5 pp

Over the same-architecture
NoPrior baseline

80.0%

Average success
3 real-world tasks

0.21M

Prior-head parameters
1.6% of the 13.22M model

01 / OVERVIEW

A small change.
A more informed start.

Learn where to start generating actions, with a lightweight prior that adapts to the robot’s own state.

Submission presentation with updated author names and affiliations; original narration and experimental slides preserved. The video reports an earlier BridgePolicy simulation score (60.8%); the paper and results below report 61.4%.

The idea

Generative robot policies typically begin with an observation-independent standard Gaussian. LeaP learns the source distribution itself. A lightweight MLP predicts the mean and state-adaptive variance of a proprioception-conditioned diagonal Gaussian. The generator starts with an informed yet stochastic sample, then refines it using visual and proprioceptive observations. The generator architecture and inference solver stay unchanged.

02 / METHOD

Learn the source.
Keep the generator.

The prior anchors generation in a state-dependent neighborhood. The generator learns the fine-grained actions.

LeaP pipeline: proprioceptive features enter a learned Gaussian prior; a sampled source and visual plus state conditioning enter the flow or diffusion generator to produce an action chunk.
Proprioception informs the source distribution. Vision and proprioception condition the downstream generator.
01

A state-informed prior

A two-layer MLP predicts a mean and diagonal variance from proprioception. Each time step samples independent noise from the shared distribution.

z₀ = μ(s) + σ(s) ⊙ ε
02

Joint, complementary training

Flow matching learns action refinement. Negative log-likelihood supervises the source density. Contrastive alignment pairs source samples with expert actions.

L = Lflow + βLNLL + αLalign
03

Generative refinement

The source sample passes to a generator conditioned on both vision and state. The same prior improves flow-matching and diffusion-bridge variants in the paper.

source → generator → action chunk
How does LeaP differ from other starting points?
Three source strategies: a standard Gaussian, a deterministic observation-derived source, and LeaP’s learned state-conditioned Gaussian.
LeaP learns both where source samples should lie and how much they should vary with state.
03 / RESULTS

A better source.
Better robot policies.

Evaluated on 15 simulation tasks and three real-world manipulation tasks. Reported results are from the paper.

Average task success

Higher is better ↑

THE SOURCE MAKES A DIFFERENCE

+25.5 pp

Same architecture.
A learned starting point.

NoPrior uses the same encoder, flow-matching generator and pipeline as LeaP, but starts from a standard Gaussian. LeaP improves average success from 56.1% to 81.6%.

Explore the full experiments
Franka Research 3 evaluation. LeaP succeeds on Pick Cube, Close Box and Pick and Place Sandbags at 70%, 80% and 90%. NoPrior scores 50%, 35% and 55%.
Real-world evaluation on a Franka Research 3: rigid-object picking, articulated manipulation and deformable-object handling. Average success: LeaP 80.0%, A2A 68.3%, VITA 56.7%, NoPrior 46.7%.
What matters in the prior? Explore the ablations.

Learning a state-adaptive variance matters: LeaP reaches 85.3% on the three-task ablation, compared with 78.0% for a mean-only source and 62.7% for a learned mean with fixed variance. The prior also improves diffusion-bridge success from 68.7% (deterministic prior mean) to 76.7%.

Ablations use three representative tasks and one random seed; they are trends, not multi-seed estimates.

Ablations comparing source input, prior distribution, training losses and generator families.

Scope: evaluated on tabletop manipulation. Transfer to vision-language-action models or dexterous hands has not yet been established.

04 / OPEN SOURCE

Start exploring
with LeaP.

Policy code, RoboTwin integration, training configuration and CPU regression tests are available on GitHub.

Get the code
CITATION

Build on this work.

@article{dai2026leap,
  title   = {Where Should Action Generation Begin? A Learnable Source
             Prior for Generative Robot Policies},
  author  = {Dai, Meipo and Zhuang, Qiyuan and Xu, He-Yang and
             Shuai, Ying-Jie and Wang, Yijun and Dou, Qi and Wei, Xiu-Shen},
  journal = {arXiv preprint arXiv:2606.17408},
  year    = {2026},
  url     = {https://arxiv.org/abs/2606.17408}
}

Accepted at CoRL 2026. The citation above links to the available arXiv version.