DOMINO success rate
42.83%
+10.26 points vs. LingBot-VA (32.57%)
RoboTwin 2.0 average success rate
92.30%
+1.09 points vs. LingBot-VA (91.21%)
DOMINO Level 1 → 2 transfer, no adaptation
29.4%
vs. 16.9% for LingBot-VA; Level 3: 9.7% vs. 6.8%
Imagination chunks per episode
4.49
vs. 6.22 for LingBot-VA, on episodes both solve
Method
We propose an event-aligned visual action reasoning framework that organizes visual prediction according to the reasoning demands of robot manipulation. Rather than treating visual prediction as a uniform continuation of the sampled trajectory, our framework structures visual reasoning around task-relevant interactions and the visual context needed to support the corresponding actions. It consists of event-aligned visual-action reasoning and execution validity prediction.
Three rules for event-aligned chunks
-
Rule 1
Event boundaries
Keep each chunk within one event: one chunk per transition, one or more per interaction.
-
Rule 2
Phase-dependent density
Keep interaction evidence at native resolution and sample surrounding motion more sparsely.
-
Rule 3
Interaction boundaries
Preserve the first and last frame of each interaction to retain its starting and ending context.
Results
We evaluate on DOMINO (35 dynamic-manipulation tasks) and RoboTwin 2.0 (50 bimanual tasks). SR denotes success rate; MS denotes manipulation score.
Benchmarks
| Method | DOMINO | RoboTwin 2.0 | |||
|---|---|---|---|---|---|
| SR | MS | Clean SR | Rand. SR | Avg. SR | |
| π0 | 8.17 | 23.96 | 65.92 | 58.40 | 62.16 |
| π0.5 | 9.63 | 26.17 | 82.74 | 76.76 | 79.75 |
| PUMA | 17.20 | 34.97 | – | – | – |
| DynamicWAM* | 38.20 | 53.20 | – | – | – |
| ABot-M0 | – | – | 81.20 | 80.40 | 80.80 |
| Motus | – | – | 88.66 | 87.02 | 87.80 |
| ImageWAM | 18.86 | 35.47 | 93.20 | 93.56 | 93.38 |
| Fast-WAM | 19.09 | 35.33 | 91.88 | 91.78 | 91.83 |
| LingBot-VA | 32.57 | 45.55 | 92.93 | 91.55 | 92.24 |
| 92.18† | 90.24† | 91.21† | |||
| Ours | 42.83 | 55.92 | 93.02 | 91.58 | 92.30 |
LingBot-VA on RoboTwin: published results (upper row) and our re-evaluation under the same inference setting as ours (†). *DynamicWAM uses 150 clean and 150 randomized demonstrations per task.
Our method reaches 42.83% SR on DOMINO, a 10.26-point gain over LingBot-VA, and 92.30% average SR on RoboTwin 2.0, up 1.09 points over LingBot-VA (91.21%) under the same evaluation setting.
Transfer across DOMINO levels
| Method | L1 → L2 | L1 → L3 | ||
|---|---|---|---|---|
| SR | MS | SR | MS | |
| PUMA | 10.5 | 30.28 | 4.6 | 20.16 |
| ImageWAM | 13.3 | 30.44 | 2.8 | 15.09 |
| Fast-WAM | 13.5 | 31.56 | 2.4 | 15.34 |
| LingBot-VA | 16.9 | 30.37 | 6.8 | 19.36 |
| Ours | 29.4 | 45.42 | 9.7 | 24.23 |
Level 1 policies evaluated on ten tasks at Levels 2 and 3. Only PUMA uses target-level adaptation (LoRA).
Our policy achieves the highest SR and MS at both levels without target-level adaptation. Level 3 success remains low across all methods.
What matters: evidence, context and validity
| Sampling | Granularity | Placement | Slots | Share in evidence | DOMINO SR | MS |
|---|---|---|---|---|---|---|
| LingBot-VA (reference) | Episode | Uniform | 291,296 | 32.35% | 45.55 | |
| Episode Uniform | Episode | Uniform | 211,664 | 38.25% | 41.32 | |
| Event Uniform | Event | Uniform | 211,664 | 47.55% | 52.26 | |
| Reduced Evidence | Event | Transition-Protected | 268,592 | 23.05% | 48.36 | |
| Reduced Context | Event | Evidence-protected | 173,216 | 67.03% | 52.99 | |
| Event-aligned (Ours) | Event | Evidence-protected | 211,664 | 54.86% | 55.92 | |
| Event-aligned w/o execution validity head | 211,664 | 52.33 | ||||
| Event-aligned w/ execution validity head | 211,664 | 55.92 | ||||
Slots include padded repetitions; share in evidence measures how many fall within interaction evidence. SR bars use a 0–50 scale; the line marks LingBot-VA (32.57). All variants except the LingBot-VA reference and no-head ablation use the validity head.
Event alignment, dense interaction evidence and surrounding context all improve performance. The execution validity head raises SR from 39.11% to 42.83%.
Fewer imagination calls, fewer commands
On 738 episodes both methods solve, ours uses fewer visual rollouts and action commands. Values are averaged across tasks.
Demos
Choose a DOMINO task to compare rollouts and our model’s imagined future.
Rollouts: LingBot-VA vs. ours
Prediction vs. reality
The real execution (left) next to our model's imagined future (right). While each chunk executes, the right half shows that chunk's last predicted frame. For visualization, predicted frames are decoded with video tokens integrated to s = 1.0; inference uses s = 0.6.
Real-world experiments
We fine-tune on a Unitree G1 humanoid with Dex1-1 grippers for three bimanual tasks: Tennis Ball Insertion, Bimanual Cup Placement and Drawer Manipulation.
| Method | Tennis Ball Insertion | Bimanual Cup Placement | Drawer Manipulation |
|---|---|---|---|
| LingBot-VA | 44% | 24% | 64% |
| Ours | 52% | 28% | 56% |
Success rate on three real-world tasks, 25 rollouts each.
Real-world task examples
BibTeX
@misc{yang2026eventaligned,
title = {Event-Aligned Visual Action Reasoning for World Action Models},
author = {Yang, Xiaomeng and Wu, Yushu and Gao, Yi and Lei, Yuhao and Zhang, Xuan and Zhao, Pu and Wang, Yanzhi},
year = {2026},
eprint = {2610.09427},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2610.09427}
}