Preprint · 2026

Event-Aligned Visual Action Reasoning for World Action Models

Northeastern University

Twelve sampled frames over a demonstration with approach, grasp, transport and release phases, shown twice. Episode Uniform, as in LingBot-VA, samples at fixed intervals, so most frames fall in the long approach and transport transitions. Event-Aligned, ours, places the frames around the grasp and release interactions and their evidence phases.

(a) Visual-action chunk granularity

Bar chart of DOMINO success rate at a matched budget of training slots: Episode Uniform 29.17 percent, Event-Aligned (ours) 42.83 percent.

(b) DOMINO success rate

Interaction-dependent temporal granularity of visual reasoning. (a) Our event-aligned visual action reasoning framework organizes visual-action targets around interaction events and transitions. (b) DOMINO success rates under a matched sequence budget of training visual-action slots.

DOMINO success rate

42.83%

+10.26 points vs. LingBot-VA (32.57%)

RoboTwin 2.0 average success rate

92.30%

+1.09 points vs. LingBot-VA (91.21%)

DOMINO Level 1 → 2 transfer, no adaptation

29.4%

vs. 16.9% for LingBot-VA; Level 3: 9.7% vs. 6.8%

Imagination chunks per episode

4.49

vs. 6.22 for LingBot-VA, on episodes both solve

Method

We propose an event-aligned visual action reasoning framework that organizes visual prediction according to the reasoning demands of robot manipulation. Rather than treating visual prediction as a uniform continuation of the sampled trajectory, our framework structures visual reasoning around task-relevant interactions and the visual context needed to support the corresponding actions. It consists of event-aligned visual-action reasoning and execution validity prediction.

Three rules for event-aligned chunks

  1. Rule 1

    Event boundaries

    Keep each chunk within one event: one chunk per transition, one or more per interaction.

  2. Rule 2

    Phase-dependent density

    Keep interaction evidence at native resolution and sample surrounding motion more sparsely.

  3. Rule 3

    Interaction boundaries

    Preserve the first and last frame of each interaction to retain its starting and ending context.

(a) Event-aligned visual-action chunk construction and joint training

(b) Visual-guided action generation with execution validity prediction

Interaction-Structured Visual Reasoning framework. Event-aligned chunks allocate visual reasoning according to interaction structure, while execution validity prediction identifies which generated actions to execute.

Results

We evaluate on DOMINO (35 dynamic-manipulation tasks) and RoboTwin 2.0 (50 bimanual tasks). SR denotes success rate; MS denotes manipulation score.

Benchmarks

MethodDOMINORoboTwin 2.0
SRMSClean SRRand. SRAvg. SR
π08.1723.9665.9258.4062.16
π0.59.6326.1782.7476.7679.75
PUMA17.2034.97–––
DynamicWAM*38.2053.20–––
ABot-M0––81.2080.4080.80
Motus––88.6687.0287.80
ImageWAM18.8635.4793.2093.5693.38
Fast-WAM19.0935.3391.8891.7891.83
LingBot-VA32.5745.5592.9391.5592.24
92.18†90.24†91.21†
Ours42.8355.9293.0291.5892.30

LingBot-VA on RoboTwin: published results (upper row) and our re-evaluation under the same inference setting as ours (†). *DynamicWAM uses 150 clean and 150 randomized demonstrations per task.

Our method reaches 42.83% SR on DOMINO, a 10.26-point gain over LingBot-VA, and 92.30% average SR on RoboTwin 2.0, up 1.09 points over LingBot-VA (91.21%) under the same evaluation setting.

Transfer across DOMINO levels

MethodL1 → L2L1 → L3
SRMSSRMS
PUMA10.530.284.620.16
ImageWAM13.330.442.815.09
Fast-WAM13.531.562.415.34
LingBot-VA16.930.376.819.36
Ours29.445.429.724.23

Level 1 policies evaluated on ten tasks at Levels 2 and 3. Only PUMA uses target-level adaptation (LoRA).

Our policy achieves the highest SR and MS at both levels without target-level adaptation. Level 3 success remains low across all methods.

What matters: evidence, context and validity

SamplingGranularityPlacementSlotsShare in evidenceDOMINO SRMS
LingBot-VA (reference)EpisodeUniform291,29632.5745.55
Episode UniformEpisodeUniform211,66429.1741.32
Event UniformEventUniform211,66439.6052.26
Reduced EvidenceEventTransition-Protected268,59236.1748.36
Reduced ContextEventEvidence-protected173,21637.8352.99
Event-aligned (Ours)EventEvidence-protected211,66442.8355.92
Event-aligned w/o execution validity head211,66439.1152.33
Event-aligned w/ execution validity head211,66442.8355.92

Slots include padded repetitions; share in evidence measures how many fall within interaction evidence. SR bars use a 0–50 scale; the line marks LingBot-VA (32.57). All variants except the LingBot-VA reference and no-head ablation use the validity head.

Event alignment, dense interaction evidence and surrounding context all improve performance. The execution validity head raises SR from 39.11% to 42.83%.

Fewer imagination calls, fewer commands

On 738 episodes both methods solve, ours uses fewer visual rollouts and action commands. Values are averaged across tasks.

Visual rollouts requested per episode
LingBot-VA6.22
Ours4.49
Action commands issued per episode
LingBot-VA167.18
Ours106.46

Demos

Choose a DOMINO task to compare rollouts and our model’s imagined future.

Rollouts: LingBot-VA vs. ours

Prediction vs. reality

The real execution (left) next to our model's imagined future (right). While each chunk executes, the right half shows that chunk's last predicted frame. For visualization, predicted frames are decoded with video tokens integrated to s = 1.0; inference uses s = 0.6.

Real-world experiments

We fine-tune on a Unitree G1 humanoid with Dex1-1 grippers for three bimanual tasks: Tennis Ball Insertion, Bimanual Cup Placement and Drawer Manipulation.

MethodTennis Ball InsertionBimanual Cup PlacementDrawer Manipulation
LingBot-VA44%24%64%
Ours52%28%56%

Success rate on three real-world tasks, 25 rollouts each.

Real-world task examples

Tennis Ball Insertion
Bimanual Cup Placement
Drawer Manipulation

BibTeX

@misc{yang2026eventaligned,
  title         = {Event-Aligned Visual Action Reasoning for World Action Models},
  author        = {Yang, Xiaomeng and Wu, Yushu and Gao, Yi and Lei, Yuhao and Zhang, Xuan and Zhao, Pu and Wang, Yanzhi},
  year          = {2026},
  eprint        = {2610.09427},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2610.09427}
}