Semantic Anchoring for Robotic Action Representations

Yuan Xu*,1, Youheng Shi*,1, Chengyang Li2,3, Wentao Zhu2, Yizhou Wang1
1Peking University, 2Eastern Institute of Technology, Ningbo, 3Shanghai Jiao Tong University
*Indicates Equal Contribution
Teaser Image

TL;DR: VLA fine-tuning on limited robot data erodes its inherited semantic structure and undermines generalization. Inspired by mirror neuron theory, we probe this erosion and reveal its correlation with model performance. We introduce a plug-and-play semantic alignment method that consistently improves performance on simulation benchmarks and real-robot tasks.

Video

Abstract

Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structure captured by pretrained encoders. Systematic probing confirms that this structure erodes during fine-tuning, and that its quality synchronizes with both task success and out-of-distribution generalization.

We introduce a plug-and-play method that anchors action representations to a semantic manifold while decomposing them into a shared semantic channel and a private execution channel—all discarded at inference, leaving the deployed model unchanged. Validated on different VLA backbones across simulation and real-world benchmarks, our method yields up to +18.7% on real-world in-distribution tasks and +21.5% on OOD generalization.

In-Distribution Evaluation

Fruit Pick & Place
"pick up the grapes and place it on the plate"

Fruit Pick & Place
"pick up the banana and place it on the plate"

Dish Stacking
"stack the cups"

Dish Stacking
"stack the plate onto the other plate"

Box Packing
"place the toy bear into the box and close both side flaps"

Box Packing
"place the grapes into the box and close both side flaps"

Cabinet Storage
"place the sponge in the cabinet and close the door"

Cabinet Storage
"place the cup in the cabinet and close the door"

OOD Generalization

Building on the pick-and-place family, we evaluate five out-of-distribution axes: Spatial Variation, Novel Object, Visual Distraction, Language Variation, and Compositional Task.

1. Spatial Variation

2. Novel Object

3. Visual Distraction

4. Language Variation

5. Compositional Task

Baseline vs. Ours

More Accurate Execution

Ours (Success)

π₀ Baseline (Failure)

Better Intention Understanding

Ours (Success)

π₀ Baseline (Failure)

BibTeX

@misc{xu2026semanticanchoringroboticaction,
      title={Semantic Anchoring for Robotic Action Representations},
      author={Yuan Xu and Youheng Shi and Chengyang Li and Wentao Zhu and Yizhou Wang},
      year={2026},
      eprint={2607.13597},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2607.13597},
}