Conference on Robot Learning · 2026

Equivariant visuotactile learning

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

WONG Lik Hang Kenny1 Yiyao Ma1 Xiu-Shen Wei2Zelong Tan1Zhuheng Song1Dongsheng Xie1 Kai Chen1 Qi Dou1

1 The Chinese University of Hong Kong

2 Southeast University

01 · Overview

Learning contact-rich skills from fewer demonstrations.

VISTA combines vision, touch, and robot state in a geometry-aware policy that transfers learned behavior across spatial configurations.

Three-minute presentation03:00

Abstract

Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a spatially robust policy under limited demonstrations a key issue. To address this, we propose VISTA, an equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning.

VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive simulation and real-world experiments show substantial gains in data efficiency and spatial robustness over strong visuotactile imitation-learning baselines.

Visuo-tactile manipulationEquivariant learningDiffusion policy
80.9%Mean simulation success
+15.9 ppOver the strongest baseline
73.3%Average real-robot success
50Demonstrations per task

02 · Method

Geometry is built into the policy.

VISTA preserves how observations and actions should transform under workspace rotations, helping the policy generalize beyond the poses present in demonstrations.

01

Lift to the sphere

Egocentric RGB and tactile images become steerable spherical tokens with a shared geometric structure.

02

Fuse contact and scene cues

Permutation-equivariant fusion lets tactile contact modulate the visual directions that matter for control.

03

Predict consistent actions

Orientation alignment and an equivariant diffusion decoder produce actions that follow workspace rotations.

Spherical fusion

Touch gains direction through vision.

The same local contact can call for a different world-frame correction depending on the end-effector pose and surrounding geometry. VISTA aligns these cues while preserving equivariance.

03 · Experiments

Spatial robustness in simulation and on hardware.

We evaluate VISTA on eight simulated tasks and three contact-rich real-robot tasks under limited demonstrations.

Simulation

Eight geometry-sensitive tasks

50 demonstrations per task

Main results

VISTA leads across tasks

Data efficiency

Strong results with half the data

Ablations

Each geometric component contributes.

Removing tactile spherical features, changing the fusion strategy, or weakening the group structure reduces performance on geometry-sensitive tasks.

04 · Real robot

Contact-rich manipulation, up close.

Choose a task to inspect successful rollouts. Paired videos keep the external and wrist-camera views synchronized.

Surface contact

VISTA maintains stable contact while following the target erasing trajectory.

Trial 01External view
Trial 02External view

05 · Citation

Build on VISTA.

The archival paper link will be added after the arXiv release.

BibTeX
@inproceedings{wong2026vista,
  title     = {Equivariant Visual-Tactile Diffusion Policy for
               Contact-Rich Manipulation},
  author    = {Wong, Lik Hang Kenny and Ma, Yiyao and Wei, Xiu-Shen
               and Tan, Zelong and Song, Zhuheng and Xie, Dongsheng
               and Chen, Kai and Dou, Qi},
  booktitle = {Conference on Robot Learning},
  year      = {2026}
}
Expanded research figure