Publication
Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning
Shilin Shan; Chuhao Zhou; Ruize Wang; Xinyan Chen; Xiangyu Chen; Xinyu Zhou; Boyu Ma; Kirishanth Chethurajah; Jingliang Li; Celeste Yuxuan Hu; Gengyu Liu-Sürücü; Guohao Chen; Tianrui Zhu; Zhen Liu; Yanjie Ze; Haoran Geng; Zhiyang Dou; Jianxin Bi; Yuejiang Liu; Jianshu Zhou; Jiachen Li; Paul Liang; Tatsuya Harada; Robert Katzschmann; Harold Soh; Na Li; Edward Johns; Danica Kragic; Jan Peters; Wojciech Matusik; Masayoshi Tomizuka; Jitendra Malik; Jianfei Yang
In: Computing Research Repository eprint Journal (CoRR), Vol. abs/2608.07558, Pages 1-53, arXiv, 2026.
Abstract
Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their
interactions with the physical world. This capability is particularly critical in contact-sensitive manip-
ulation, where successful task execution depends not only on visual perception and motion generation,
but also on force regulation and adaptive control. In this context, recent robot learning methods have
made substantial progress by integrating force, tactile, vision, language, and proprioceptive sensing
into learned manipulation policies. In parallel, many systems adopt multi-phase architectures that
combine high-level policies, action-refinement modules, and low-level controllers to bridge semantic
task understanding with reactive physical execution. Despite these advances, existing surveys have
not explicitly reviewed force- and tactile-aware robot learning from a unified perspective that jointly
captures multimodal sensing and multi-phase system design. This survey addresses this gap by
proposing TF-ART, a Tactile/Force-Aware Robot learning Taxonomy for multimodal and multi-phase
frameworks, which maps individual methods into a unified hierarchical structure. The framework
characterizes how recent works organize observation modalities, encode and fuse heterogeneous sensory
inputs, generate and refine actions across multiple phases, and connect learned policies to reactive
robot-end control. Building on this methodological view, we further examine the task settings and
infrastructure requirements of physical interaction, thereby integrating both algorithmic and practical
perspectives on force- and tactile-aware robot learning.
