Skip to main navigation Skip to search Skip to main content

Visuo-Tactile Zero-Shot Object Recognition with Vision-Language Model

  • Shiori Ueda
  • , Atsushi Hashimoto
  • , Masashi Hamaya
  • , Kazutoshi Tanaka
  • , Hideo Saito

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Tactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach leverages the zero-shot capability of VLMs to infer tactile properties from the names of tactilely similar objects. The proposed method translates tactile data into a textual description solely by annotating object names for each tactile sequence during training, making it adaptable to various contexts with low training costs. The proposed method was evaluated on the FoodReplica and Cube datasets, demonstrating its effectiveness in recognizing objects that are difficult to distinguish by vision alone.

Original languageEnglish
Title of host publication2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2024
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages7243-7250
Number of pages8
ISBN (Electronic)9798350377705
DOIs
Publication statusPublished - 2024
Event2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2024 - Abu Dhabi, United Arab Emirates
Duration: 2024 Oct 142024 Oct 18

Publication series

NameIEEE International Conference on Intelligent Robots and Systems
ISSN (Print)2153-0858
ISSN (Electronic)2153-0866

Conference

Conference2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2024
Country/TerritoryUnited Arab Emirates
CityAbu Dhabi
Period24/10/1424/10/18

ASJC Scopus subject areas

  • Control and Systems Engineering
  • Software
  • Computer Vision and Pattern Recognition
  • Computer Science Applications

Fingerprint

Dive into the research topics of 'Visuo-Tactile Zero-Shot Object Recognition with Vision-Language Model'. Together they form a unique fingerprint.

Cite this