Pith. sign in

Paper Citation Record · LEDGER

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

As of 20 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 1 inbound Pith citation observation for arXiv:2605.17336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17336 v1

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T12:52:15.138790Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T08:56:16.261660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 119 outbound references displayed

  • verified exact35
  • verified fuzzy64
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73b197c5-73c6-4e9d-93e1-e22169c2f196 · outbound

This paper cites Multimodal visual- tactile representation learning through self-supervised con- trastive pre-training.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Multimodal visual- tactile representation learning through self-supervised con- trastive pre-training

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.313189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:364d373e18c9521103d300f889e6c9f243460dea5d7b8372e4eb6bcd26b0e4ce

Observation adf387c3-8eda-4de0-aa5f-9ee7ba3ee031 · outbound

This paper cites Bind- ing touch to everything: Learning unified multimodal tactile representations.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Bind- ing touch to everything: Learning unified multimodal tactile representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.332574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:9749104aad37a32787713b50807f0b8d5f52fc67d866f8ceeb06861bfedf696b

Observation 2b9098bd-d404-46b5-8fe2-18a1a02738fa · outbound

This paper cites A Touch, Vision, and Language Dataset for Multimodal Alignment.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms A Touch, Vision, and Language Dataset for Multimodal Alignment

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.469458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:868033f6a2bd0bdaa2877f560a183d77d34a555e609643d2fdffa15549a77050

Observation 6ba9885e-dfb5-41ea-9ee5-fdb1faf2e382 · outbound

This paper cites Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.444946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:7425a85a0deb9fb6bfcc7a03b3edb8aa43b150356716c886720ead20f5ac3517

Observation f713ced7-563e-4082-9f38-a98503453aac · outbound

This paper cites CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.441507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:6c5d46128d541388805d1b0165eda47e23a82c070377388f5259624d9aa5bfb8

Observation b0b65d4d-e417-4d76-bf5c-82aa7c0b8a5c · outbound

This paper cites Touch100k: A large-scale touch-language-vision dataset for touch-centric multimodal representation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Touch100k: A large-scale touch-language-vision dataset for touch-centric multimodal representation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.261989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:173f9850800e861fae55550cd776f12dd0e202034041d8437a70bad44ed66117

Observation 5cc4fca7-2ba3-4457-8197-de2dad359cf1 · outbound

This paper cites VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.455775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:aef0c9cedd67ba3e89ba4612fc9c1995fdaced54dad39cae60d366c3c8185089

Observation 0fdca6f4-85a3-4da9-adb4-c103fec5f06d · outbound

This paper cites Universal Visuo-Tactile Video Understanding for Embodied Interaction.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.369400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:c4588cba43eb722707c4b0a9679a54fae6992685e1dbd94a7b58e2322ddc9a30

Observation 053b0f1b-396c-4ab1-b1bd-ea3a3a42a4e7 · outbound

This paper cites AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.424927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:b1e1fc11288c438abc62a2920c9960eeb8593459ee070d101af0918830233b18

Observation 2f92aeea-fef1-49d9-8b22-ff6079330066 · outbound

This paper cites Vitac: Feature sharing between vision and tactile sensing for cloth texture recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Vitac: Feature sharing between vision and tactile sensing for cloth texture recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.242822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:cbe4429b90925086e0b116c5e7879663d67d65eee27f45caf44bb70041d8d626

Observation f6060e86-4763-4035-89aa-975db3b13e26 · outbound

This paper cites Can vision feel touch? tactile-aware visual grasping for transparent objects.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Can vision feel touch? tactile-aware visual grasping for transparent objects

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.256409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:32685d43066b6520cbf4fffc2079f510142d72050e1a1d29fb596a3d1b5f8810

Observation 876a8f97-99a6-4170-ad27-e55f9a8fa318 · outbound

This paper cites Surformer v1: Transformer-based surface classification using 18 tactile and vision features.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Surformer v1: Transformer-based surface classification using 18 tactile and vision features

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.272654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:977773b238d79abe9d12d145141dbb4ddae95e527f7e6ac621a5112813616cab

Observation 3123f774-dbc7-4db8-88c4-c134580d2e6d · outbound

This paper cites Ra- touch: Retrieval-augmented touch understanding with enriched visual data.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Ra- touch: Retrieval-augmented touch understanding with enriched visual data

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.335353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:82b7f0429d22a62fcfa5c9c46c48ccc835b38d2a8944ccd44ed9d8bcbe3cb5b6

Observation 2d1e9918-18cd-4fe4-a8ac-22a3e43ac760 · outbound

This paper cites A survey of deep learning and its applications: a new paradigm to machine learning.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms A survey of deep learning and its applications: a new paradigm to machine learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.224263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:79651ce5bcdcf78a8722284c17171ee01105869a58f759c981d8eba1ddaa827c

Observation 024ce2a7-5c48-4f46-b36a-383dafab748d · outbound

This paper cites Attention is all you need.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Attention is all you need

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.219299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:3eb3c79576c4511079b982bf797df486226c07ae8026a5765ef4595b0c069c21

Observation b9daba0d-2468-4aa1-9842-3f7e6767f438 · outbound

This paper cites A Survey of Large Language Models.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms A Survey of Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.476296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:085094f430035b6d8fb87e297b9ffb27905c789e66cb06410e19fca81314b69f

Observation 8c5af53e-3c63-4831-935e-a6b0015ab036 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Imagebind: One embedding space to bind them all

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.226512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:19d5eed8ca8cd7b114118a7464f195d366401c3ae7ea114c6e78144ac433d3b0

Observation 13763f84-16cc-469e-874b-16ebf216c2cd · outbound

This paper cites Transformer in Touch: A Survey.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Transformer in Touch: A Survey

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.358964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:cc38ebc9bd5190c582b66165bdd98f9c09207df73858977ee36c4fe2378c61ec

Observation 4d8ed6f1-95f4-42eb-8df0-916f05e72178 · outbound

This paper cites Tactile data generation and applications based on visuo-tactile sensors: A review.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Tactile data generation and applications based on visuo-tactile sensors: A review

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.214178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:9f45195b9e098cd38e581489b34a89ef97ef926dfae87d38765a8664aca9250c

Observation 4fa798e2-829e-4e91-9529-28a09ec146d7 · outbound

This paper cites Deep residual learning for image recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Deep residual learning for image recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.216941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:3aad83a67de5a2aa2297b61fee6094437a326f94c448eaa373e65a8ce65f1b3b

Observation 02b30ea2-6b8b-447e-ae3e-5542016dec47 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.431364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:7420bf9b7d555bbd5f40dc94b6c3f07a96f3862c6654f0750491909253bb85da

Observation 3fa16c51-3ef9-4360-918e-0f31498fd5fd · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.206747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:6338f123fc1924b32e94ac747c211632b3017dae1fcfb067f5490e1102657874

Observation 55a30c0e-0784-4986-823e-f12d8a458fb9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Learning transferable visual models from natural language supervision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.211914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:e69fd2117c1815fc38ba226579ca13eae6b18c58c338284db6e91fb94c262f80

Observation 2a553c4a-33c5-49f5-8367-5d1545baf7db · outbound

This paper cites Recent progress in pressure and temperature tactile sensors: principle, classification, integration and outlook.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Recent progress in pressure and temperature tactile sensors: principle, classification, integration and outlook

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.162201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:e5157be96da1240cc370e6657f1efb306af92a9145d7078f71a90e87a99f5ee0

Observation 1670e543-c02c-4337-94df-03731d6fb6f0 · outbound

This paper cites Classification of vision-based tactile sensors: A review.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Classification of vision-based tactile sensors: A review

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.195251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:2ffac1f3fd049a7bc276967f979cfc6ede6427ca560da719b0d6a992e14cf37f

Observation a55e4c2d-8e4f-48cb-a9a3-60d138ce9de3 · outbound

This paper cites Tactile sensors: A review.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Tactile sensors: A review

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.197831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:b91d92b9dda68193fb076795de5ba08e414f96af28d04cc93d772ce67ca03b77

Observation e5dacb2a-2de9-4efb-b22e-c8f45905d93f · outbound

This paper cites Recent progresses on flexi- ble tactile sensors.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Recent progresses on flexi- ble tactile sensors

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.203467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d3371eeb8887129255ceb739576eafc19c33ba2c29b5ef35c3130fc1d6dd848b

Observation 58e2d5c0-6e37-4e9a-a7c6-318571f8119d · outbound

This paper cites A review of tactile information: Perception and action through touch.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms A review of tactile information: Perception and action through touch

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.200276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:e7a52d151e2d561b87ac65376731e0109f7ad57e88e3544f1a78f6324172de3f

Observation 8de26369-785f-4e6e-aa26-05687eef0d4a · outbound

This paper cites Biomimetic tactile sensors and signal processing with spike trains: A review.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Biomimetic tactile sensors and signal processing with spike trains: A review

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.154933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:392bb9610d85f70eb05de7ecd4c2955f035a169afa35eb1ef237369220a7b6b0

Observation ac83c240-11db-4286-b9ae-8071b23269a5 · outbound

This paper cites Gelsight: High- resolution robot tactile sensors for estimating geometry and force.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Gelsight: High- resolution robot tactile sensors for estimating geometry and force

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.186907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:9b1e521430448467c49f949d80d26e2de6bf960d470d3d7a9df8550bccd654a5

Observation 154c7375-7705-4291-8b9a-f76092fc3eea · outbound

This paper cites Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.179856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:374dd7f2101cd9a1ea797ef651cb05878ed9b38795cd9024cc13fadfcda40f3c

Observation 25da002a-eba1-4c43-a234-b8f8fb15ea81 · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.397992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d55374e91899ca536cea7fbc0ecddcddd8617ce481403718003b611c136d0b59

Observation 672426bd-70d6-4276-9f0b-0731d5ecf46a · outbound

This paper cites Gelstereo 2.0: An improved gelstereo sensor with multimedium refractive stereo calibration.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Gelstereo 2.0: An improved gelstereo sensor with multimedium refractive stereo calibration

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.174749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:67c8cfd24348864ccb4b12660e666419f0ca12980e51a20fbb467fcea6e65fa4

Observation 3713c724-ab22-4357-85fc-9d536ec1a20b · outbound

This paper cites Gelslim 3.0: High- resolution measurement of shape, force and slip in a compact tactile-sensing finger.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Gelslim 3.0: High- resolution measurement of shape, force and slip in a compact tactile-sensing finger

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.177123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:31253d0f872770cd47b2e14c569fd11a65dbed4b58d14941b9d3450cf32a4698

Observation c96ad874-460b-41c1-8e57-85232037f780 · outbound

This paper cites Omnitact: A multi-directional high-resolution touch sensor.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Omnitact: A multi-directional high-resolution touch sensor

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.182095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:87f5c64337642444e54ec1f70793a4d31e9f6bfff470e76faa4b8fac077ebc72

Observation 39a22fcc-26b9-4f02-a8b7-9dafdda2de5e · outbound

This paper cites Seeing through your skin: Rec- ognizing objects with a novel visuotactile sensor.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Seeing through your skin: Rec- ognizing objects with a novel visuotactile sensor

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.189509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:af0523d95e3232d75b26b5fc88ce271f2a938eb70672844b3f15c6bd0f70cb9d

Observation 57cd68ab-3690-44bc-9600-ebf94cba0b68 · outbound

This paper cites Multimodal alignment and fusion: A survey.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Multimodal alignment and fusion: A survey

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.438009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:b0a29faa78e310b45021a8a90daec2ff6c76275779bf0d69c4ce7c9506c46311

Observation 810e2937-408d-49bb-aa1f-adc64b995e36 · outbound

This paper cites A survey on multimodal large language models.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms A survey on multimodal large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.141359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d4e8e0d99653fb468e98e170ef1d5b7006fd22fcd801917ce5bfbc73b53692f6

Observation cb804b5d-9a17-4448-acfe-77880908d79d · outbound

This paper cites Vhtformer: A joint query perception method for visual-haptic- textual information based on transformer.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Vhtformer: A joint query perception method for visual-haptic- textual information based on transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.169867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:c27b819613846df5b66ccd97f981b7f06f4573cf51651c4f56e551c194921a5b

Observation 94f0dc05-0666-430c-a358-2320e47a5858 · outbound

This paper cites Visual–tactile fusion for object recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Visual–tactile fusion for object recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.167227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:de69c42eb5aefec72a4a7acd86fabb1a279ac43e826b6f2acbdd2b5e0c138ffd

Observation ef08472e-e22d-4dcd-b267-3936354e20a2 · outbound

This paper cites The Feeling of Success: Does Touch Sensing Help Predict Grasp Outcomes?.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms The Feeling of Success: Does Touch Sensing Help Predict Grasp Outcomes?

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.452177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:ef84a0de570effd13bd1e030abee0d075bf6220fa9e6228f8dd9ba5decadf4e4

Observation c5dcf87d-371b-4b67-9834-c0b6e66ba50b · outbound

This paper cites Connecting look and feel: Associating the visual and tactile properties of physical materials.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Connecting look and feel: Associating the visual and tactile properties of physical materials

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.164939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:184a76d0e8461e9678afc81458f1fdd6d05b89c2e4b56da0e8c42753d6c49f87

Observation 7e38ce2b-29ea-43aa-8d73-900ee2a87e8f · outbound

This paper cites More than a feeling: Learning to grasp and regrasp using vision and touch.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms More than a feeling: Learning to grasp and regrasp using vision and touch

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.172050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:23de49a25888b7221912983d98e8b6b9e54a8ed439e082b30b93c97584ade821

Observation 6b032316-4610-407e-88d1-68530e4039e7 · outbound

This paper cites Multimodal grasp data set: A novel visual–tactile data set for robotic manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Multimodal grasp data set: A novel visual–tactile data set for robotic manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.184429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:e6b3ee9a216c676540971c77d70ddd6c047ec18da1aeeff495d1c3bc982e8da9

Observation dbf6ca89-2639-49f2-be1f-5757b726d16d · outbound

This paper cites Connecting touch and vision via cross-modal prediction.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Connecting touch and vision via cross-modal prediction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.209331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:252f45d9c5f4560d14d1c97a0046a71c6a1bfc63e869e82b6d470f2f67c09eca

Observation 46b90da7-386d-4afa-adcf-67ed04b1e28d · outbound

This paper cites Touch and Go: Learning from Human-Collected Vision and Touch.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Touch and Go: Learning from Human-Collected Vision and Touch

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.394578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:3a867c03824469f49168cb98cbdbf785e0888a2375cc22ee57bd7ee014d6f61f

Observation 99c83d13-dc32-464c-9978-3cb1add6a24f · outbound

This paper cites Controllable visual-tactile synthesis.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Controllable visual-tactile synthesis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.152105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:4ed57be8537a1f3a948b8eb634621a1947af1a14cca2be92b574666ac6bc2eed

Observation ca883d37-fb45-4b21-b9b0-b38c67f60d0f · outbound

This paper cites Learning to jointly understand visual and tactile signals.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Learning to jointly understand visual and tactile signals

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.267627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d1e5da75380653d762de285491d57b27f1fe232d3eb94bd22e94938a1f78a9a9

Observation c18a54dc-71f2-4cef-aca0-1266a1a652c7 · outbound

This paper cites Touch in the wild: Learning fine-grained manipulation with a portable visuo-tactile gripper.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Touch in the wild: Learning fine-grained manipulation with a portable visuo-tactile gripper

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.427943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:083caf2c636043420f6718beb1de62486039755e3ba012b3daf39a99dce8fbdf

Observation 0fc18a5f-ebd9-477d-94df-75c80cce0ad2 · outbound

This paper cites Octopi: Object Property Reasoning with Large Tactile-Language Models.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Octopi: Object Property Reasoning with Large Tactile-Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.490680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:94a3045b3b78f7a694debcb5c17be788a9b62fc5832208deb235b51a281dd8de

Observation 00e1411f-9230-4952-9ec1-a50cb38c57c2 · outbound

This paper cites Stola: Self- adaptive touch-language framework with tactile common- sense reasoning in open-ended scenarios.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Stola: Self- adaptive touch-language framework with tactile common- sense reasoning in open-ended scenarios

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.466118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:5e4aa4998d45a37ff9d98d4a4e603b289b11fe3a156b77475f02f0057465697d

Observation 92fb07c9-f006-4e30-8b3c-e5f9f7d16262 · outbound

This paper cites Multi-modal representation learning with tactile data.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Multi-modal representation learning with tactile data

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.324386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:73bd0ae85d81cf9d0fdbee6e92268052f53343c63ade40834ccf51969e1bc37b

Observation a4d80c0f-39c1-45a5-a622-c590365a7c51 · outbound

This paper cites Damf: A semantic-guided dynamic attention frame- work for visual-haptic-textual multimodal fusion.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Damf: A semantic-guided dynamic attention frame- work for visual-haptic-textual multimodal fusion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.326884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:01d00d8311e9b8097d12452b10cddaf03eb796a58b3136296d9abbd61bd3fe60

Observation cc9b4178-7991-4d81-b02e-5187e6643a23 · outbound

This paper cites Tvt- transformer: A tactile-visual-textual fusion network for object recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Tvt- transformer: A tactile-visual-textual fusion network for object recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.329658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:f5e7b8834b34470d1d68afd231a8afebb72e5d6b346f5943dbe78096b131f534

Observation 29dc1dd8-680d-40e4-afc9-dece60c3fdf3 · outbound

This paper cites Omnivtla: Vision- tactile-language-action model with semantic-aligned tactile sensing.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Omnivtla: Vision- tactile-language-action model with semantic-aligned tactile sensing

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.473136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:ee8281517b5c9bacf47ca5486621c14c07f09044b107a09f050ea4ccb1cbceb3

Observation 76baf2c7-8471-43f3-baca-62f51dc01150 · outbound

This paper cites ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.487475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:76ef855db68a93f781fe27a652f8fe1eff385083538d96533fb87e34f4a4f951

Observation 03e78128-9def-422a-aaf8-1e7deaf1c2e4 · outbound

This paper cites Objectfolder 2.0: A multisensory object dataset for sim2real transfer.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Objectfolder 2.0: A multisensory object dataset for sim2real transfer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.321376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:0a7d18941574c1b1d4c8dd2b153270d6e10a0bc6347ca9bed722d5080631d3d8

Observation 62d8a576-bf60-45dd-b160-f93a21456c43 · outbound

This paper cites The objectfolder benchmark: Multisensory learning with neural and real objects.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms The objectfolder benchmark: Multisensory learning with neural and real objects

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.318283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:c4023b9f68c9246885d08b0adb42341e3ada95808eac0fa9f9be0d90e852bb47

Observation 587a85ed-f281-4bd7-806c-d76dc70a4c44 · outbound

This paper cites TLA: Tactile-Language-Action Model for Contact-Rich Manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms TLA: Tactile-Language-Action Model for Contact-Rich Manipulation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.411212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:2597b3f29b5419f17ae68715abb0aaecc16331b63c6c0e74eeddb572635b8c31

Observation 4d1a7b0d-0d05-4d9c-bd25-c2e2545c1ab4 · outbound

This paper cites Freetacman: Robot-free visuo-tactile data col- lection system for contact-rich manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Freetacman: Robot-free visuo-tactile data col- lection system for contact-rich manipulation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.484237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d4f635ab1cd4b920be1f5c1864ca627104b3a88e63d0c9167ab132bef7270ca4

Observation e4824277-f129-43dd-895c-25278fc67ee7 · outbound

This paper cites Opentouch: Bring- ing full-hand touch to real-world interaction.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Opentouch: Bring- ing full-hand touch to real-world interaction

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.421571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:751bd0959c15599c37d6011b87801ed15a9b1878b531f25bbce42449e9aeaedd

Observation 0d1f0295-37c6-419f-9af6-75f66bb10aba · outbound

This paper cites Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.434522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:5872d7d606a2b29470cef851829b3d99e555779aa1faae45a4005d029bd806bd

Observation afef0867-5d2e-4ce2-a4c5-4c8db5e74efc · outbound

This paper cites VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.383723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:2422e2a22e963e804ea505ec8905bccdb2b044757373f9f28ac0cf1e1aefa1c5

Observation 928bc7d3-1046-4f83-88b2-e81db2d2f579 · outbound

This paper cites Omnivta: Visuo-tactile world modeling for contact- rich robotic manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Omnivta: Visuo-tactile world modeling for contact- rich robotic manipulation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.307814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:3bdcb06efacb1593393ec267c5c3873d9f0f713f20f210c1cfe88dc3e7dcf45e

Observation 0ed9662e-9984-4db7-8834-bc51abbcdbf8 · outbound

This paper cites Imagenet clas- sification with deep convolutional neural networks.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Imagenet clas- sification with deep convolutional neural networks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.302359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:79d9af4197a89976e94745ab442efbae4674a47e40afebafbe38ffa06571e2a8

Observation 6e05a811-cefa-4c8a-a210-a712e6e05c0f · outbound

This paper cites See, feel, act: Hierarchical learning for complex manipula- tion skills with multisensory fusion.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms See, feel, act: Hierarchical learning for complex manipula- tion skills with multisensory fusion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.305163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:bb93a12f80ea4d77329fddb304852682a9b5842a2eae7a03be6620565fdcce6b

Observation 4142ea65-3cbc-497d-9d0f-3f37736e257c · outbound

This paper cites Mask r-cnn.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Mask r-cnn

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.315602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:fe338c7bb8cabbe3f31be1ca673c28d19e0b2d153dadc583ca9b9a187b357091

Observation edbbfebe-4634-4f12-8006-3b9d88f57b71 · outbound

This paper cites Bayesian neural networks.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Bayesian neural networks

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.296806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d42aa7565b8ba1351ef0683031b85a262ac5d0e21b992698df695bc80b94f995

Observation b7954fb9-28fa-4d0c-8c9e-4251fec120ff · outbound

This paper cites Learning cross-modal visual-tactile representation using ensembled generative adver- sarial networks.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Learning cross-modal visual-tactile representation using ensembled generative adver- sarial networks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.289223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:bdb291b34eedf1e2ab107a3ffbf87ada9b4621553549534367db37d4d0575bf1

Observation 2a3d21c6-9589-4fe5-bef3-0b4ed57a0875 · outbound

This paper cites Gen- erative adversarial nets.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Gen- erative adversarial nets

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.294276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:ef1eaf273df361f9b14fefa07cd81ddd715f35f6bb375ec5c39ec8e44af317f1

Observation 39b1df4a-cf04-472f-9b81-bc74ab882b0d · outbound

This paper cites Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.310483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:f4d15d75df6da00186566daf09b858f3b813438623fa28b2deeead7e1e102153

Observation acd8b079-2cae-4a91-ba0c-7ef6555bd397 · outbound

This paper cites Flownet: Learning optical flow with convolutional networks.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Flownet: Learning optical flow with convolutional networks

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.340402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:13bbcf73113588ed8c04f7a5403d5afcbcd9789b680b28bec3941dcab20fd69d

Observation bb738793-d227-4401-a657-29d7b91489bf · outbound

This paper cites Lifelong visual-tactile cross- modal learning for robotic material perception.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Lifelong visual-tactile cross- modal learning for robotic material perception

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.286644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:6bdb4e3723e8df3f5b2bd5a25959327d789d362f5763124b4310eff6cdc2ee40

Observation 858a1ae4-b5eb-49e9-8fa7-57e6cf0d2721 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.387438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:4373e4f07253e28c55f8eefb48befe8e43b9bc9185a4c7c08f12098894a430de

Observation f82398ff-0081-4fb1-ab27-f82298b64b71 · outbound

This paper cites Visuo-Tactile Transformers for Manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Visuo-Tactile Transformers for Manipulation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.351857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:f0321e9e8db5a617d7a35686aef9a27b7a8e28bcf5ef646eb8da7fe1c388f798

Observation b9f4ce5c-6c55-47f2-aa9d-e1554e458f4e · outbound

This paper cites Visuotactile-rl: Learning multimodal manipulation policies with deep reinforcement learning.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Visuotactile-rl: Learning multimodal manipulation policies with deep reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.157338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:a2f931cc1ba10b91755669fdf0e55c099ef70513505f7641f188dda9a18e2d74

Observation 688ff8ed-21af-4c57-b6bb-5504b6602930 · outbound

This paper cites Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:53:17.480649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:fe8a65016e7d06b6280e59eb81f10a987988dee8391fca7319bc36ac07f46f0f

Observation f15f9726-0bce-40a0-ad86-4338bd2a6a2a · outbound

This paper cites Vito-transformer: a visual-tactile fusion network for object recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Vito-transformer: a visual-tactile fusion network for object recognition

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.159840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:88f9a0e801f69a4e078f5d470c2824d59c7c863809aee97ebf343e7d7585f019

Observation a8bfa834-3a4a-4838-81ee-1240108815e0 · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Mlp-mixer: An all-mlp architecture for vision

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.149818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:5530c3f74b5193f14506558a78c2a584deceed636707192e7d60a110b352a16e

Observation e7453b5b-b37f-4ec4-9c3a-ebf087eda3c1 · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Fine-tuned clip models are efficient video learners

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.192454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:b1ba25a7a4bd98a988be1c4ba6577c3a8d9ff957f0abfb36bee9daef811f8ac7

Observation be6e87de-a4df-48fe-b935-cd669fa50c71 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.144579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:b24caed98fc4e20d327b0ee0304f0755db7fd1b6cbac64551eb1d70b6a202a75

Observation b6ef88c6-d68f-453e-80d0-539728edb12d · outbound

This paper cites Bidirectional visual-tactile cross-modal generation using latent feature space flow model.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Bidirectional visual-tactile cross-modal generation using latent feature space flow model

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.147226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:84ba6639fe5f3d8c2e8640769e1fe20ab746e59edc3bc496b02bd1770fa1b0c9

Observation 4c68853a-ae98-46d6-871b-cec59999d52f · outbound

This paper cites Demonstrating the Octopi-1.5 Visual-Tactile-Language Model.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Demonstrating the Octopi-1.5 Visual-Tactile-Language Model

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.407959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:f62fa769c54fd70129573b0efac69ada2e08f34c0d1f1692475d59ae4f25bfc4

Observation 7c6014b1-8264-4496-a9cd-47a79b93e7f4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.355261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:01f04dceb3c1b59d681f60afad497306a1c2733d2141a8b5dd9de9dd25dbae03

Observation 1c403072-4317-447c-b75e-67178daadf6d · outbound

This paper cites ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.362226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:bf79252a7cbdb5ad0dbe7407df1b59e3746d6f65811c665957c2dfe6df83b655

Observation 8a5b76d7-4c90-4e98-84c7-e92f8ae67e85 · outbound

This paper cites ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.414391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:42a858bd6b84a7c75b0740970f40c2c54037c76048cef6bfdca0a8c02e5c21b0

Observation c945e8cd-2c47-4621-b909-37441dfbb7a9 · outbound

This paper cites Object attribute recognition method integrating visual-tactile data and multi- task learning.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Object attribute recognition method integrating visual-tactile data and multi- task learning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.221772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:d49eb6782405ea13d141484260b8fa3763b83c4c91bbedaddc5807c450f1ed9f

Observation 0ace1f1b-26c1-499b-baa1-0b0bfe490ffc · outbound

This paper cites Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.404600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:5097da25e3bb2d81aad94dc9d1b13d6fbc25145fe6482a9c8da2680e8f350d43

Observation f593d785-cc20-467e-90bb-6c264d71bfcc · outbound

This paper cites Efficientnetv2: Smaller models and faster training.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Efficientnetv2: Smaller models and faster training

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.291936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:5f5d1dd45d62b4beb0dd852350b17a972de086cc78c59a536b18ba86fd9bcbfc

Observation 1e6cdd75-690d-436f-9b25-4f6cb92e5d26 · outbound

This paper cites Ulip- 2: Towards scalable multimodal pre-training for 3d under- standing.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Ulip- 2: Towards scalable multimodal pre-training for 3d under- standing

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.299484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:85ae55920d9faef99bed30082d3a1ebd12df32eb9317d066ebf2c611bb7b6d03

Observation c06afac8-7083-4411-baa8-72f1f7235935 · outbound

This paper cites ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.418098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:aea43a419f4415ffd7569887efff8daa43a8929a0dc33f79a29a5143c998867d

Observation a849b8cf-c4b0-4057-a1ff-15b7ca65662d · outbound

This paper cites Vtlg: A vision-tactile-language grasp generation method ori- ented towards task.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Vtlg: A vision-tactile-language grasp generation method ori- ented towards task

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.343169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:16e17c910236f7cf19c20601ce17c996ec9230ae0d7da23dd3b3829530d9b40a

Observation d79f0b80-3cdf-41dd-becb-0fa697b3a754 · outbound

This paper cites Dt-transformer: A text-tactile fusion network for object recognition.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Dt-transformer: A text-tactile fusion network for object recognition

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.280132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:99db831ea3462e1fd84b3813c3aa629a88897dcc773ad2456dfaa36a952ac9d8

Observation 9e82bd9d-515c-4190-a492-f3664865aed8 · outbound

This paper cites Densely connected convolutional networks.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Densely connected convolutional networks

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.229033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:7f7969ecf4663799598b3e3125a7f96d9dfedddc344a1eff53975036eeef5ab6

Observation de4803ad-5f51-431a-b09f-6bdcd8812763 · outbound

This paper cites Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.448473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:987726e7d69daa1bda3c6ed8a6ca2c935bcb21d80a1f8b1e239af2954049b6dc

Observation 7c76f681-ae5b-4355-982d-e4760d8409f4 · outbound

This paper cites VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.376885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:fa08b633554e4b5b25067093cec891eaffdedc5baca69ac11de775fc19171c68

Observation bf190bc9-9ad9-4306-a3eb-5bebdcbfd2e9 · outbound

This paper cites GPT-4o System Card.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms GPT-4o System Card

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.458924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:336f20d448618ebb00a8e4a2fb31a2173786c5260e94e902ab96e95d32070de5

Observation 75175fc8-efd1-480f-bcc8-b5e9a7c88b67 · outbound

This paper cites Dextac: Learning contact-aware visuotactile policies via hand-by-hand teaching.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Dextac: Learning contact-aware visuotactile policies via hand-by-hand teaching

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.391165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:dd8e15bd012bcaf3cc7fb62baadbd1789ee9e0ed43b19f03c6d9b7818e953733

Observation a105f1ae-f48d-4825-aa06-815d8f554219 · outbound

This paper cites Vtam: Video-tactile-action models for complex physical interaction beyond vlas.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Vtam: Video-tactile-action models for complex physical interaction beyond vlas

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.233845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:0828f1e4bdd0dba88c1f590de6032bfff45939b4d7eca7869257ab22f731b2a1

Observation 1fef7431-2e6a-4122-a5f1-569d44bc8729 · outbound

This paper cites Visual–tactile fusion material identification using dictionary learning.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Visual–tactile fusion material identification using dictionary learning

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:18.248330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:4ca9003041b8b98bd636d3908363aadff39b08c9d06bd9baa3a0182369474644

Pith citing papers

Observation 8bb4da49-7a63-44f1-a5c3-543fd3315938 · inbound

TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction cites this paper.

TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

Reference 3307

Resolution
unresolved
no resolver link, observed 2026-07-31T08:56:16.261660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:56:16.261660Z digest=sha256:58eef22f82795de6823ff6adae8655ff229ad38f25e9b2652275e159628c4890