Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:26.504997Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 5 inbound Pith citation observations for arXiv:2505.22566.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:26.504997Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T08:20:49.438388Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T10:05:40.612441Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43cfddc7-1d6d-41b0-af88-f4bf9e6f1ad5 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction A review of tactile information: Perception and action through touch
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbfc6331-cc9d-439c-a6d0-4d4eab314da0 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Task and material properties interac- tively affect softness explorations along different dimensions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a8dadab-e645-4958-b1e7-f92336479787 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4f5cbb6-2d5b-4f5f-8785-90ec4adb4fd0 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0301aa58-7e92-46ab-8142-1e633397736b · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d442cf6-3280-4c99-8959-021be48de50b · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction High- resolution image synthesis with latent diffusion models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98429ef1-db62-4cff-bb88-2344d067e4cb · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1106966-9d92-452e-9005-3dcd31fb4d2c · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c811a921-8343-401f-ae5f-591f7be8ffad · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a6ffca7-6dd9-4dc7-9421-2c0cd24964fb · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Gelsight: High-resolution robot tactile sensors for estimating geometry and force
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfea851b-8f91-4e67-a99d-ace2f346f1a7 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation acebc66d-26bc-4d24-abc6-ef6bd4c66b39 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ef809c-43e9-4f30-8cf1-93e8632aa829 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 624d1181-696b-404c-ab7e-f6392e065556 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Transferable tactile transformers for representation learning across diverse sensors and tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 793c2a97-ea4c-4375-88da-a04937c09976 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Octopi: Object Property Reasoning with Large Tactile-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdda796b-c152-44fa-914f-0ee8ff653e78 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction A touch, vision, and language dataset for multimodal alignment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7fbd7ce-31e4-40ec-bb6b-37ec7f3bed1b · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Binding touch to everything: Learn- ing unified multimodal tactile representations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea52084-156f-4997-a910-f59c5759fb43 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Sparsh: Self-supervised touch representations for vision-based tactile sensing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ebbfd123-0d6e-4b15-a32c-6d0014f2f269 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Visuo- tactile affordances for cloth manipulation with local control
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbec3edb-3dd7-4112-afbc-bf4784575848 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction A Survey of Embodied Learning for Object-Centric Robotic Manipulation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208acdd1-c51d-49ff-aee7-c0666b5a9c53 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch and go: learning from human-collected vision and touch
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 392610d5-dcb9-4574-992d-ee451a827144 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d1a6651-3972-4388-a3d7-4f43fee1f54d · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder 2.0: A multisensory object dataset for sim2real transfer
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6dacc2bd-93f8-45eb-9a08-28d9e34f710e · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction See, hear, and feel: Smart sensory fusion for robotic manipulation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2901758c-fe3c-4126-b33d-0b0937768079 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Active clothing material perception using tactile sensing and deep learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c16cee2-1899-4b26-ad14-2bd330bd1bca · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beac2993-c198-4208-8264-2b4af7e285c2 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab5930d-261a-4dc9-a0c9-682cbba1406e · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae v2: Scaling video masked autoencoders with dual masking
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cdc1cd11-5c75-4fe2-b7cf-614719443e24 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Sigma: Sinkhorn-guided masked video modeling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f57f20a-4a5a-4dbb-8f77-630bc069549c · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Mgmae: Motion guided masking for video masked autoencoding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8890e7dc-b217-4dbd-9e61-553736bbc3c4 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9cbbb563-ba2a-4b99-a2b9-4a5c98629509 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomac: Video masked autoencoders meet convnets
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6118caa9-3120-4001-9b1f-e5fe9de1fba9 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9995e569-a2b3-4ea4-9e92-b43426a67bcd · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Vipergpt: Visual inference via python execution for reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3013d9a-5e2f-4baa-895f-36b61cf644e3 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97709027-603c-493b-a6dc-e93839b22fad · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Lora: Low-rank adaptation of large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 352771ac-e717-4add-8dd5-0cdf78a2e61e · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97ba6137-fecf-4e2f-b213-67f56a72c03b · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078e2ba6-a78d-4edd-a9b7-f6b6dd95d72f · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Improved baselines with visual instruction tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76a937d-66db-41a6-9a14-10efc16017b2 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d471dd7-0154-42c7-9e64-3e2961686ada · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b73640-6507-408d-a905-1fce1cd35252 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fa4dfe-7599-43b1-b803-b66e80ffddb1 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f3347f-591d-4657-a834-23fbaa070590 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4426ec64-3622-445e-ae22-47279bff05cd · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931b65ea-f382-4776-81ba-6b669b203d5d · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Cubic spline interpolation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3d1544c-7f2d-4937-94c2-cc7062024bc7 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction An image is worth 16x16 words: Transformers for image recognition at scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efad683c-0329-41d4-8300-4b83db121f71 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian Error Linear Units (GELUs)
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0cadcd-cdd2-481d-b2d0-2be2fd96087a · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Masked autoencoders are scalable vision learners
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c45fc7-ed47-49be-9795-9d6c2fc851f6 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian mixture models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e15719c8-7422-4632-91e9-557f79f72825 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Raft: Recurrent all-pairs field transforms for optical flow
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a60081-4ffb-4412-b2f8-1304ccf66c42 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Forward and backward warping for optical flow-based frame interpolation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6bd2af6-b91d-445a-b71f-16d16cb38fe9 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 362bd598-074d-4d98-8f7c-4449b1589cb9 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Cross-entropy loss functions: Theoretical analysis and applications
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b16cc7da-a388-4856-8a25-5582d5a000fd · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction GPT-4o System Card
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a304cf9d-0f65-4eef-af4e-93345d8f906c · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction gemini-2.5-pro-preview-05-06
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3a13097-92b7-47d9-b443-a76e1e1f8917 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-OneVision: Easy Visual Task Transfer
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120b31ac-4054-4861-882b-b452c3ddd294 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51cfa79-1d7a-46af-bebb-b92b42132eb1 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee21c921-e493-4641-8370-9e04b885dc58 · outbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5-VL Technical Report
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 742e49d2-d20c-4ff6-bfd8-ff1b8067ab7b · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38e50fb9-092a-48b4-b69d-ba9e2c9a760a · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0fdca6f4-85a3-4da9-adb4-c103fec5f06d · inbound
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Universal Visuo-Tactile Video Understanding for Embodied Interaction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34d78103-deb5-4e72-84f6-8f3bac3128f1 · inbound
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation Universal Visuo-Tactile Video Understanding for Embodied Interaction
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48f125a2-bfc2-4656-84fb-a1c77d22f31a · inbound
TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios Universal Visuo-Tactile Video Understanding for Embodied Interaction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.