Pith. sign in

Paper Citation Record · LEDGER

Universal Visuo-Tactile Video Understanding for Embodied Interaction

As of 17 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 6 inbound Pith citation observations for arXiv:2505.22566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22566 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:26.504997Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:47:20.035585Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:05:40.612441Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43cfddc7-1d6d-41b0-af88-f4bf9e6f1ad5 · outbound

This paper cites A review of tactile information: Perception and action through touch.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A review of tactile information: Perception and action through touch

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.774596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:19.560774Z digest=sha256:40376a8ff6ea38daf06f8d657999767ac907cc99d3c287f1b6c10d184b0a49d2

Observation fbfc6331-cc9d-439c-a6d0-4d4eab314da0 · outbound

This paper cites Task and material properties interac- tively affect softness explorations along different dimensions.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Task and material properties interac- tively affect softness explorations along different dimensions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.547098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:19.672847Z digest=sha256:1c8aefc3d22d91b7934025ba9906a7cf769c9cdcf53d3fbc901dcf2dd4ce243c

Observation 8a8dadab-e645-4958-b1e7-f92336479787 · outbound

This paper cites Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.284124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:19.825495Z digest=sha256:c4599381ef89b1af8551ce90c6fb36bf6d859b092f6fa205e0e122f5acc4f03a

Observation c4f5cbb6-2d5b-4f5f-8785-90ec4adb4fd0 · outbound

This paper cites Qwen Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:19.929844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:19.929844Z digest=sha256:58a308552e77a945aa15aeafe9f36975955a22642ad3c463754834276118dae2

Observation 0301aa58-7e92-46ab-8142-1e633397736b · outbound

This paper cites Qwen2.5 Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.070560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.070560Z digest=sha256:5303d91ed00b2313c5e1964b65729328e51c37f1e83462e18d942ba0298ac1ba

Observation 2d442cf6-3280-4c99-8959-021be48de50b · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction High- resolution image synthesis with latent diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.167450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.167450Z digest=sha256:b8e00ab88a47a93477c0ee305b99f0ff7f00c9f669aafd2840851261dd536d66

Observation 98429ef1-db62-4cff-bb88-2344d067e4cb · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.291420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.291420Z digest=sha256:c0e41115d4fc2f7a851b5df38b8179f9fa3df3deb5fc7c1e5ab25896564399fc

Observation e1106966-9d92-452e-9005-3dcd31fb4d2c · outbound

This paper cites CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis.

Universal Visuo-Tactile Video Understanding for Embodied Interaction CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:27.119977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:20.431420Z digest=sha256:d181ece7ff1962978cc1b907e2b5831aaad7d6978ef2d8860f2a533c553d8eb9

Observation c811a921-8343-401f-ae5f-591f7be8ffad · outbound

This paper cites When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective.

Universal Visuo-Tactile Video Understanding for Embodied Interaction When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.073625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:20.545737Z digest=sha256:c8f846d2a03d0a801d3572adba9301367ba06f22e8794613772c6aee6091998e

Observation 1a6ffca7-6dd9-4dc7-9421-2c0cd24964fb · outbound

This paper cites Gelsight: High-resolution robot tactile sensors for estimating geometry and force.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gelsight: High-resolution robot tactile sensors for estimating geometry and force

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.833076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:20.676655Z digest=sha256:0caa24e7e5f54d9c78d1e4680332013e7ab39b38bb9a3851013bcb5f306d9487

Observation dfea851b-8f91-4e67-a99d-ace2f346f1a7 · outbound

This paper cites Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.535134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:20.786246Z digest=sha256:6bbf97aa19eb4deb5edde037af23525a0890d629c192abb1dce0a5aab1cec092

Observation acebc66d-26bc-4d24-abc6-ef6bd4c66b39 · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.888419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.888419Z digest=sha256:7a53aefde796ffabc2cff2826e70c966d0de6000c3581f5a777c692566d2a99e

Observation 76ef809c-43e9-4f30-8cf1-93e8632aa829 · outbound

This paper cites Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.238115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:21.005344Z digest=sha256:374cc9d57aae2014622674b24235309088c427c6b5114f393c749106f97faaa2

Observation 624d1181-696b-404c-ab7e-f6392e065556 · outbound

This paper cites Transferable tactile transformers for representation learning across diverse sensors and tasks.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Transferable tactile transformers for representation learning across diverse sensors and tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.979642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:21.115468Z digest=sha256:c69080b2dfafe854cde7c0e8da3932c3dd14a6d574e6dbca3d8020d2b27a398e

Observation 793c2a97-ea4c-4375-88da-a04937c09976 · outbound

This paper cites Octopi: Object Property Reasoning with Large Tactile-Language Models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Octopi: Object Property Reasoning with Large Tactile-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.199940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.199940Z digest=sha256:6d8013baf9af6064298558e520d7a769bb77d8789e0248cda0661d966b1b8b01

Observation cdda796b-c152-44fa-914f-0ee8ff653e78 · outbound

This paper cites A touch, vision, and language dataset for multimodal alignment.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A touch, vision, and language dataset for multimodal alignment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.655885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:21.303672Z digest=sha256:03eaf4e13c028b852ad7e2ff8dafe63a991b6fbd135f1c760374592256feee1f

Observation f7fbd7ce-31e4-40ec-bb6b-37ec7f3bed1b · outbound

This paper cites Binding touch to everything: Learn- ing unified multimodal tactile representations.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Binding touch to everything: Learn- ing unified multimodal tactile representations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.437900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.437900Z digest=sha256:6529057dfa5bfe1c3098ab8bed41c3c7fd7761988700ce0c4a452acc19cc226e

Observation 7ea52084-156f-4997-a910-f59c5759fb43 · outbound

This paper cites Sparsh: Self-supervised touch representations for vision-based tactile sensing.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Sparsh: Self-supervised touch representations for vision-based tactile sensing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.436273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:21.550534Z digest=sha256:f0efa813efe09746733011d7c0b47ac59e9c624c049386889c45d8646bab2983

Observation ebbfd123-0d6e-4b15-a32c-6d0014f2f269 · outbound

This paper cites Visuo- tactile affordances for cloth manipulation with local control.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Visuo- tactile affordances for cloth manipulation with local control

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.215381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:21.673000Z digest=sha256:e47ec94f31dd891424465734024148754978fbf5ac7c3b054135c90a632af3f9

Observation cbec3edb-3dd7-4112-afbc-bf4784575848 · outbound

This paper cites A Survey of Embodied Learning for Object-Centric Robotic Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A Survey of Embodied Learning for Object-Centric Robotic Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.815777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.815777Z digest=sha256:4c4dbf0a2545b76eb38b206707b851a6acfddec2c357281e424ebf75bada1683

Observation 208acdd1-c51d-49ff-aee7-c0666b5a9c53 · outbound

This paper cites Touch and go: learning from human-collected vision and touch.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch and go: learning from human-collected vision and touch

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.953063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:21.941510Z digest=sha256:f218b3cd8e962f3f988876692218bac3905734cc45fdddf6f1cb5d280377e272

Observation 392610d5-dcb9-4574-992d-ee451a827144 · outbound

This paper cites Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.696731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:22.066880Z digest=sha256:050d0a86f0ea41a2979281fdee7feea359bcb8bc6809b35a61c825244997c605

Observation 0d1a6651-3972-4388-a3d7-4f43fee1f54d · outbound

This paper cites Objectfolder 2.0: A multisensory object dataset for sim2real transfer.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder 2.0: A multisensory object dataset for sim2real transfer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.516242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:22.198067Z digest=sha256:86b6caa938c18cc737482e23fae4c127781ab02731a7cbfc26febf6e486104ee

Observation 6dacc2bd-93f8-45eb-9a08-28d9e34f710e · outbound

This paper cites See, hear, and feel: Smart sensory fusion for robotic manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction See, hear, and feel: Smart sensory fusion for robotic manipulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.352472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:22.334747Z digest=sha256:0fb96d77195b8e09cb8f9ffca2d45854d0c714250a73577f00859925f422a750

Observation 2901758c-fe3c-4126-b33d-0b0937768079 · outbound

This paper cites Active clothing material perception using tactile sensing and deep learning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Active clothing material perception using tactile sensing and deep learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.101643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:22.466478Z digest=sha256:b483f876c6bb810405c0e686efc8ad95b070c6927339ec396c8000417bdfcd11

Observation 8c16cee2-1899-4b26-ad14-2bd330bd1bca · outbound

This paper cites Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:22.601918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:22.601918Z digest=sha256:28ca78d90d5f4ad2e265ddf7c0b8d464a9517fec8a7c00b26f12781c546aff65

Observation beac2993-c198-4208-8264-2b4af7e285c2 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:22.755242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:22.755242Z digest=sha256:cdb92f8beb0f8bb231e9d50e1771df5803bd703b8f1771d950df5758034613b1

Observation eab5930d-261a-4dc9-a0c9-682cbba1406e · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae v2: Scaling video masked autoencoders with dual masking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.809973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:22.850972Z digest=sha256:fff246bc9903dad4686733dad92cc0c557d0253e8d0f22a00c1b2fb8c99d7c36

Observation cdc1cd11-5c75-4fe2-b7cf-614719443e24 · outbound

This paper cites Sigma: Sinkhorn-guided masked video modeling.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Sigma: Sinkhorn-guided masked video modeling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.582233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:22.991207Z digest=sha256:7a11b152d487ff71b4876e66456cb6c126f6b979c7c4d03d3075ef88a31088f3

Observation 2f57f20a-4a5a-4dbb-8f77-630bc069549c · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Mgmae: Motion guided masking for video masked autoencoding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.355424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:23.156250Z digest=sha256:9d83b9daca1ec39979e3aae59327a4750035ca55c8cc82a0adcf291c56c08fb8

Observation 8890e7dc-b217-4dbd-9e61-553736bbc3c4 · outbound

This paper cites VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining.

Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:26.914536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:23.238850Z digest=sha256:113b3343bb4a77fc109e5fe4644012d1e0883930769340dff556862e6e07caba

Observation 9cbbb563-ba2a-4b99-a2b9-4a5c98629509 · outbound

This paper cites Videomac: Video masked autoencoders meet convnets.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomac: Video masked autoencoders meet convnets

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.125395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:23.337796Z digest=sha256:07c3531961f273b483d6f1f1f8d4860b86d5cf7d1890e5d27925546dd89b8706

Observation 6118caa9-3120-4001-9b1f-e5fe9de1fba9 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.432848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.432848Z digest=sha256:f85b1a5dca5a0b61fa0651ba7c14c1aec5fdeb73bbe5b55a0ac8a7fe6dc76ff3

Observation 9995e569-a2b3-4ea4-9e92-b43426a67bcd · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Vipergpt: Visual inference via python execution for reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.595338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.595338Z digest=sha256:67065e87a1d14e4e03889140b8d7bf26391bf2d1f2428cd572b5ecdfa7f0ddf6

Observation a3013d9a-5e2f-4baa-895f-36b61cf644e3 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.681350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.681350Z digest=sha256:aa2514fe07d670c009a27080a08306a84c7e6b83724800ce5575dd7f205cc177

Observation 97709027-603c-493b-a6dc-e93839b22fad · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Lora: Low-rank adaptation of large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.786995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.786995Z digest=sha256:02823cb8b4084102f746d8bb968d0af33e6835b42b31832de90d6bba56b65ee1

Observation 352771ac-e717-4add-8dd5-0cdf78a2e61e · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.868777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:23.865887Z digest=sha256:76acd8169593d793ef6d59008e29acd146dca945f5724326d55c32516274687d

Observation 97ba6137-fecf-4e2f-b213-67f56a72c03b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.964185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.964185Z digest=sha256:61f388d9db796cc3365bc059fef64d92349b2aa1fdb2aa2153a843b9ecaa59ef

Observation 078e2ba6-a78d-4edd-a9b7-f6b6dd95d72f · outbound

This paper cites Improved baselines with visual instruction tuning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Improved baselines with visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.036960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.036960Z digest=sha256:fd8560a1908f84722614ae57ea0dabbc3bc21ed757d4eee53242b04c6773fa2b

Observation d76a937d-66db-41a6-9a14-10efc16017b2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.134433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.134433Z digest=sha256:f798f53b778bd178ce343483a76982bbed4eb1c312c5e31fe2fe9a7abdcfffa8

Observation 4d471dd7-0154-42c7-9e64-3e2961686ada · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.251376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.251376Z digest=sha256:cd9294aeeb51478735846e0c56b2b149cdfdda1887d14b8528d2c2b821bbd66c

Observation d3b73640-6507-408d-a905-1fce1cd35252 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.344367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.344367Z digest=sha256:2f4098f7258d1008b0457cd36ed9e30bef595044fae536c06db069a8962d35d8

Observation 50fa4dfe-7599-43b1-b803-b66e80ffddb1 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.408542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.408542Z digest=sha256:8ba8c817cdcae6ed4edf149a60528c414ffe6e443b8925b7d5ce41331c9217cc

Observation a3f3347f-591d-4657-a834-23fbaa070590 · outbound

This paper cites Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.502357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.502357Z digest=sha256:c817651b131ce97dd429eae9c5b859f071c7fff6e5a542be18667b18713a9213

Observation 4426ec64-3622-445e-ae22-47279bff05cd · outbound

This paper cites Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.596314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.596314Z digest=sha256:e76cd4f75f0e767f57f5b1a83404e19dcca1a2db22b57b8bcd95ab77e49a27fc

Observation 931b65ea-f382-4776-81ba-6b669b203d5d · outbound

This paper cites Cubic spline interpolation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Cubic spline interpolation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.617259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:24.715771Z digest=sha256:b669fe1a3279b68791b26161324fd044066703c369ff2f3f10922062c7cf39b9

Observation d3d1544c-7f2d-4937-94c2-cc7062024bc7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Universal Visuo-Tactile Video Understanding for Embodied Interaction An image is worth 16x16 words: Transformers for image recognition at scale

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.313082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:24.840567Z digest=sha256:0c11557bf3823f0395e6e3288090bd7e34a5992e32ce8ac75f728d733e9aa474

Observation efad683c-0329-41d4-8300-4b83db121f71 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian Error Linear Units (GELUs)

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.983127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.983127Z digest=sha256:84c2b76e751d29076e404df7c651bcd18b81267dfdf6b7705cc79161560cd7b8

Observation 6b0cadcd-cdd2-481d-b2d0-2be2fd96087a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Masked autoencoders are scalable vision learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.062498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.062498Z digest=sha256:82d864cf409e39113c36ff7dbc7f00298111d59392126a27de87f2e60f5f464e

Observation d9c45fc7-ed47-49be-9795-9d6c2fc851f6 · outbound

This paper cites Gaussian mixture models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian mixture models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.079862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:25.153871Z digest=sha256:923b1e405ee353a3b9ef16ce50f105e9e62bca5434eb121080151f58657fcb15

Observation e15719c8-7422-4632-91e9-557f79f72825 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Raft: Recurrent all-pairs field transforms for optical flow

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.284624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.284624Z digest=sha256:44f0c2a881efa5ed2ba3c116402b0e37f15b0f9774472a2c81d332794709d2c1

Observation e6a60081-4ffb-4412-b2f8-1304ccf66c42 · outbound

This paper cites Forward and backward warping for optical flow-based frame interpolation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Forward and backward warping for optical flow-based frame interpolation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.860591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:25.397975Z digest=sha256:3271877ed0873540076de7ec9cbaf959dc16546ebafeb89175b77f545df47036

Observation d6bd2af6-b91d-445a-b71f-16d16cb38fe9 · outbound

This paper cites Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.612526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:25.529812Z digest=sha256:d29889dded04528935d8ee2d9b3f40e8db4240a915fe2844874824cb6005f56e

Observation 362bd598-074d-4d98-8f7c-4449b1589cb9 · outbound

This paper cites Cross-entropy loss functions: Theoretical analysis and applications.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Cross-entropy loss functions: Theoretical analysis and applications

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.692625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.692625Z digest=sha256:2965165d5a33ac45bcaf539e29db92c2b74ed4532de7580eb9c8f3c07cab8999

Observation b16cc7da-a388-4856-8a25-5582d5a000fd · outbound

This paper cites GPT-4o System Card.

Universal Visuo-Tactile Video Understanding for Embodied Interaction GPT-4o System Card

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.820834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.820834Z digest=sha256:07d718ff6839619f0e9c2105e5b55eae255d4fec774095d886742ffb38fbbec3

Observation a304cf9d-0f65-4eef-af4e-93345d8f906c · outbound

This paper cites gemini-2.5-pro-preview-05-06.

Universal Visuo-Tactile Video Understanding for Embodied Interaction gemini-2.5-pro-preview-05-06

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.360930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T13:09:25.923745Z digest=sha256:d824ba1374479e1332e32067ba68084af1cb836e78061903c12b1e3c0d097b26

Observation f3a13097-92b7-47d9-b443-a76e1e1f8917 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-OneVision: Easy Visual Task Transfer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.070117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.070117Z digest=sha256:0fbf48e5bdc74bfa16e9a7c4ca0c215f0f06bf3d142a927e6e7aa36d48a0ccf3

Observation 120b31ac-4054-4861-882b-b452c3ddd294 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.216112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.216112Z digest=sha256:29e2749001f6da7ac86970ecab1a7c884df76bc66c2fcf3a1399fdc411eae8cf

Observation c51cfa79-1d7a-46af-bebb-b92b42132eb1 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.367593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.367593Z digest=sha256:6760eb268927a04a8cc7c6760257d9d0ea60f3467a1dfc4b431a595069bbb123

Observation ee21c921-e493-4641-8370-9e04b885dc58 · outbound

This paper cites Qwen2.5-VL Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5-VL Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.504997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.504997Z digest=sha256:f44fe4a9f36fdaae90dccd20456e3c133a2429fcd031b4e7a0d628573ca10b77

Pith citing papers

Observation abc599b5-bf7f-4fd9-bc19-3353ec6d9c16 · inbound

SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes cites this paper.

SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:47:20.035585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:47:20.035585Z digest=sha256:df19e370b439716324ff3b75fed421fd9459e35906547f52abbfe23bf1855bbb

Observation 742e49d2-d20c-4ff6-bfd8-ff1b8067ab7b · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:11.718986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:2da1c62bf90ddc6325641709ef960112a8369e80d4affb03ef52e1d925f4371e

Observation 38e50fb9-092a-48b4-b69d-ba9e2c9a760a · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.117773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:cbc0a68ee830214ad590b4817c0b37f95e89e5bd7ad28fe5ba8711db016170f8

Observation 0fdca6f4-85a3-4da9-adb4-c103fec5f06d · inbound

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms cites this paper.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.369400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:6f61ae6eed6fd4ef90dbcaa47484b001528339018ec9b176f80cc47beec8c0c0

Observation 34d78103-deb5-4e72-84f6-8f3bac3128f1 · inbound

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation cites this paper.

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.614055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T05:56:03.597839Z digest=sha256:b3d602dd942362f4e550d8386a70c8effd6711d051409cd3fe87fd817af27074

Observation 48f125a2-bfc2-4656-84fb-a1c77d22f31a · inbound

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios cites this paper.

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:20:49.438388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:20:49.438388Z digest=sha256:e0e0ddc52578150dba1d25934156829912c4cc5ad2ffc3419f97a49619b2dcaa