Pith. sign in

Paper Citation Record · LEDGER

Universal Visuo-Tactile Video Understanding for Embodied Interaction

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 5 inbound Pith citation observations for arXiv:2505.22566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22566 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:26.504997Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T08:20:49.438388Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:05:40.612441Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43cfddc7-1d6d-41b0-af88-f4bf9e6f1ad5 · outbound

This paper cites A review of tactile information: Perception and action through touch.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A review of tactile information: Perception and action through touch

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.774596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:19.560774Z digest=sha256:e0c5dcab879f91195656e5dfaacb0ca22e96defcfa62451f5b48836155106d7d

Observation fbfc6331-cc9d-439c-a6d0-4d4eab314da0 · outbound

This paper cites Task and material properties interac- tively affect softness explorations along different dimensions.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Task and material properties interac- tively affect softness explorations along different dimensions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.547098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:19.672847Z digest=sha256:d689d750ac0796017c782d6e0646e916f163b024870a67382fc9da85c038c79d

Observation 8a8dadab-e645-4958-b1e7-f92336479787 · outbound

This paper cites Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Predicting perceptual haptic attributes of textured surface from tactile data based on deep cnn-lstm network

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.284124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:19.825495Z digest=sha256:781f752dc7cc042d5b0002ef565ec44a71c0b37e66f2629fd62fe94a50c8bd4c

Observation c4f5cbb6-2d5b-4f5f-8785-90ec4adb4fd0 · outbound

This paper cites Qwen Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:19.929844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:19.929844Z digest=sha256:f78bd4d81f986f88547c2edcc389dc599eda71d9a9604f4d195f8a82d10ac42c

Observation 0301aa58-7e92-46ab-8142-1e633397736b · outbound

This paper cites Qwen2.5 Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.070560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.070560Z digest=sha256:8cc195cd0c59f465a2eceb2e594a911a86dba8fa009fde91f127b795013e71ba

Observation 2d442cf6-3280-4c99-8959-021be48de50b · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction High- resolution image synthesis with latent diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.167450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.167450Z digest=sha256:783e257d810e11944fe17e4a77e96033da236992c872a0e93cf0218aaddc4057

Observation 98429ef1-db62-4cff-bb88-2344d067e4cb · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.291420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.291420Z digest=sha256:7aa23feb9b59a99edc0b88512dc941a2e0edd02de575314ba516258a1ccb7906

Observation e1106966-9d92-452e-9005-3dcd31fb4d2c · outbound

This paper cites CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis.

Universal Visuo-Tactile Video Understanding for Embodied Interaction CCIS-Diff: A Generative Model with Stable Diffusion Prior for Controlled Colonoscopy Image Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:27.119977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:20.431420Z digest=sha256:07f8d8d8e30454eeb0cdbb97f49a4cc237ae5b77ea244912b1ca833f318b196c

Observation c811a921-8343-401f-ae5f-591f7be8ffad · outbound

This paper cites When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective.

Universal Visuo-Tactile Video Understanding for Embodied Interaction When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:33.073625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:20.545737Z digest=sha256:3ff033c1d6e3ab9884553cec8876dfb2ff9b628501cbc8497796843e6135e9c6

Observation 1a6ffca7-6dd9-4dc7-9421-2c0cd24964fb · outbound

This paper cites Gelsight: High-resolution robot tactile sensors for estimating geometry and force.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gelsight: High-resolution robot tactile sensors for estimating geometry and force

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.833076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:20.676655Z digest=sha256:8d810387b68cbf900c642af63c71f4fa4fe774962cf6c2d231d6af7dd7a59a25

Observation dfea851b-8f91-4e67-a99d-ace2f346f1a7 · outbound

This paper cites Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.535134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:20.786246Z digest=sha256:765b3e5952fd763a42bbb4adb4458475342b042677de072f5658c79488a78860

Observation acebc66d-26bc-4d24-abc6-ef6bd4c66b39 · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:20.888419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:20.888419Z digest=sha256:e6ebf841c0cfcfbe13939de3555702ecac9a81a922c83cc8dc6f9e4d60544eff

Observation 76ef809c-43e9-4f30-8cf1-93e8632aa829 · outbound

This paper cites Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:32.238115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:21.005344Z digest=sha256:3bdffa64985a658798cb701aa8701c206fab14eeea3a101fa8fd56e42b72525b

Observation 624d1181-696b-404c-ab7e-f6392e065556 · outbound

This paper cites Transferable tactile transformers for representation learning across diverse sensors and tasks.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Transferable tactile transformers for representation learning across diverse sensors and tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.979642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:21.115468Z digest=sha256:8d235b9fb06f10a2b0374a0a9c257e65436c511fedbea6a796317671fbec5cba

Observation 793c2a97-ea4c-4375-88da-a04937c09976 · outbound

This paper cites Octopi: Object Property Reasoning with Large Tactile-Language Models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Octopi: Object Property Reasoning with Large Tactile-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.199940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.199940Z digest=sha256:45ebab2b88ee25e9a9ee55cd80457934d58a2890ad3e516d1ec4691d3902e8db

Observation cdda796b-c152-44fa-914f-0ee8ff653e78 · outbound

This paper cites A touch, vision, and language dataset for multimodal alignment.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A touch, vision, and language dataset for multimodal alignment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.655885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:21.303672Z digest=sha256:40fe824bbcffcf5e3f594fc808d7ca5f86d2d4e6c18e790ba49e37aed6613ac5

Observation f7fbd7ce-31e4-40ec-bb6b-37ec7f3bed1b · outbound

This paper cites Binding touch to everything: Learn- ing unified multimodal tactile representations.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Binding touch to everything: Learn- ing unified multimodal tactile representations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.437900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.437900Z digest=sha256:e36ccc85726f7dc58ccd8e312aa94b7f35247379bd34bacabc207a55e0df5705

Observation 7ea52084-156f-4997-a910-f59c5759fb43 · outbound

This paper cites Sparsh: Self-supervised touch representations for vision-based tactile sensing.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Sparsh: Self-supervised touch representations for vision-based tactile sensing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.436273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:21.550534Z digest=sha256:4f792b152f54246ec2c145da7f40591f4defe1c592ae5d673ebcc2d088373897

Observation ebbfd123-0d6e-4b15-a32c-6d0014f2f269 · outbound

This paper cites Visuo- tactile affordances for cloth manipulation with local control.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Visuo- tactile affordances for cloth manipulation with local control

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:31.215381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:21.673000Z digest=sha256:1e96b1aafc5793369c401bf502c54a14e51c06b8017afb7b0e0ee71660e367e1

Observation cbec3edb-3dd7-4112-afbc-bf4784575848 · outbound

This paper cites A Survey of Embodied Learning for Object-Centric Robotic Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction A Survey of Embodied Learning for Object-Centric Robotic Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:21.815777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:21.815777Z digest=sha256:c52f31a199f9da63cf3ec308df744eebaf42abb63c47e387be5a1e5f234b96e9

Observation 208acdd1-c51d-49ff-aee7-c0666b5a9c53 · outbound

This paper cites Touch and go: learning from human-collected vision and touch.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch and go: learning from human-collected vision and touch

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.953063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:21.941510Z digest=sha256:c9739f2def0e76cf5caaffb752f45bd915558a4dc60ccc447fff4b93277054c7

Observation 392610d5-dcb9-4574-992d-ee451a827144 · outbound

This paper cites Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.696731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:22.066880Z digest=sha256:abe21b57477711e09aff1f4d7dc3d6b7f7cbfdb6e00f2cb5928ae8aae4d7e253

Observation 0d1a6651-3972-4388-a3d7-4f43fee1f54d · outbound

This paper cites Objectfolder 2.0: A multisensory object dataset for sim2real transfer.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Objectfolder 2.0: A multisensory object dataset for sim2real transfer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.516242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:22.198067Z digest=sha256:3227a1d57ac2928828ba8d8380cbf499122a0749141bf175c08028e7ba070ea9

Observation 6dacc2bd-93f8-45eb-9a08-28d9e34f710e · outbound

This paper cites See, hear, and feel: Smart sensory fusion for robotic manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction See, hear, and feel: Smart sensory fusion for robotic manipulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.352472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:22.334747Z digest=sha256:47252a0e6d3a7dc18d1f41fa6939e817e5b3329f1fe107a3099815744d79e0c6

Observation 2901758c-fe3c-4126-b33d-0b0937768079 · outbound

This paper cites Active clothing material perception using tactile sensing and deep learning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Active clothing material perception using tactile sensing and deep learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:30.101643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:22.466478Z digest=sha256:cc0f6a1059092b6ddc0d93606d864beeb7ab3d455ee9bdac62db67fa43ef3420

Observation 8c16cee2-1899-4b26-ad14-2bd330bd1bca · outbound

This paper cites Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:22.601918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:22.601918Z digest=sha256:0d62e45cd22cfcd385cbcaa371e2f5588ed56ec4e6e589b46180bda1aee3671e

Observation beac2993-c198-4208-8264-2b4af7e285c2 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:22.755242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:22.755242Z digest=sha256:9d09c3bd2c1f6b938abcab3c224204fe88f61038a953e5c852305070dd74076e

Observation eab5930d-261a-4dc9-a0c9-682cbba1406e · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomae v2: Scaling video masked autoencoders with dual masking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.809973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:22.850972Z digest=sha256:5b6c2017f2f1956dbf945eb05cfc879ee3dacc0b6501a705ed1d87da6966a121

Observation cdc1cd11-5c75-4fe2-b7cf-614719443e24 · outbound

This paper cites Sigma: Sinkhorn-guided masked video modeling.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Sigma: Sinkhorn-guided masked video modeling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.582233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:22.991207Z digest=sha256:a7e39e0ef429fb291ed4157b4132d3579349c6c8881d9129951448988eefdb39

Observation 2f57f20a-4a5a-4dbb-8f77-630bc069549c · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Mgmae: Motion guided masking for video masked autoencoding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.355424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:23.156250Z digest=sha256:132847e252dabe5f3025c1aae622b9d9c5b5f80c3ff5b7bc9325be5547ef941a

Observation 8890e7dc-b217-4dbd-9e61-553736bbc3c4 · outbound

This paper cites VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining.

Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:09:26.914536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:23.238850Z digest=sha256:de9302626dce6edf352afa69a360f1fe1e156086b96be31248430092093bc351

Observation 9cbbb563-ba2a-4b99-a2b9-4a5c98629509 · outbound

This paper cites Videomac: Video masked autoencoders meet convnets.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Videomac: Video masked autoencoders meet convnets

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:29.125395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:23.337796Z digest=sha256:0b44c0b44186fc2eb8f986b3aa7998c311e4378d301d52f14665204dc0282c5e

Observation 6118caa9-3120-4001-9b1f-e5fe9de1fba9 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.432848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.432848Z digest=sha256:72ad1b4a6f911b233862cf5abfa481518b829d3124934a9161c565fca8bc669e

Observation 9995e569-a2b3-4ea4-9e92-b43426a67bcd · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Vipergpt: Visual inference via python execution for reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.595338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.595338Z digest=sha256:834adee637a54d6d81daf132e14e5a538738cc49c3f6b9269a89e9e08ce29352

Observation a3013d9a-5e2f-4baa-895f-36b61cf644e3 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.681350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.681350Z digest=sha256:2e3edf42bcb5c2d22833d9b8502700eb713894fd8f085d76dee02f954016987d

Observation 97709027-603c-493b-a6dc-e93839b22fad · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Lora: Low-rank adaptation of large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.786995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.786995Z digest=sha256:2aa19662782762fc92f0b0b294a3f6a3252d3e736180f2436622c37a100a2b23

Observation 352771ac-e717-4add-8dd5-0cdf78a2e61e · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.868777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:23.865887Z digest=sha256:a5ee0aeeb0f2f8e4fa2cb6bb3dfff1b08085f07a8012e9c573f6466c2acf7387

Observation 97ba6137-fecf-4e2f-b213-67f56a72c03b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:23.964185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:23.964185Z digest=sha256:e6cdf844db66605f34fe61061341d19da6d23b648d6583bafc06d6c89f606618

Observation 078e2ba6-a78d-4edd-a9b7-f6b6dd95d72f · outbound

This paper cites Improved baselines with visual instruction tuning.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Improved baselines with visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.036960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.036960Z digest=sha256:a226afe7628a373768b265c2513a70bb6792866fb0cd278b714ca523fd30513f

Observation d76a937d-66db-41a6-9a14-10efc16017b2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.134433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.134433Z digest=sha256:4756387970abc56b164c81fc091ac1e7fff0c50122074b370a940e2babc33b4a

Observation 4d471dd7-0154-42c7-9e64-3e2961686ada · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.251376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.251376Z digest=sha256:d8d25cfc1665f1d2647bfa21b4c97fc778f43cdd2d85c536f894bed8900556e7

Observation d3b73640-6507-408d-a905-1fce1cd35252 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.344367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.344367Z digest=sha256:2ed77f596e5f9b6ea3d212635626dfe709a052d3e45770e1fe7b03d7bc270f58

Observation 50fa4dfe-7599-43b1-b803-b66e80ffddb1 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.408542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.408542Z digest=sha256:b5041f08a85215a6462ca8ae9f2f915796b9f8a5637a9ae5f35d67d3bca6269c

Observation a3f3347f-591d-4657-a834-23fbaa070590 · outbound

This paper cites Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.502357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.502357Z digest=sha256:7d25873ca9854c2e0b797fdd53962043bd6ca9e5b3fc5b22c7224f624e0b4dde

Observation 4426ec64-3622-445e-ae22-47279bff05cd · outbound

This paper cites Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.596314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.596314Z digest=sha256:fe0e82176e483b285ba2559e10b91651c62f6c2fab3cfce32a753ec518913718

Observation 931b65ea-f382-4776-81ba-6b669b203d5d · outbound

This paper cites Cubic spline interpolation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Cubic spline interpolation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.617259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:24.715771Z digest=sha256:9cd1d9e4241d13fef435855aea34da182e80baaeb5c22d83cd98396b67323fe2

Observation d3d1544c-7f2d-4937-94c2-cc7062024bc7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Universal Visuo-Tactile Video Understanding for Embodied Interaction An image is worth 16x16 words: Transformers for image recognition at scale

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.313082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:24.840567Z digest=sha256:1b02df0dac95862c202d8f258a7d32d9f8b0ad5dec00e604deb03ad3e954482e

Observation efad683c-0329-41d4-8300-4b83db121f71 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian Error Linear Units (GELUs)

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.983127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.983127Z digest=sha256:8561634c9ef83529ee4017d93265da013b9d08034d4cc39731340b8576833d97

Observation 6b0cadcd-cdd2-481d-b2d0-2be2fd96087a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Masked autoencoders are scalable vision learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.062498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.062498Z digest=sha256:e48a47a9f17da6063c5c3098366b8608569a5db79c28f4daa994727ca01f2076

Observation d9c45fc7-ed47-49be-9795-9d6c2fc851f6 · outbound

This paper cites Gaussian mixture models.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Gaussian mixture models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:28.079862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:25.153871Z digest=sha256:bc1c886253d14c767cba89798c9bd437752073129b757f6717f1acdd1446f2b3

Observation e15719c8-7422-4632-91e9-557f79f72825 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Raft: Recurrent all-pairs field transforms for optical flow

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.284624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.284624Z digest=sha256:c4316aed6fa02f5b96a6053bb70fdbf8479941eafd823faa1bee67571fbf065a

Observation e6a60081-4ffb-4412-b2f8-1304ccf66c42 · outbound

This paper cites Forward and backward warping for optical flow-based frame interpolation.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Forward and backward warping for optical flow-based frame interpolation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.860591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:25.397975Z digest=sha256:0e789d80143127eb6f28872540831687762fb023b46ecc5ff60e04a95b244891

Observation d6bd2af6-b91d-445a-b71f-16d16cb38fe9 · outbound

This paper cites Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Extrapolation-based video retargeting with backward warping using an image-to-warping vector generation network

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.612526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:25.529812Z digest=sha256:7e4516acfd547e24a8cf8848e02dd50533e1041dd5dbdfd6b0c05e3b2f8a5019

Observation 362bd598-074d-4d98-8f7c-4449b1589cb9 · outbound

This paper cites Cross-entropy loss functions: Theoretical analysis and applications.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Cross-entropy loss functions: Theoretical analysis and applications

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.692625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.692625Z digest=sha256:92d9b0ae42aa9d008f797c82eb0235923a952c39f63ecfc7cbc0e5f07dbcba70

Observation b16cc7da-a388-4856-8a25-5582d5a000fd · outbound

This paper cites GPT-4o System Card.

Universal Visuo-Tactile Video Understanding for Embodied Interaction GPT-4o System Card

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:25.820834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:25.820834Z digest=sha256:695237edffc478c8119b3b0d4fe64700e0252d261a43ac2febed30451996d691

Observation a304cf9d-0f65-4eef-af4e-93345d8f906c · outbound

This paper cites gemini-2.5-pro-preview-05-06.

Universal Visuo-Tactile Video Understanding for Embodied Interaction gemini-2.5-pro-preview-05-06

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:09:27.360930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:09:25.923745Z digest=sha256:5e19f636a8b07528cab4e906ef64124e40b94a6cdc84272104bfd7fcc77a04c5

Observation f3a13097-92b7-47d9-b443-a76e1e1f8917 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-OneVision: Easy Visual Task Transfer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.070117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.070117Z digest=sha256:d63a2eadc2a1aa8c07d9bbec2b9602659ac8b02a296612fb022808af563d3100

Observation 120b31ac-4054-4861-882b-b452c3ddd294 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.216112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.216112Z digest=sha256:5516b9b464f6b1789203657058b025113f206a4a4e0d69611c0f17ae1f9b8494

Observation c51cfa79-1d7a-46af-bebb-b92b42132eb1 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.367593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.367593Z digest=sha256:d1070ea02919c8f92f37ee6bad16a82ed7668429f7f21c65daa11324e31c4613

Observation ee21c921-e493-4641-8370-9e04b885dc58 · outbound

This paper cites Qwen2.5-VL Technical Report.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Qwen2.5-VL Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.504997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.504997Z digest=sha256:25fed17c9ed15f08f32a93aeff83d7878c07706b10c52ee4eacddfca8ae52b8c

Pith citing papers

Observation 742e49d2-d20c-4ff6-bfd8-ff1b8067ab7b · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:11.718986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:93438b4b14bdfbdca4bf107abcc786be193a3797286410e194fe05ad7c907090

Observation 38e50fb9-092a-48b4-b69d-ba9e2c9a760a · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.117773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:4a9ffaf77e9c6d0149bfc446e742374695215c16da474099e14adcd30b215ec1

Observation 0fdca6f4-85a3-4da9-adb4-c103fec5f06d · inbound

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms cites this paper.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.369400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:14d52de4d4b3dd6ffea5fba39e3463c14f5c6c4ddd155096eda8b7e04e4fb2f2

Observation 34d78103-deb5-4e72-84f6-8f3bac3128f1 · inbound

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation cites this paper.

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.614055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:56:03.597839Z digest=sha256:50e03de870fcd8b4563618b6629cffbce67cc1fd23c157f7f1c972a531e8efca

Observation 48f125a2-bfc2-4656-84fb-a1c77d22f31a · inbound

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios cites this paper.

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios Universal Visuo-Tactile Video Understanding for Embodied Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:20:49.438388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:20:49.438388Z digest=sha256:2971218d3bb023c2c8e7835fd1a4865e1344aa3fe210c02260e7057fa34219f3