Pith. sign in

Paper Citation Record · LEDGER

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2510.13808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.13808 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:44:50.965141Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 703c74ad-0bdf-40fd-a576-612ff0538a68 · outbound

This paper cites Thinking with images, April 2025.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Thinking with images, April 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:44.695818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:44.695818Z digest=sha256:a2619610a16808622f5b66ce91c8afc6b5a03ecc3b4d8972178f1f1901080cfb

Observation 3fa84e41-81e5-41b2-90b9-f9ebd3e2e573 · outbound

This paper cites Qwen2.5-VL Technical Report.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:44.802882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:44.802882Z digest=sha256:e1c737189983b514760e48012b42e920ff3f2ec616e55b5f708caa0845ff7882

Observation 5347d4a3-8a61-4f5b-8d5d-eacbc4be409d · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:44.942510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:44.942510Z digest=sha256:321f176e043ff9d37febfe9b9c8ec61e24b78338d3c66408529247edf05a6c09

Observation dd10a2b4-2dba-47cc-9d04-db6d0ee37960 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models, 2025.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models xgen-mm (blip-3): A family of open large multimodal models, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.114367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.114367Z digest=sha256:5193e92a0cede184feb8fe575a891f328f6c7cfb48bbdab8b6e5afc16b65176b

Observation 92163ffc-9d70-48ac-8fc1-a11a9b1faae7 · outbound

This paper cites Socratic models: Composing zero-shot multimodal rea- soning with language.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Socratic models: Composing zero-shot multimodal rea- soning with language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.284291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.284291Z digest=sha256:cc1b8149ac194071fd210940714f60d700f070e2bc1cb336211f7e9c5e050f82

Observation ed888a4f-484b-4e1e-af2c-f7f422d28372 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, 2024.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.435292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.435292Z digest=sha256:3a52e9357324d123e6fee6ddf9b0c0d84e5cfef6208edefe38a59cc922836cb1

Observation e27b91bd-f254-44cf-a1a9-9939fb1c3a6e · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models LISA: Reasoning Segmentation via Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.580737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.580737Z digest=sha256:86432235629dc0821104ff0478b607114d4a569082ed078124fc80b1d3ea3656

Observation ee02fc50-6892-439a-87d4-6324647577fd · outbound

This paper cites Ryoo, and Tsung- Yu Lin.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Ryoo, and Tsung- Yu Lin

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.721037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.721037Z digest=sha256:e9ba70bc1f091f643c3c839f7e7d78b84dd1c1a29bbcde98e228242cdd1be242

Observation 731e11fd-50a3-4db6-a776-4718e223b13b · outbound

This paper cites Qwen2.5 technical report, 2025.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Qwen2.5 technical report, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:45.904826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:45.904826Z digest=sha256:9ed80027dac2268f2f2a1efac8df8432e3bfda6eb65069b32dac50cf1e6f54fd

Observation b98e510d-70ab-4803-9b9d-1a529ffe0993 · outbound

This paper cites The llama 3 herd of models, 2024.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models The llama 3 herd of models, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.020920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.020920Z digest=sha256:47f41277bfdc79e98f9a2f8f15fb9b1312cf1e6dbdf086f64249a03552af6cee

Observation 0438d23b-29dc-4097-a7ed-9215bd13511c · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Learning transferable visual models from natural language supervision, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.119853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.119853Z digest=sha256:7b784e01b6dc2d209b6ad9845548b0ff07dbe8266426f20b86451d9f1a67fd16

Observation 2d80083f-a11d-42f7-b433-4dca7311ed2f · outbound

This paper cites Sigmoid loss for language image pre-training.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Sigmoid loss for language image pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.302660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.302660Z digest=sha256:a43ca7284a6685019a0bef81838921ad26e26be17930102772aca75ba0033c58

Observation eb838706-b226-4911-99cc-02e620c5c3bc · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Video instruction tuning with synthetic data, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.421954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.421954Z digest=sha256:4376d327b3d60300787425ea1da73c835534e8166f7a46bf9f4a64b6a43a48c3

Observation ed10b1be-5ae6-4fc7-853e-f3e962310adb · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Sharegpt4video: Improving video understanding and generation with better captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.502225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.502225Z digest=sha256:aa5652d4e1a9f1d53c0fff6d71029d42700fb9e87f899eb08c34043385cd0aa7

Observation 9d0c5737-7e21-4090-a776-6cce7783a518 · outbound

This paper cites Videogpt+: Integrating image and video encoders for enhanced video understanding, 2024.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Videogpt+: Integrating image and video encoders for enhanced video understanding, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.690847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.690847Z digest=sha256:6158d1c227bc811c11c278eb5915c41d84739ec6bd74dd61b7e8efdcb55e6d41

Observation 6d41108d-5aca-4786-80bb-73884693a5e2 · outbound

This paper cites Cinepile: A long video question answering dataset and benchmark, 2024.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Cinepile: A long video question answering dataset and benchmark, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.821253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.821253Z digest=sha256:42afbe50b7b076f4188d4cd4ec92ead558ef1c85934a6a474337afd88afb24c3

Observation 7953156d-0ae6-4d7e-8ad7-e2a80f007008 · outbound

This paper cites Is space-time attention all you need for video understanding? InProceedings of the International Conference on Machine Learning (ICML), July 2021.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Is space-time attention all you need for video understanding? InProceedings of the International Conference on Machine Learning (ICML), July 2021

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:46.942298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:46.942298Z digest=sha256:521a2ef1b6a081fb180660d9b1eae68952ad71825e4a25007153edaeaae52277

Observation 8db839ba-1d39-44d9-8644-febbdf9291ff · outbound

This paper cites Vivit: A video vision transformer.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Vivit: A video vision transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.098323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.098323Z digest=sha256:443f4797c7b07ccb9e6cda3220b80162a85b291eaf7a688ddfb38a3f9fd84126

Observation bd8bd5ff-67ff-4883-95d7-b0b4a446bfc9 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.204008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.204008Z digest=sha256:b35778a33e3b29d1c9380d72112046d311f57dd400017cc0c7028481e7aae69f

Observation 2345d0e5-69ca-4c56-9360-965e88d689ed · outbound

This paper cites Aim: Adapting image models for efficient video understanding.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Aim: Adapting image models for efficient video understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.321751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.321751Z digest=sha256:4bff333e4c3a14a58f2df2eeec9804bab6d7430f696525ec5278f6e4079e5807

Observation 2d01dee1-66ea-4f8b-a6ff-7c34bced0fd1 · outbound

This paper cites Overcoming the pitfalls of vision- language model finetuning for ood generalization.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Overcoming the pitfalls of vision- language model finetuning for ood generalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.479768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.479768Z digest=sha256:693e73227ce91b49e79ca96908d38b452c64b0c7de034cde5e0c00ee04fa1b95

Observation 7cce8824-06e1-4cf5-96ef-194973dfc908 · outbound

This paper cites Vision- language model fine-tuning via simple parameter-efficient modification.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Vision- language model fine-tuning via simple parameter-efficient modification

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.594715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.594715Z digest=sha256:a541bbf8b34b0e4e292a278fb5e76d3b2c8d538f379f468cb9f12d5d0d4e2ef8

Observation 0e8ac2bd-34db-4e79-a892-f56c42d21e74 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.732506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.732506Z digest=sha256:64c3be55f1052801b1f9f425a232f8c8b42dd34c73106774460560cbe6f77976

Observation 1eb3cbb1-652e-4426-a733-35feb873af7f · outbound

This paper cites Video swin transformer.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3192–3201, 2021.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Video swin transformer.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3192–3201, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.848413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.848413Z digest=sha256:b8913b64631ffd671a4f74a77e0c2103bbe3c6e9d3b221176a14de05f67885ad

Observation 3cba4304-3b6a-4786-8306-ec03f06328aa · outbound

This paper cites an unresolved cited work.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:47.989069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:47.989069Z digest=sha256:e354fd1936e4850d87cb9c0b300b3bc1e42f931b71b97e0f1aab207ed0814759

Observation bf39505f-b5d5-462e-84ea-8653c103f11d · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bi´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bi´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.073445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.073445Z digest=sha256:8360d5a2c86d63d61fed81f16702abe4f8d5c22d9c9637efb24a8bcaeec3c747

Observation e43e9f4f-e94f-4afe-88c8-bb00605e6cb2 · outbound

This paper cites Fusion of domain-adapted vision and language models for medical visual question answering.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Fusion of domain-adapted vision and language models for medical visual question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.193133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.193133Z digest=sha256:9b5d162461087d414778387c0cdb21bde7a9be7013b5b70977e1906e069dc86d

Observation 2d6c48cc-da5e-40b1-b426-1d0db4bd8bf5 · outbound

This paper cites Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Jongwoo Park, Kanchana Ranasinghe, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Jongwoo Park, Kanchana Ranasinghe, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.274249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.274249Z digest=sha256:7218baaa945daaa8e8ca2302f45dd88e86961ca23d932d4e51978f7099631fe1

Observation 6ffeb72b-f20c-4262-bb38-97926e745109 · outbound

This paper cites Finetuned clip models are efficient video learners.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Finetuned clip models are efficient video learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.411022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.411022Z digest=sha256:3a5c247d401ef34dc2e42ec9e531ff8954bb13342b51db8584debeca8c35b830

Observation f7a88768-743a-46db-b621-68d364b46111 · outbound

This paper cites Conditional prompt learning for vision-language models.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Conditional prompt learning for vision-language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.474795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.474795Z digest=sha256:7f01cd8c15a9460c3e7ea5907fa12b5c029ce87537a15161da6d2d8d8a37664a

Observation cd45331a-3824-462b-b964-8453fe39cbfa · outbound

This paper cites Prompt-aligned gradient for prompt tuning.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Prompt-aligned gradient for prompt tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.524319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.524319Z digest=sha256:d7aa5ac9c6f5d3276667631c1fca3b67186a9823cf4187cf8dd8fcfe4ca6c363

Observation ba1b7fb0-a1c5-402e-9230-de2f5f689cf6 · outbound

This paper cites Visual-language prompt tuning with knowledge- guided context optimization.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Visual-language prompt tuning with knowledge- guided context optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.626790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.626790Z digest=sha256:935e4d5891dbd2107eeebb0c103a4c059c774ecd5e5b6dcfcf51a49e68ce3236

Observation 5d6a533d-0a5d-4f9f-bfa6-484eba016fea · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Clip-adapter: Better vision-language models with feature adapters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.704116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.704116Z digest=sha256:4babe6dc61c54fce1ed93c87cdc27cb88d45850cc9a8579a766be02328d7de52

Observation c254138b-8778-4437-ad37-f91c40ae2657 · outbound

This paper cites On domain-adaptive post-training for multimodal large language models.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models On domain-adaptive post-training for multimodal large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.805476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.805476Z digest=sha256:396d985c6e8a89f15b2da492508129e63ccf5d75411bef5d1b89c6fd78785bff

Observation 72584479-cbe5-4b20-ac57-f90012f8e66a · outbound

This paper cites an unresolved cited work.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:48.875216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:48.875216Z digest=sha256:661d653d9a7463841f722513c5c7091648186ad8859b7081b9960e9050833531

Observation 4d5f8f65-81cd-47d5-be42-1e57e3bab511 · outbound

This paper cites Llavidal: A large language-vision model for daily activities of living.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Llavidal: A large language-vision model for daily activities of living

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.032642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.032642Z digest=sha256:cc33a2115933335bb0daf8cae9c1e9f062a5f11e35323dbecdf5975621fbec5d

Observation eea5c746-c53e-4855-8090-c92daa703f02 · outbound

This paper cites Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.092729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.092729Z digest=sha256:9c00c95c5c8e599ff2a42431ab180a29203e5bebae1fefc3521a5d4f4033dbb6

Observation cba64c5a-d775-44d0-9055-84d17570ff1d · outbound

This paper cites Visual instruction tuning.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Visual instruction tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.152203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.152203Z digest=sha256:c50fdb0efc3d1c667522c515ecd33e797ad9b3b13f88fd9c2154d55bfc9f07b0

Observation 2356a3cb-984d-442c-a650-2133a2114045 · outbound

This paper cites Apollo: An exploration of video understanding in large multimodal models.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Apollo: An exploration of video understanding in large multimodal models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.258254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.258254Z digest=sha256:0ab9f87855a659603ae67a966c05bc31c7fe434cb616f7afab55fab9e2f15b90

Observation 8a76d039-9da7-4f66-9b87-337588268ab5 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models LoRA: Low-rank adaptation of large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.306602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.306602Z digest=sha256:eba461d6ee73270722737382b9bb63dbb78508d72d09221af7d338df62eac662

Observation 9b48874e-5625-435b-8025-d4447f24fde7 · outbound

This paper cites an unresolved cited work.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.418322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.418322Z digest=sha256:23c74ecf8f681932c7b01d11c0a29fcd6c61711b58fa1f0bdbceb3b65538cbcd

Observation 85c45f6a-4bd0-4d03-9e82-a930608e22fb · outbound

This paper cites From my view to yours: Ego-augmented learning in large vision language models for understanding exocentric daily living activities, 2025.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models From my view to yours: Ego-augmented learning in large vision language models for understanding exocentric daily living activities, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.535653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.535653Z digest=sha256:215c0bb5005bb49290980b0674ae1cc6ea340375d6a4a7152169e987011ab38f

Observation 889de24a-01e2-47bb-8d9f-c647e51b9421 · outbound

This paper cites Depth anything v2.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Depth anything v2

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.589103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.589103Z digest=sha256:954ec21b39e8d279c45db6238e927a64c6001ef5d3fe4f2d7895d52eafab86e7

Observation a2e3786e-38c0-4edb-b7fa-325de65507eb · outbound

This paper cites Vima: General robot manipulation with multimodal prompts.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Vima: General robot manipulation with multimodal prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.700819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.700819Z digest=sha256:73ac95e1bd0ba77ab6d58bea0357fc39b949ca740971484030d589a47d0974ef

Observation 4f4cd366-8ca7-48f2-a15d-76c755a24c24 · outbound

This paper cites an unresolved cited work.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.777926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.777926Z digest=sha256:b259bebc5c14e42fd664a557d68a85ec6a0437926e685f9f56e2c96fe6ddf30b

Observation 520aedaa-1907-4e17-a493-5a2fa1ef0193 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:49.888970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:49.888970Z digest=sha256:62cf2f2a2e1477ac020697612b90c1f0a6e4d4b870e3a63aca6c45feb7e94ec7

Observation 16bee8ce-67c4-466a-b3cd-22101435a773 · outbound

This paper cites an unresolved cited work.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.010830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.010830Z digest=sha256:62f16a18370e4eb667a25c35fb8cab3dfbb731f4530d7605ee00191913a9d1c4

Observation 2338c354-6ad5-456e-bb39-f6b3928ffbfd · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Next-qa: Next phase of question- answering to explaining temporal actions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.090316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.090316Z digest=sha256:5d526f5d8129af9be86ba2be9479063d21254084aa64127e398f30bc5f760cb8

Observation 167acb43-5eb7-4d4b-b6eb-60226a32dd53 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal large language models in video analysis.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal large language models in video analysis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.170468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.170468Z digest=sha256:12575ac12eeee5ab2b3746f07d59f92ab59be98c06f122336f48ea3db05e70e8

Observation c5a9ad19-db59-474a-a511-7e57ad6b1db5 · outbound

This paper cites Toyota smarthome: Real-world activities of daily living.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Toyota smarthome: Real-world activities of daily living

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.237677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.237677Z digest=sha256:78ea07b74bea5316e5366cc7da7ede36f39f0e574fcf7244b649e3cc3fb0ef84

Observation d5da6090-b0a0-4783-92be-d4f377fa44bc · outbound

This paper cites Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.365277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.365277Z digest=sha256:77d5fd65da08f6e8b53b5ed383325133e19196f6af973aa434901e52f761f188

Observation 46374306-4d2d-4cea-a22a-3dc8184c17ad · outbound

This paper cites Lemma: A multi- view dataset for learning multi-agent multi-task activities.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Lemma: A multi- view dataset for learning multi-agent multi-task activities

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.510994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.510994Z digest=sha256:035731a1c51a635e2cdf272e25822e7acaaef76ca4d8c6260b562396aaca7e07

Observation 82eab911-937f-4ef1-ad36-ce4e4b3d5f3d · outbound

This paper cites Toyota smarthome untrimmed: Real-world untrimmed videos for activity detection.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Toyota smarthome untrimmed: Real-world untrimmed videos for activity detection.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.665501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.665501Z digest=sha256:ecec422f269dacebb055f3f3c8f570145e00ff0ce63063d583ac1b4891463844

Observation f9b593c0-f945-42dc-b6bd-5918f1372bcd · outbound

This paper cites Quantifying attention flow in transformers.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Quantifying attention flow in transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.805308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.805308Z digest=sha256:e76fd5ff4f1be156e60d059d8d7515b82b9ef79d2d64e0039b88b4696f666885

Observation f4673158-ce69-4f31-92c7-f9e615754da7 · outbound

This paper cites Place the carrot on the plate.

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Place the carrot on the plate

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:50.965141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:50.965141Z digest=sha256:938d4030d8bed64d6de146465d9636d4928443c6099a605f755b114ac7fb9943

Pith citing papers

No inbound Pith citation observations are available.