Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:44:50.965141Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2510.13808.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:44:50.965141Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 703c74ad-0bdf-40fd-a576-612ff0538a68 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Thinking with images, April 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa84e41-81e5-41b2-90b9-f9ebd3e2e573 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5347d4a3-8a61-4f5b-8d5d-eacbc4be409d · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd10a2b4-2dba-47cc-9d04-db6d0ee37960 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models xgen-mm (blip-3): A family of open large multimodal models, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92163ffc-9d70-48ac-8fc1-a11a9b1faae7 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Socratic models: Composing zero-shot multimodal rea- soning with language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed888a4f-484b-4e1e-af2c-f7f422d28372 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27b91bd-f254-44cf-a1a9-9939fb1c3a6e · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models LISA: Reasoning Segmentation via Large Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee02fc50-6892-439a-87d4-6324647577fd · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Ryoo, and Tsung- Yu Lin
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 731e11fd-50a3-4db6-a776-4718e223b13b · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Qwen2.5 technical report, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98e510d-70ab-4803-9b9d-1a529ffe0993 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models The llama 3 herd of models, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0438d23b-29dc-4097-a7ed-9215bd13511c · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Learning transferable visual models from natural language supervision, 2021
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d80083f-a11d-42f7-b433-4dca7311ed2f · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Sigmoid loss for language image pre-training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb838706-b226-4911-99cc-02e620c5c3bc · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Video instruction tuning with synthetic data, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed10b1be-5ae6-4fc7-853e-f3e962310adb · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Sharegpt4video: Improving video understanding and generation with better captions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0c5737-7e21-4090-a776-6cce7783a518 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Videogpt+: Integrating image and video encoders for enhanced video understanding, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d41108d-5aca-4786-80bb-73884693a5e2 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Cinepile: A long video question answering dataset and benchmark, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7953156d-0ae6-4d7e-8ad7-e2a80f007008 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Is space-time attention all you need for video understanding? InProceedings of the International Conference on Machine Learning (ICML), July 2021
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db839ba-1d39-44d9-8644-febbdf9291ff · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Vivit: A video vision transformer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8bd5ff-67ff-4883-95d7-b0b4a446bfc9 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2345d0e5-69ca-4c56-9360-965e88d689ed · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Aim: Adapting image models for efficient video understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d01dee1-66ea-4f8b-a6ff-7c34bced0fd1 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Overcoming the pitfalls of vision- language model finetuning for ood generalization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cce8824-06e1-4cf5-96ef-194973dfc908 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Vision- language model fine-tuning via simple parameter-efficient modification
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e8ac2bd-34db-4e79-a892-f56c42d21e74 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb3cbb1-652e-4426-a733-35feb873af7f · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Video swin transformer.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3192–3201, 2021
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cba4304-3b6a-4786-8306-ec03f06328aa · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf39505f-b5d5-462e-84ea-8653c103f11d · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bi´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e43e9f4f-e94f-4afe-88c8-bb00605e6cb2 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Fusion of domain-adapted vision and language models for medical visual question answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6c48cc-da5e-40b1-b426-1d0db4bd8bf5 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Jongwoo Park, Kanchana Ranasinghe, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ffeb72b-f20c-4262-bb38-97926e745109 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Finetuned clip models are efficient video learners
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a88768-743a-46db-b621-68d364b46111 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Conditional prompt learning for vision-language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd45331a-3824-462b-b964-8453fe39cbfa · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Prompt-aligned gradient for prompt tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1b7fb0-a1c5-402e-9230-de2f5f689cf6 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Visual-language prompt tuning with knowledge- guided context optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6a533d-0a5d-4f9f-bfa6-484eba016fea · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Clip-adapter: Better vision-language models with feature adapters
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c254138b-8778-4437-ad37-f91c40ae2657 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models On domain-adaptive post-training for multimodal large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72584479-cbe5-4b20-ac57-f90012f8e66a · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5f8f65-81cd-47d5-be42-1e57e3bab511 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Llavidal: A large language-vision model for daily activities of living
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea5c746-c53e-4855-8090-c92daa703f02 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba64c5a-d775-44d0-9055-84d17570ff1d · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Visual instruction tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2356a3cb-984d-442c-a650-2133a2114045 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Apollo: An exploration of video understanding in large multimodal models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a76d039-9da7-4f66-9b87-337588268ab5 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models LoRA: Low-rank adaptation of large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b48874e-5625-435b-8025-d4447f24fde7 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c45f6a-4bd0-4d03-9e82-a930608e22fb · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models From my view to yours: Ego-augmented learning in large vision language models for understanding exocentric daily living activities, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889de24a-01e2-47bb-8d9f-c647e51b9421 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Depth anything v2
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e3786e-38c0-4edb-b7fa-325de65507eb · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Vima: General robot manipulation with multimodal prompts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4cd366-8ca7-48f2-a15d-76c755a24c24 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520aedaa-1907-4e17-a493-5a2fa1ef0193 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16bee8ce-67c4-466a-b3cd-22101435a773 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2338c354-6ad5-456e-bb39-f6b3928ffbfd · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Next-qa: Next phase of question- answering to explaining temporal actions
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 167acb43-5eb7-4d4b-b6eb-60226a32dd53 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal large language models in video analysis
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a9ad19-db59-474a-a511-7e57ad6b1db5 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Toyota smarthome: Real-world activities of daily living
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5da6090-b0a0-4783-92be-d4f377fa44bc · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46374306-4d2d-4cea-a22a-3dc8184c17ad · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Lemma: A multi- view dataset for learning multi-agent multi-task activities
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82eab911-937f-4ef1-ad36-ce4e4b3d5f3d · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Toyota smarthome untrimmed: Real-world untrimmed videos for activity detection.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b593c0-f945-42dc-b6bd-5918f1372bcd · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Quantifying attention flow in transformers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4673158-ce69-4f31-92c7-f9e615754da7 · outbound
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models Place the carrot on the plate
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.