Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:13.889873Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.24346.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:13.889873Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2ac25e71-3ae8-449b-9bd8-4267519269bd · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Mm-vit: Multi-modal video transformer for compressed video action recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 408e9a2d-3b76-43e1-8d19-c0db8a1f8b40 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3ca7c93-d783-4fc9-8c8f-d91f16364647 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f4a7285-0452-42b3-a448-a88371dc3454 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Swinbert: End-to-end transformers with sparse attention for video captioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d18ed918-683d-4e62-9802-8732f95e455e · outbound
VUDG: A Dataset for Video Understanding Domain Generalization End-to-end generative pretraining for multimodal video captioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dd63dcc-8a37-4eba-92da-5b710be2637a · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Text with knowledge graph augmented transformer for video captioning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c44d19c-02c3-41b2-b73a-b3873a620ef5 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da4843da-867d-484e-a70b-f61e44a77040 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Invariant grounding for video question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a94c8fd1-48a4-4d3b-b56e-525276fc14cb · outbound
VUDG: A Dataset for Video Understanding Domain Generalization From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39c95ccd-7480-4c1c-947a-a64ad30d46af · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Morevqa: Exploring modular reasoning models for video question answering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719fcd8c-da12-49b7-952c-787b07e52646 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6ae223e-649a-476b-b129-640cd0b54210 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aa4dada-0a55-4abf-a65a-93b597e49f3f · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Meta-causal learning for single domain generalization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5de17fab-d8c3-44db-9614-aade6df7d16b · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3642f322-b4d7-420c-99d3-6c0c4b4e89ee · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408e7f61-f69b-46cd-97ec-da0311486d57 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d7565d3-43ba-40ea-bf00-649e69218ad5 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04193e7e-8181-459f-9af9-e29e13903829 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfef86a0-5c3d-43d5-acfa-ccba42c555b1 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f001e191-3bef-465f-873e-5f0544e86273 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Internvid: A large-scale video-text dataset for multimodal understanding and generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca3c054d-30ba-46a5-a16e-d55863fa1781 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8fcc48-396c-45eb-8395-952d6e62d9e5 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Qwen2.5-VL Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e1aad9-c1a1-46fb-9736-104483cb5b61 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Activitynet: A large-scale video benchmark for human activity understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f30c54c-a707-413a-acfd-957ef53fdc5a · outbound
VUDG: A Dataset for Video Understanding Domain Generalization The Kinetics Human Action Video Dataset
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e75b38-7b5e-4f43-91bf-3494471c519f · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Hollywood in homes: Crowdsourcing data collection for activity understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3daff457-2f87-4fc7-b2e4-9a7c19385452 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization TVQA: Localized, Compositional Video Question Answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abd516f-781c-4d5d-bd2e-5341b2b2387d · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video question answering via gradually refined attention over appearance and motion
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe10182d-2607-4e6b-a8cc-aa468c18279e · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Msr-vtt: A large video description dataset for bridging video and language
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a479c01a-fbd6-4be4-90ed-34515963f71f · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afd06ccc-503f-4e77-90a8-fad64cd4d8f5 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-audio domain generalization via confounder disentanglement
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8ae2171-134a-421e-a12a-96bcb2751ee4 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83944137-ba67-43fa-80a5-680fe3cc3151 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88bec4d3-31fa-4d79-bf07-42b717047a3a · outbound
VUDG: A Dataset for Video Understanding Domain Generalization What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3173452b-aaee-44b7-8dec-102394f24a74 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Ego4d: Around the world in 3,000 hours of egocentric video
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50943904-c01d-455a-9c31-7a1f84084fe4 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea7fe981-ba3d-46fc-b507-43ec8e51e7b8 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1acd7fa0-0319-4d83-a723-8ff7b95ac313 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec74c2f-daa6-4af3-993d-154d4e839773 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57437514-aca4-4203-9fdb-6bd9ed8012ea · outbound
VUDG: A Dataset for Video Understanding Domain Generalization TempCompass: Do Video LLMs Really Understand Videos?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e93b29a3-6889-478a-9775-24288b4fe2fe · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e01d63-607a-4dd7-86f7-e68e3e54acb9 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Poem: polarization of embeddings for domain-invariant representations
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de7d5e04-b281-46b1-8e1c-6643da3feda6 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3122d34-d08f-4471-b2fb-b56761562758 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Clifton, and Jie Chen
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88d004b5-88ee-426f-90fa-505684e4b023 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b324e589-68a1-4714-8d08-3deb4559a04d · outbound
VUDG: A Dataset for Video Understanding Domain Generalization MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ddc9ed9-fd2d-4530-ba5f-289c8eae72c8 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization VideoChat: Chat-Centric Video Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52315005-d983-45bb-afc4-a0374560c0ed · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f53ff2bb-efdb-4858-b175-7b01a95b1168 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56e612ab-1b36-4961-b76d-e4097da982a5 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5057a45-94c8-41c0-9308-1e4fcef91d8a · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd6c106-ca4c-4be1-9780-1baf1543045a · outbound
VUDG: A Dataset for Video Understanding Domain Generalization Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ee116f-83bc-4a1e-8d60-984fb2908a3f · outbound
VUDG: A Dataset for Video Understanding Domain Generalization An image is worth 16x16 words: Transformers for image recognition at scale
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f291b6-cb27-4868-aac0-081effaf01d9 · outbound
VUDG: A Dataset for Video Understanding Domain Generalization C- pack: Packed resources for general chinese embeddings
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.