Pith. sign in

Paper Citation Record · LEDGER

VUDG: A Dataset for Video Understanding Domain Generalization

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.24346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24346 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:13.889873Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ac25e71-3ae8-449b-9bd8-4267519269bd · outbound

This paper cites Mm-vit: Multi-modal video transformer for compressed video action recognition.

VUDG: A Dataset for Video Understanding Domain Generalization Mm-vit: Multi-modal video transformer for compressed video action recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.898598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.169622Z digest=sha256:ed3655e579aa3dc66d4a370e44d35bcd5ac3f838c3ed96c82569f9663ad74976

Observation 408e9a2d-3b76-43e1-8d19-c0db8a1f8b40 · outbound

This paper cites Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.765665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.271553Z digest=sha256:ba4b6547bde1a85d0f1fe67d41fa6a3389a4f2b97da688b6d1b32a1875a9da2f

Observation d3ca7c93-d783-4fc9-8c8f-d91f16364647 · outbound

This paper cites Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025.

VUDG: A Dataset for Video Understanding Domain Generalization Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.342343Z digest=sha256:5b8bfdfa035a1695e0731500e3996b2eb823711b757e69f4d9f75a433a8aa739

Observation 0f4a7285-0452-42b3-a448-a88371dc3454 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization Swinbert: End-to-end transformers with sparse attention for video captioning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.266711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.421238Z digest=sha256:3f435ecd6e0d009814a960a25a2993255629bfd8ef5a51e653f5ab85395faf90

Observation d18ed918-683d-4e62-9802-8732f95e455e · outbound

This paper cites End-to-end generative pretraining for multimodal video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization End-to-end generative pretraining for multimodal video captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.135728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.512771Z digest=sha256:2321b1d3b4041f0561cf932bd7fe2d88c284f5af1e9f25d97010229a42f8fe4c

Observation 2dd63dcc-8a37-4eba-92da-5b710be2637a · outbound

This paper cites Text with knowledge graph augmented transformer for video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization Text with knowledge graph augmented transformer for video captioning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.924672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.605822Z digest=sha256:f858a86b83d8391703d2c6a8557e832f185d59a8f8a29d643b7c5f1127257f0b

Observation 3c44d19c-02c3-41b2-b73a-b3873a620ef5 · outbound

This paper cites Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.778516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.711402Z digest=sha256:c11403c791e99d9cb41de48f4a006c7c24cfd571d4010ee9a409e801ec3a5334

Observation da4843da-867d-484e-a70b-f61e44a77040 · outbound

This paper cites Invariant grounding for video question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Invariant grounding for video question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.551486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.807715Z digest=sha256:deb870a9d4a8d6893bd1e2bb09ea5fc963d3817881c26e8810de3268ccbf4140

Observation a94c8fd1-48a4-4d3b-b56e-525276fc14cb · outbound

This paper cites From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

VUDG: A Dataset for Video Understanding Domain Generalization From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.195284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:09.885422Z digest=sha256:d9e93e7adbdb0b7b4c75f0942a357de287e68e36701a0e8084872ef158aa93e2

Observation 39c95ccd-7480-4c1c-947a-a64ad30d46af · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Morevqa: Exploring modular reasoning models for video question answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:09.991658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:09.991658Z digest=sha256:998b0dfb91d5ec6e00395f9fe556f921a5df436635b07c18c2e472c2d1aca9e2

Observation 719fcd8c-da12-49b7-952c-787b07e52646 · outbound

This paper cites Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.929225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:10.103329Z digest=sha256:cf34f5f8c68f3f06a1c306dd7c8ad8a6c1f97fce9383224754acb706fc597451

Observation c6ae223e-649a-476b-b129-640cd0b54210 · outbound

This paper cites A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.172317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.172317Z digest=sha256:eec9353b60bea8aaf0f60afe6652812d74b5fed8fee1100c11d862f378c72e6d

Observation 9aa4dada-0a55-4abf-a65a-93b597e49f3f · outbound

This paper cites Meta-causal learning for single domain generalization.

VUDG: A Dataset for Video Understanding Domain Generalization Meta-causal learning for single domain generalization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.764934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:10.246214Z digest=sha256:1d201a7c5dc45a8b830d58cb6fd4290cf66568c6584c6f05e7fdd1e3a6a0c406

Observation 5de17fab-d8c3-44db-9614-aade6df7d16b · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.645225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:10.308652Z digest=sha256:aff8bc8a2cfd56b11e9fcc79d95dc09f342b96a018736de94fbcea2bee268a84

Observation 3642f322-b4d7-420c-99d3-6c0c4b4e89ee · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VUDG: A Dataset for Video Understanding Domain Generalization Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.406120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.406120Z digest=sha256:f63a654b4cc01fd7e585783db6890011906a4b792aa4b78a2e29cf0a568f0bfa

Observation 408e7f61-f69b-46cd-97ec-da0311486d57 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VUDG: A Dataset for Video Understanding Domain Generalization VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.535033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.535033Z digest=sha256:269d73b5cb79f5aff30b714114ba68f6493797f67ce5c89baf05f9d5614b2d07

Observation 4d7565d3-43ba-40ea-bf00-649e69218ad5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VUDG: A Dataset for Video Understanding Domain Generalization Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.659981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.659981Z digest=sha256:377395275122e799faec9cdc56b1c27eacd3b2e0e01a3b57517ddf2f1d4ffe36

Observation 04193e7e-8181-459f-9af9-e29e13903829 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

VUDG: A Dataset for Video Understanding Domain Generalization Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.532575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:10.737475Z digest=sha256:2df76f0492677a7383f4791e7c3398cae8b55c756deb6ed63dbaa04891a26bd5

Observation bfef86a0-5c3d-43d5-acfa-ccba42c555b1 · outbound

This paper cites Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments.

VUDG: A Dataset for Video Understanding Domain Generalization Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:32:14.306658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:10.839221Z digest=sha256:2ccf8648d03402020aa5c8e7f683fe2da44d62d0e4e7f7317cf0b9865b3264c9

Observation f001e191-3bef-465f-873e-5f0544e86273 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

VUDG: A Dataset for Video Understanding Domain Generalization Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.400235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:10.912842Z digest=sha256:3c7eb7493a57a61b86a368abd58d2aa45e02606089ec27f176e3d2b9bdbd439c

Observation ca3c054d-30ba-46a5-a16e-d55863fa1781 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.012551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.012551Z digest=sha256:25e2bbfb135f663cb02cbd731d2275c84ed4da8629f4a2937fea04341522d9c1

Observation aa8fcc48-396c-45eb-8395-952d6e62d9e5 · outbound

This paper cites Qwen2.5-VL Technical Report.

VUDG: A Dataset for Video Understanding Domain Generalization Qwen2.5-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.135409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.135409Z digest=sha256:5b2125d371d5f071c84922fa7fe04dd3fc86ea0775de94450b9d49b090479345

Observation 47e1aad9-c1a1-46fb-9736-104483cb5b61 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VUDG: A Dataset for Video Understanding Domain Generalization Activitynet: A large-scale video benchmark for human activity understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.277723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.277723Z digest=sha256:96c8ec29f580c0396123b7cc8b9c8735b15bddcb78325eec488c4a1874ac7878

Observation 0f30c54c-a707-413a-acfd-957ef53fdc5a · outbound

This paper cites The Kinetics Human Action Video Dataset.

VUDG: A Dataset for Video Understanding Domain Generalization The Kinetics Human Action Video Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.404740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.404740Z digest=sha256:1ee4d3aa4a730f2511e592397a57b8fcd09f4df62621fcb1c87b551c35587b41

Observation 12e75b38-7b5e-4f43-91bf-3494471c519f · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

VUDG: A Dataset for Video Understanding Domain Generalization Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.537568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.537568Z digest=sha256:167ea2c222807f6731ea596869fef3837d2a4d94b94e7b3747772e2e3b017671

Observation 3daff457-2f87-4fc7-b2e4-9a7c19385452 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

VUDG: A Dataset for Video Understanding Domain Generalization TVQA: Localized, Compositional Video Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.633968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.633968Z digest=sha256:59f594a277e6ccda1d543ec5b1854adc252f6adbca1b28f004c0270b94e0a60e

Observation 3abd516f-781c-4d5d-bd2e-5341b2b2387d · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VUDG: A Dataset for Video Understanding Domain Generalization Video question answering via gradually refined attention over appearance and motion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.713809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.713809Z digest=sha256:d383827d2887e9b8188bc267008dda248e0e3aaa22f779bdf5f6e5f341ab9044

Observation fe10182d-2607-4e6b-a8cc-aa468c18279e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

VUDG: A Dataset for Video Understanding Domain Generalization Msr-vtt: A large video description dataset for bridging video and language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.808333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.808333Z digest=sha256:fbb33063e025dab0adf66853a747a5e97638c09ae00de07ef9d5998d143fb195

Observation a479c01a-fbd6-4be4-90ed-34515963f71f · outbound

This paper cites Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.181783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:11.907963Z digest=sha256:b6b120610b489205b48ee9a625c521d5f32ed22832e8170573aaa9a621e43bdc

Observation afd06ccc-503f-4e77-90a8-fad64cd4d8f5 · outbound

This paper cites Video-audio domain generalization via confounder disentanglement.

VUDG: A Dataset for Video Understanding Domain Generalization Video-audio domain generalization via confounder disentanglement

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.076518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:11.958263Z digest=sha256:cf8258eba21ef3cb816093e569cb9118a587a40d321626353b5eb7849386b8ac

Observation e8ae2171-134a-421e-a12a-96bcb2751ee4 · outbound

This paper cites Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021.

VUDG: A Dataset for Video Understanding Domain Generalization Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.938790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:12.051541Z digest=sha256:0264d81ba2871ee6b58f1aa7ddd1d5245b96db1022994c16e7b56b9111352a8e

Observation 83944137-ba67-43fa-80a5-680fe3cc3151 · outbound

This paper cites Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022.

VUDG: A Dataset for Video Understanding Domain Generalization Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.819775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:12.169976Z digest=sha256:1b61dab2a8d0775269f8335d67899c5b80527fbc5f0c16752b0dc41e40d7e1c2

Observation 88bec4d3-31fa-4d79-bf07-42b717047a3a · outbound

This paper cites What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations.

VUDG: A Dataset for Video Understanding Domain Generalization What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.731642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:12.279011Z digest=sha256:840b5349fef57aa6b3b30a5df42a1f24e1b78d914e4ffcab7e14509bd294d13b

Observation 3173452b-aaee-44b7-8dec-102394f24a74 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VUDG: A Dataset for Video Understanding Domain Generalization Ego4d: Around the world in 3,000 hours of egocentric video

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.349303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.349303Z digest=sha256:0d01c0e8eb0fd396d1e13376006cea8b686d03dd5d5e9133f7973e4a63adb991

Observation 50943904-c01d-455a-9c31-7a1f84084fe4 · outbound

This paper cites Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection.

VUDG: A Dataset for Video Understanding Domain Generalization Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.579291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:12.469345Z digest=sha256:434919ad7c42267f7ffdceb87d1532433c734268a09dba4f99cc0100d542b472

Observation ea7fe981-ba3d-46fc-b507-43ec8e51e7b8 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.589851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.589851Z digest=sha256:081a2ab4f77fecf459c9bd75b6362b2bd5e30ee09234cdf5c9aa3d670f95db75

Observation 1acd7fa0-0319-4d83-a723-8ff7b95ac313 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.660743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.660743Z digest=sha256:5629957307f45cc3a41a4734c20c0b46c341c8f1d3c407ae823343d76dbdd9ee

Observation dec74c2f-daa6-4af3-993d-154d4e839773 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VUDG: A Dataset for Video Understanding Domain Generalization Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.770007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.770007Z digest=sha256:7c941de4212b49f457c4962b1cdfa82db582972f376f1319e6dfb5e09ef3c7f5

Observation 57437514-aca4-4203-9fdb-6bd9ed8012ea · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VUDG: A Dataset for Video Understanding Domain Generalization TempCompass: Do Video LLMs Really Understand Videos?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.883906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.883906Z digest=sha256:894f4d22d492a4c397d16baa2be96f524b773476f3517474c793318109bf5832

Observation e93b29a3-6889-478a-9775-24288b4fe2fe · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.990057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.990057Z digest=sha256:3e4c67c531008d513ba8590930d182b4bc4e1d4699bb6e467345b97ffc21656a

Observation b0e01d63-607a-4dd7-86f7-e68e3e54acb9 · outbound

This paper cites Poem: polarization of embeddings for domain-invariant representations.

VUDG: A Dataset for Video Understanding Domain Generalization Poem: polarization of embeddings for domain-invariant representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.456435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:13.106658Z digest=sha256:ce489f6360a9a88c076cb978d0e79d91b0fc41fb1b70415edf85fc3cb7bdbacd

Observation de7d5e04-b281-46b1-8e1c-6643da3feda6 · outbound

This paper cites Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning.

VUDG: A Dataset for Video Understanding Domain Generalization Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.297134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:13.180703Z digest=sha256:c9426cd29268e42cb4aed6ef088e7cb62e32843ca1610c0d227bf4682bf48181

Observation f3122d34-d08f-4471-b2fb-b56761562758 · outbound

This paper cites Clifton, and Jie Chen.

VUDG: A Dataset for Video Understanding Domain Generalization Clifton, and Jie Chen

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.126826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:13.233194Z digest=sha256:f0a8d7064bd78d04952df44144a8839eb034002bcf9ec3482c381538004576c1

Observation 88d004b5-88ee-426f-90fa-505684e4b023 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VUDG: A Dataset for Video Understanding Domain Generalization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.326003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.326003Z digest=sha256:b47c271f191df21f42f653a52d1c2464a7ddcf3055742d76cb9313be7452d291

Observation b324e589-68a1-4714-8d08-3deb4559a04d · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

VUDG: A Dataset for Video Understanding Domain Generalization MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.398642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.398642Z digest=sha256:c16f170f93d61774ad841123e40b26dd653e52737d1c3ad8d1e64aa9cca07a75

Observation 7ddc9ed9-fd2d-4530-ba5f-289c8eae72c8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VUDG: A Dataset for Video Understanding Domain Generalization VideoChat: Chat-Centric Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.451229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.451229Z digest=sha256:260e97c7d135d1cf0a3fe0869c7603641dd9266bdd236aae6fe2333a3ac47d57

Observation 52315005-d983-45bb-afc4-a0374560c0ed · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VUDG: A Dataset for Video Understanding Domain Generalization Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.505372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.505372Z digest=sha256:cf57874c6ac30e88c41205171784324ea1ad9db3fa560826ad70cf2984418e55

Observation f53ff2bb-efdb-4858-b175-7b01a95b1168 · outbound

This paper cites mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models.

VUDG: A Dataset for Video Understanding Domain Generalization mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.916081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:13.548397Z digest=sha256:c724c657cac7eeca1fa8dda7405fa0a95dbaeee11cdffcb454adef23507af3bb

Observation 56e612ab-1b36-4961-b76d-e4097da982a5 · outbound

This paper cites Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.721883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:13.607365Z digest=sha256:e7b41d07a754572d95cf2b7ce9c3aa50185d3f8b1beccc0ed71875043d25c1c1

Observation e5057a45-94c8-41c0-9308-1e4fcef91d8a · outbound

This paper cites Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025.

VUDG: A Dataset for Video Understanding Domain Generalization Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.692363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.692363Z digest=sha256:ca58c0a2374a24b256e019fd7af4bc6d76751854b60180655ee173ffd46ba3a4

Observation 1cd6c106-ca4c-4be1-9780-1baf1543045a · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

VUDG: A Dataset for Video Understanding Domain Generalization Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.775272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.775272Z digest=sha256:07f0851756438b05b24127e342c37060a813951eae6346063bab82a082478c5e

Observation 40ee116f-83bc-4a1e-8d60-984fb2908a3f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VUDG: A Dataset for Video Understanding Domain Generalization An image is worth 16x16 words: Transformers for image recognition at scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.822308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.822308Z digest=sha256:599efe8e7c262e29088131f6b6f9dafbb83e2963e2779d64b8a81ca241463e06

Observation b7f291b6-cb27-4868-aac0-081effaf01d9 · outbound

This paper cites C- pack: Packed resources for general chinese embeddings.

VUDG: A Dataset for Video Understanding Domain Generalization C- pack: Packed resources for general chinese embeddings

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.508279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:32:13.889873Z digest=sha256:d6a858e050646c1f64a195b0e7a5a5c7724553eb1f202c49865444c0f4ec39c8

Pith citing papers

No inbound Pith citation observations are available.