Pith. sign in

Paper Citation Record · LEDGER

VUDG: A Dataset for Video Understanding Domain Generalization

As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.24346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24346 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:13.889873Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ac25e71-3ae8-449b-9bd8-4267519269bd · outbound

This paper cites Mm-vit: Multi-modal video transformer for compressed video action recognition.

VUDG: A Dataset for Video Understanding Domain Generalization Mm-vit: Multi-modal video transformer for compressed video action recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.898598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.169622Z digest=sha256:4b43ec2f5c778a9342f75a96318ef421ffec4a619ea7570d3c6270997608d8e9

Observation 408e9a2d-3b76-43e1-8d19-c0db8a1f8b40 · outbound

This paper cites Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.765665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.271553Z digest=sha256:6da650ec940d5957a8948e46d56e8ee9d6309a4c683f8c787e91e4f13fa229b5

Observation d3ca7c93-d783-4fc9-8c8f-d91f16364647 · outbound

This paper cites Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025.

VUDG: A Dataset for Video Understanding Domain Generalization Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.342343Z digest=sha256:c4e2f2d9d8991aa002bde9192dd62fad4b035c2ea927b3a1693f93abdfdc2ce2

Observation 0f4a7285-0452-42b3-a448-a88371dc3454 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization Swinbert: End-to-end transformers with sparse attention for video captioning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.266711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.421238Z digest=sha256:edb9c5b6af0efc98fa513cbe645d268432c0065efffbf29152118e71fa584159

Observation d18ed918-683d-4e62-9802-8732f95e455e · outbound

This paper cites End-to-end generative pretraining for multimodal video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization End-to-end generative pretraining for multimodal video captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.135728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.512771Z digest=sha256:013c8243fa3737da5dc1e1b618afeb480bf5ee2274bdac3baa8a24c0b8dbcfaa

Observation 2dd63dcc-8a37-4eba-92da-5b710be2637a · outbound

This paper cites Text with knowledge graph augmented transformer for video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization Text with knowledge graph augmented transformer for video captioning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.924672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.605822Z digest=sha256:773acd6075ac87a2e8810a8f190a530abd55ecbbe65a105407b250d882100a3f

Observation 3c44d19c-02c3-41b2-b73a-b3873a620ef5 · outbound

This paper cites Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.778516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.711402Z digest=sha256:bda13c148a5620851aab76e9ab2cdc7b38d8aa97abafefca0a05fce891b814c0

Observation da4843da-867d-484e-a70b-f61e44a77040 · outbound

This paper cites Invariant grounding for video question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Invariant grounding for video question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.551486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.807715Z digest=sha256:0aea6ba3734a3ad2095f12986d24cc1f6e8224751e28af275d67c083414ba54d

Observation a94c8fd1-48a4-4d3b-b56e-525276fc14cb · outbound

This paper cites From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

VUDG: A Dataset for Video Understanding Domain Generalization From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.195284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:09.885422Z digest=sha256:7a043481d8dd382bb3a24d812dd4229aff60084bae195f553f9e363f9adfda41

Observation 39c95ccd-7480-4c1c-947a-a64ad30d46af · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Morevqa: Exploring modular reasoning models for video question answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:09.991658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:09.991658Z digest=sha256:e2eadc2b93d34013d146fd9949b11f6fa0b0e97d5fbeb0b5d7b9045833d57fa3

Observation 719fcd8c-da12-49b7-952c-787b07e52646 · outbound

This paper cites Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.929225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:10.103329Z digest=sha256:d438f63ce3e6177a0316a53a8ef4341c24d2fc32451c524caf05cb9821bc6208

Observation c6ae223e-649a-476b-b129-640cd0b54210 · outbound

This paper cites A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.172317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.172317Z digest=sha256:41cb4797bd02776420ab8bf1de81f4b4c059ee77b92e22dff6afdaf6b39ef238

Observation 9aa4dada-0a55-4abf-a65a-93b597e49f3f · outbound

This paper cites Meta-causal learning for single domain generalization.

VUDG: A Dataset for Video Understanding Domain Generalization Meta-causal learning for single domain generalization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.764934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:10.246214Z digest=sha256:601c2776b46179eb74495c8bb37fa15a71f3e9489b16d78d7d627df9f0f128d8

Observation 5de17fab-d8c3-44db-9614-aade6df7d16b · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.645225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:10.308652Z digest=sha256:a11c50155adecd5b3424fddd5a459f902e2330e2ac6661c77fe3cab3c879a402

Observation 3642f322-b4d7-420c-99d3-6c0c4b4e89ee · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VUDG: A Dataset for Video Understanding Domain Generalization Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.406120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.406120Z digest=sha256:8e453ae572eb401a821f33cd25e134a281c86b6030d1f29008677daa595813d4

Observation 408e7f61-f69b-46cd-97ec-da0311486d57 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VUDG: A Dataset for Video Understanding Domain Generalization VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.535033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.535033Z digest=sha256:28a5b3b174e37db7e2a0d7c71ce3b0769c0e11ffa17211d75dff5ee859216fbb

Observation 4d7565d3-43ba-40ea-bf00-649e69218ad5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VUDG: A Dataset for Video Understanding Domain Generalization Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.659981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.659981Z digest=sha256:5d74643d2915d607acbaec977eb875c1cafc718b452ddb2dd5f78f6022cbc5d5

Observation 04193e7e-8181-459f-9af9-e29e13903829 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

VUDG: A Dataset for Video Understanding Domain Generalization Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.532575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:10.737475Z digest=sha256:e2ca3b555a0c1e5defa5d344c37e93b503aa66c33b76c8c3e4ec3c0e9e4a24fc

Observation bfef86a0-5c3d-43d5-acfa-ccba42c555b1 · outbound

This paper cites Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments.

VUDG: A Dataset for Video Understanding Domain Generalization Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:32:14.306658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:10.839221Z digest=sha256:ece71ee0a70c192ce57a5a6484f74240d086f8c504951345f10d42476808a313

Observation f001e191-3bef-465f-873e-5f0544e86273 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

VUDG: A Dataset for Video Understanding Domain Generalization Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.400235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:10.912842Z digest=sha256:b6c85989e1889a7210846acbaf17ae8de5546dea3f4faea57f0ff2a61546303d

Observation ca3c054d-30ba-46a5-a16e-d55863fa1781 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.012551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.012551Z digest=sha256:389ab729320cca974473be2dd0077b37bbe17ee309904902898bba1c5f94b239

Observation aa8fcc48-396c-45eb-8395-952d6e62d9e5 · outbound

This paper cites Qwen2.5-VL Technical Report.

VUDG: A Dataset for Video Understanding Domain Generalization Qwen2.5-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.135409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.135409Z digest=sha256:f936863b913f26601cde143403a36f0d04491a71a8ebd24e9b79a7a4dcfebe41

Observation 47e1aad9-c1a1-46fb-9736-104483cb5b61 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VUDG: A Dataset for Video Understanding Domain Generalization Activitynet: A large-scale video benchmark for human activity understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.277723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.277723Z digest=sha256:a5acdd2f9b500304e8a1138ae7961868bed40a0b2c143dca40a4b9d7d2777f40

Observation 0f30c54c-a707-413a-acfd-957ef53fdc5a · outbound

This paper cites The Kinetics Human Action Video Dataset.

VUDG: A Dataset for Video Understanding Domain Generalization The Kinetics Human Action Video Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.404740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.404740Z digest=sha256:1ad6d821d7657d5bea50e078ba20bfd6d383da244e7cdaabd088a2bb8771d6a4

Observation 12e75b38-7b5e-4f43-91bf-3494471c519f · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

VUDG: A Dataset for Video Understanding Domain Generalization Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.537568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.537568Z digest=sha256:3cd93c3c4f4ea8de05b31aaec9a17affecfd62b2dfe027a43591aab386aa1fc0

Observation 3daff457-2f87-4fc7-b2e4-9a7c19385452 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

VUDG: A Dataset for Video Understanding Domain Generalization TVQA: Localized, Compositional Video Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.633968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.633968Z digest=sha256:ab93630eaf8afc5e39c0c44b23c369172219f36613fb0fb86688c84280c35de2

Observation 3abd516f-781c-4d5d-bd2e-5341b2b2387d · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VUDG: A Dataset for Video Understanding Domain Generalization Video question answering via gradually refined attention over appearance and motion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.713809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.713809Z digest=sha256:84449368cb7fca353302cede7b6a553fc4833740680d91d477fe536b3e1271e2

Observation fe10182d-2607-4e6b-a8cc-aa468c18279e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

VUDG: A Dataset for Video Understanding Domain Generalization Msr-vtt: A large video description dataset for bridging video and language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.808333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.808333Z digest=sha256:05aeba093c4e002825dcf89c558cbff237f115ebfa794b9af879841fd29b5fba

Observation a479c01a-fbd6-4be4-90ed-34515963f71f · outbound

This paper cites Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.181783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:11.907963Z digest=sha256:c6e94fe07c501db5c4a5616dcb3252f05e5eccbdf10c66876722d3a7425940f3

Observation afd06ccc-503f-4e77-90a8-fad64cd4d8f5 · outbound

This paper cites Video-audio domain generalization via confounder disentanglement.

VUDG: A Dataset for Video Understanding Domain Generalization Video-audio domain generalization via confounder disentanglement

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.076518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:11.958263Z digest=sha256:5817b701079d942bca032c11b4856752e769f5907188efe06252f989bb68cc28

Observation e8ae2171-134a-421e-a12a-96bcb2751ee4 · outbound

This paper cites Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021.

VUDG: A Dataset for Video Understanding Domain Generalization Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.938790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:12.051541Z digest=sha256:226e208ca2cde595aec684749d499695b932faac51c2ff170f227568019815c8

Observation 83944137-ba67-43fa-80a5-680fe3cc3151 · outbound

This paper cites Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022.

VUDG: A Dataset for Video Understanding Domain Generalization Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.819775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:12.169976Z digest=sha256:bb3bac868df51781d8b4eea5dce1253c17e2229a1bee0cd5d176f3eeb7698dd5

Observation 88bec4d3-31fa-4d79-bf07-42b717047a3a · outbound

This paper cites What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations.

VUDG: A Dataset for Video Understanding Domain Generalization What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.731642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:12.279011Z digest=sha256:baf7060c5f157ce5ecb7eda0b695b3a78c402d739a9fe56f48b1edd23215095f

Observation 3173452b-aaee-44b7-8dec-102394f24a74 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VUDG: A Dataset for Video Understanding Domain Generalization Ego4d: Around the world in 3,000 hours of egocentric video

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.349303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.349303Z digest=sha256:2244a9ec5366445eebb474c448bf4e53087550ef59a771bdafa6602830f10a1e

Observation 50943904-c01d-455a-9c31-7a1f84084fe4 · outbound

This paper cites Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection.

VUDG: A Dataset for Video Understanding Domain Generalization Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.579291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:12.469345Z digest=sha256:bec336103d87ccbd3d78886a1371a38d8f58cd51569d696ec34dd7119848c253

Observation ea7fe981-ba3d-46fc-b507-43ec8e51e7b8 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.589851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.589851Z digest=sha256:e1cd29b789ad71bb120557bff28239c7f3e44674272a4b0b451693bd57276f9c

Observation 1acd7fa0-0319-4d83-a723-8ff7b95ac313 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.660743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.660743Z digest=sha256:147637cc1cc1bff3eec3d1ffa922c90f2b0eaa25df5c6ba044daabd5923106d0

Observation dec74c2f-daa6-4af3-993d-154d4e839773 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VUDG: A Dataset for Video Understanding Domain Generalization Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.770007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.770007Z digest=sha256:873a0e8347389010f0f1e9deadcbf51c8df2305b08cf2a4be4718fbd07de2c9e

Observation 57437514-aca4-4203-9fdb-6bd9ed8012ea · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VUDG: A Dataset for Video Understanding Domain Generalization TempCompass: Do Video LLMs Really Understand Videos?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.883906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.883906Z digest=sha256:43a18df82ce02f3fa1ae1af65dbcd6b2a2ff160f93c89b9f0b662aee9244ed1e

Observation e93b29a3-6889-478a-9775-24288b4fe2fe · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.990057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.990057Z digest=sha256:03350ea1417d3cade55cbfd3d1bbe82f7ceacda2b46381cf5ff630660f572002

Observation b0e01d63-607a-4dd7-86f7-e68e3e54acb9 · outbound

This paper cites Poem: polarization of embeddings for domain-invariant representations.

VUDG: A Dataset for Video Understanding Domain Generalization Poem: polarization of embeddings for domain-invariant representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.456435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:13.106658Z digest=sha256:89d6a9388e5df3d02410338a5a78e6cd9a530348adda7a725e499ec9acc5d338

Observation de7d5e04-b281-46b1-8e1c-6643da3feda6 · outbound

This paper cites Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning.

VUDG: A Dataset for Video Understanding Domain Generalization Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.297134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:13.180703Z digest=sha256:2566e61431996e4e7ff153e2aa90ef60d31342b7e51fdce87733583553bf4a84

Observation f3122d34-d08f-4471-b2fb-b56761562758 · outbound

This paper cites Clifton, and Jie Chen.

VUDG: A Dataset for Video Understanding Domain Generalization Clifton, and Jie Chen

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.126826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:13.233194Z digest=sha256:5e695169a97cfcf2bbe5cee32e778dcb7ba527bf76192c2315f13477a1800978

Observation 88d004b5-88ee-426f-90fa-505684e4b023 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VUDG: A Dataset for Video Understanding Domain Generalization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.326003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.326003Z digest=sha256:35c41cb6e0f5d35d0be231847436b23b057cb5247b17c664e85632627393bff5

Observation b324e589-68a1-4714-8d08-3deb4559a04d · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

VUDG: A Dataset for Video Understanding Domain Generalization MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.398642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.398642Z digest=sha256:4bf631cb25e169cdffa802bea8a2a4d6fcf8b3aac034de62ba8ac7597a1e9e45

Observation 7ddc9ed9-fd2d-4530-ba5f-289c8eae72c8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VUDG: A Dataset for Video Understanding Domain Generalization VideoChat: Chat-Centric Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.451229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.451229Z digest=sha256:e3b8b331ed9dcd714c74d7f3a4a254bb7d08441b7374dfea5fbe5d354f5610a9

Observation 52315005-d983-45bb-afc4-a0374560c0ed · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VUDG: A Dataset for Video Understanding Domain Generalization Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.505372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.505372Z digest=sha256:e755d4e7b0871f8dc1511978742fa171d97632893f19bd538384eb4546973cbd

Observation f53ff2bb-efdb-4858-b175-7b01a95b1168 · outbound

This paper cites mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models.

VUDG: A Dataset for Video Understanding Domain Generalization mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.916081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:13.548397Z digest=sha256:82c579db86cd87937b3a36dda35bf0bdfb7568498c92ecf6341fd497e36ecaa5

Observation 56e612ab-1b36-4961-b76d-e4097da982a5 · outbound

This paper cites Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.721883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:13.607365Z digest=sha256:e51f6b01d2ad8afb030638697771f856cab970c558f415baf60d83919350ec5a

Observation e5057a45-94c8-41c0-9308-1e4fcef91d8a · outbound

This paper cites Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025.

VUDG: A Dataset for Video Understanding Domain Generalization Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.692363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.692363Z digest=sha256:16250e4a2451ad89582e80883adeb7411d165e38c53fbbecd2e18a606a89a83a

Observation 1cd6c106-ca4c-4be1-9780-1baf1543045a · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

VUDG: A Dataset for Video Understanding Domain Generalization Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.775272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.775272Z digest=sha256:c9cbadc5be5ba48c5be469675277b4402fb6ad8ebeff081b0b6ca0e1950c87a0

Observation 40ee116f-83bc-4a1e-8d60-984fb2908a3f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VUDG: A Dataset for Video Understanding Domain Generalization An image is worth 16x16 words: Transformers for image recognition at scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.822308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.822308Z digest=sha256:f39b64dea7f77f84e5c98273fa077935a4c04b057b9c90ec6b9f1a1bbde2c4c1

Observation b7f291b6-cb27-4868-aac0-081effaf01d9 · outbound

This paper cites C- pack: Packed resources for general chinese embeddings.

VUDG: A Dataset for Video Understanding Domain Generalization C- pack: Packed resources for general chinese embeddings

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.508279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:32:13.889873Z digest=sha256:28d8bd34f361aa5cd13dbe2ba776e327ddfb84b20efe6e67b17848cfeee5fdbf

Pith citing papers

No inbound Pith citation observations are available.