Pith. sign in

Paper Citation Record · LEDGER

Video Understanding by Design: How Datasets Shape Video Models

As of 22 August 2026, this Paper Citation Record lists 100 of 250 outbound references and 0 inbound Pith citation observations for arXiv:2509.09151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09151 v2

Coverage vector

measured 100 of 250 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:37:41.849450Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 250 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af737d4e-a076-4b85-807d-542a4c852b3e · outbound

This paper cites Large-scale video classification with convolutional neural networks,.

Video Understanding by Design: How Datasets Shape Video Models Large-scale video classification with convolutional neural networks,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.562439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.562439Z digest=sha256:fa81a154e5047b2e91a8a769c62b27fe6d04116a61e88f66fee4b72c8a109804

Observation 3cf5d4d4-2926-499f-8c04-c9f76b5665c2 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

Video Understanding by Design: How Datasets Shape Video Models Learning spatiotemporal features with 3d convolutional networks,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.566161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.566161Z digest=sha256:326e6fc95b53efb8e63dfd2b107d8e560d2a297667d490b86f91bebaad42303d

Observation b7b83d8b-aaa4-4957-8ed0-89dbe96e8c9c · outbound

This paper cites Slowfast networks for video recognition,.

Video Understanding by Design: How Datasets Shape Video Models Slowfast networks for video recognition,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.568966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.568966Z digest=sha256:53bed4c83f8dfc19e75293bf910eda3066ddc984cdc275db3d742cb826ed1cdb

Observation 61ed61d3-9c5c-4a18-9fff-f604b1d8a41c · outbound

This paper cites Motion meets attention: Video motion prompts,.

Video Understanding by Design: How Datasets Shape Video Models Motion meets attention: Video motion prompts,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.572025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.572025Z digest=sha256:7fcab49d8005614085e4c43d304ca90ad771690b3a878f7316113c3f50e261f5

Observation 6a9fe7c2-6a88-4b2e-b361-8a1710f8472c · outbound

This paper cites Taylor videos for action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Taylor videos for action recognition,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.574797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.574797Z digest=sha256:aba18c0c46fa9529093f484b40ed56294d36b4b1ff5ebfac790000483dc96e32

Observation 698eff6a-6fdd-4dd2-aa22-1c94c745bf3b · outbound

This paper cites Learnable expansion of graph operators for multi-modal feature fusion,.

Video Understanding by Design: How Datasets Shape Video Models Learnable expansion of graph operators for multi-modal feature fusion,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.578114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.578114Z digest=sha256:9f55a0135c9c52b802ce78754962bf71c2ba06c6eaf672c364187a4ee376a0d4

Observation bc36bfda-5915-4ca3-8b9e-d8733b267178 · outbound

This paper cites Meet jeanie: a similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,.

Video Understanding by Design: How Datasets Shape Video Models Meet jeanie: a similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.580999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.580999Z digest=sha256:0aa72833ee10c8164c35d49a4785c9cc438f09fb9f530dd529c92f6b3cb019f8

Observation 88974bff-5613-40a0-a9d5-c013942bb74a · outbound

This paper cites Evolving skeletons: Motion dynamics in action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Evolving skeletons: Motion dynamics in action recognition,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.583680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.583680Z digest=sha256:cd3927f84beecc17c45cf90291d85893188596eb084d379f7ad65175a6f8cbfc

Observation d5f942de-b118-4514-ad83-2a2d41453837 · outbound

This paper cites Feature hallucination for self-supervised action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Feature hallucination for self-supervised action recognition,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.586255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.586255Z digest=sha256:0dcf750c3c94a34e56c2af14532a8ae01e0a187147b750812de3133f643d5605

Observation aa1e96b8-b01f-4e79-801a-75d6ac127e2a · outbound

This paper cites Do language models understand time?.

Video Understanding by Design: How Datasets Shape Video Models Do language models understand time?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.589156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.589156Z digest=sha256:e2681d12292e55bc2fef96b1b020c3ed6dc132ab1b7dc68ab5259b417f76c48e

Observation a144318c-9692-4e1a-b4cc-a8343442d465 · outbound

This paper cites Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight.

Video Understanding by Design: How Datasets Shape Video Models Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.591934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.591934Z digest=sha256:35ebf8623f4bad8163f0ad823fe44c033eb68c046c7ae89f0fd52a6eb12764b7

Observation e564258f-e0eb-4074-8427-11388387187a · outbound

This paper cites The journey of action recognition,.

Video Understanding by Design: How Datasets Shape Video Models The journey of action recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.595267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.595267Z digest=sha256:c85d181fa3ca962c06de307c5a1f743de02e9278280d18225715ea574b378629

Observation 05e508ef-6369-42da-b8dc-847b6e89cf23 · outbound

This paper cites Representation-centric survey of skeletal action recognition and the anubis benchmark,.

Video Understanding by Design: How Datasets Shape Video Models Representation-centric survey of skeletal action recognition and the anubis benchmark,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.598047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.598047Z digest=sha256:5a4fd2d0ce40baf2554f5fb3a37337b6ac8cc722856bc4f100554577e60302e6

Observation 217712fc-9fa7-468a-8a99-833d6bb663f6 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

Video Understanding by Design: How Datasets Shape Video Models Foundation Models for Video Understanding: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.600749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.600749Z digest=sha256:9ed808360d3d7b6c87e4249c890a7578ff6ba22d558bb07560a68cc20b5bb1aa

Observation 6885caf8-9ccd-4b8a-bb03-9733f743dce4 · outbound

This paper cites Video understanding with large language models: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Video understanding with large language models: A survey,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.603790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.603790Z digest=sha256:e9bef2f27a27f46a45dc9f28b449c1419e831b386773e5edabacc38031f80cbd

Observation 3a40fdea-d68d-43ba-86ac-02ca168bea99 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video Understanding by Design: How Datasets Shape Video Models The Kinetics Human Action Video Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.606822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.606822Z digest=sha256:051a184247c19311c0881f90d96b68982e4032369583dee3b76cace0124c5e5a

Observation f95e75f5-7ea5-4e24-af56-dbd9cb00ae5a · outbound

This paper cites The” something something.

Video Understanding by Design: How Datasets Shape Video Models The” something something

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.610179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.610179Z digest=sha256:9bd844a51f344ed4fb780523455b4862e3c2e1cf8364cbf788b784878160e675

Observation f2f0f853-6642-498a-adf2-59609f1058df · outbound

This paper cites Activi- tynet: A large-scale video benchmark for human activity understanding,.

Video Understanding by Design: How Datasets Shape Video Models Activi- tynet: A large-scale video benchmark for human activity understanding,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.612970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.612970Z digest=sha256:63bdfb302574922e642f307bb2ec890ef56bed54a14b8c3c36fd1009260dc154

Observation 1ebc26bc-7b7d-4b17-9189-c162a414c5c0 · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding,.

Video Understanding by Design: How Datasets Shape Video Models Hollywood in homes: Crowdsourcing data collection for activity understanding,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.615770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.615770Z digest=sha256:0e9748cd12b4ab052e6d95f446e19182872059914c969c3fa8282ef4fdb5feac

Observation d1c36766-1ff9-4d51-a775-fae9c971cb8b · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

Video Understanding by Design: How Datasets Shape Video Models Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.618568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.618568Z digest=sha256:c28263ad78906e25d611bd10e138b389c2107b7be461cb4f72b7c1bff2ce067a

Observation 1f3b6bb5-81c6-46bb-be11-5340198aa558 · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions,.

Video Understanding by Design: How Datasets Shape Video Models Ava: A video dataset of spatio-temporally localized atomic visual actions,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.621704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.621704Z digest=sha256:7323aafc8377d4ba3c35ecc8e81063a254095c1e24c656412926047b70035990

Observation 81b15cac-a4d8-427a-a154-3a2047258c66 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,.

Video Understanding by Design: How Datasets Shape Video Models Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.624382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.624382Z digest=sha256:4d5b998a65d2d65d088b9c1c1ba62bf161dca48efc90f74ab42df8eb5a025c77

Observation 58436897-7102-4e39-b35d-b1844c8ad7a7 · outbound

This paper cites Graph based skeleton motion representation and similarity measurement for action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Graph based skeleton motion representation and similarity measurement for action recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.626977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.626977Z digest=sha256:6460863d015e4b24bf666abb8aef1cc16cf9a5228e0565b1c89e157852262869

Observation c7b717d8-4c90-4cf4-be6c-e95a231864d8 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

Video Understanding by Design: How Datasets Shape Video Models Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.629850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.629850Z digest=sha256:1d575b458b519129f6ebc239ad5bedeed7f6a3c16a45950f7a7a741d2caf2a0c

Observation 409b2f64-7add-48cc-b93e-907490fc1b3f · outbound

This paper cites Non-local neural networks,.

Video Understanding by Design: How Datasets Shape Video Models Non-local neural networks,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.632560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.632560Z digest=sha256:b1695fc8731221b88c68d152e1d7053d02dfa7e68c28ca49ba1aaa64f1babd18

Observation 805c6334-b76e-4731-8353-3d4e1412f939 · outbound

This paper cites Spatial temporal graph convolutional networks for skeleton-based action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Spatial temporal graph convolutional networks for skeleton-based action recognition,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.635238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.635238Z digest=sha256:b266799f708e798ad10398f5362d8876bf3c8a1210e2a6c35b8d578b8af7be87

Observation 34a16798-62c1-4b3b-bb86-7acda8cb0242 · outbound

This paper cites Vivit: A video vision transformer,.

Video Understanding by Design: How Datasets Shape Video Models Vivit: A video vision transformer,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.637882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.637882Z digest=sha256:7083637e2a41d56bc2239ecc9fa710515478ff0ffa672f4e423cb4893335c3a3

Observation fac54d79-ad37-4493-a215-5d8ba1158541 · outbound

This paper cites Is space-time attention all you need for video understanding?.

Video Understanding by Design: How Datasets Shape Video Models Is space-time attention all you need for video understanding?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.640602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.640602Z digest=sha256:d3f129262e07c350e1b6b75afe7db9c6f3bd5b1181da91f5cc2afc4407e1205d

Observation 1df2569b-70fe-4751-bfb6-68b4ae774b40 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,.

Video Understanding by Design: How Datasets Shape Video Models Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.643269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.643269Z digest=sha256:c1c18041d3c363ffe1d3158f1bb9c786d0650670e3adb43b287206c2486015f9

Observation a0fc6081-e0ab-4d7b-aab9-69f4f9ca65f5 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Video Understanding by Design: How Datasets Shape Video Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.645943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.645943Z digest=sha256:37a3db2169dee2021be679517e503a36aa6f8496656a54dcaeb2fb9c6e682128

Observation 55acbbb5-9c64-4323-b605-b92249014f24 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking,.

Video Understanding by Design: How Datasets Shape Video Models Videomae v2: Scaling video masked autoencoders with dual masking,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.648984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.648984Z digest=sha256:0a711c36183e96103642bc23e261558aa78ac56dd3a978c9246de8ff89a0e474

Observation 18331253-4077-4f84-a368-b481ae0ab25b · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding,.

Video Understanding by Design: How Datasets Shape Video Models Internvideo2: Scaling foundation models for multimodal video understanding,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.651787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.651787Z digest=sha256:cc2a0655d8a35d31107aa6fca7e36b5f6015b5db9ac7361304f6f08d5651583c

Observation 8197175e-9cfb-4a28-8582-471dc2d84705 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Video Understanding by Design: How Datasets Shape Video Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.654342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.654342Z digest=sha256:84316986b88b1b2c6d085d4719e507f93ee22c0cdb53d71ce459d753cd6ddaf9

Observation 16c3da7b-43bb-4d12-aed1-a73a304493db · outbound

This paper cites Hmdb: a large video database for human motion recognition,.

Video Understanding by Design: How Datasets Shape Video Models Hmdb: a large video database for human motion recognition,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.657439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.657439Z digest=sha256:1aa8d114f76d1599a5b288ba3f78d3be586bd2817d97db8602f71ba4258627f3

Observation 35516b6a-f3be-4ac2-a7d1-686491f95699 · outbound

This paper cites A spatio-temporal descriptor based on 3d-gradients,.

Video Understanding by Design: How Datasets Shape Video Models A spatio-temporal descriptor based on 3d-gradients,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.660022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.660022Z digest=sha256:067a0b760be78562a3a7fb32957d767ddf31d1c97fd90332adc5d3e00f08e46f

Observation 5a8a63d9-f1f4-456a-a99c-f4a1312c2794 · outbound

This paper cites Action recognition with improved trajectories,.

Video Understanding by Design: How Datasets Shape Video Models Action recognition with improved trajectories,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.662771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.662771Z digest=sha256:8a216071361ecc9fea95493aa900dba9e5732427777997d98048b945ca9d029e

Observation 63fc5404-8e68-459b-8621-267e1d8f186e · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset,.

Video Understanding by Design: How Datasets Shape Video Models Scaling egocentric vision: The epic-kitchens dataset,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.665375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.665375Z digest=sha256:d9d4c20abf3b9e3ac3ecf2787a407d0c7bb2aad60e4c314fc3d05bafb25e0a10

Observation ad772c31-d47f-48bf-91cc-bcf662bebb83 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos,.

Video Understanding by Design: How Datasets Shape Video Models Two-stream convolutional networks for action recognition in videos,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.667946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.667946Z digest=sha256:3f9ab4fdd2f996af16c38c388a128fa9d3cf5cd8853aba19414f33e9e20ce37c

Observation d2dae815-9191-46da-885b-c4957829917f · outbound

This paper cites Omnivl: One foundation model for image-language and video-language tasks,.

Video Understanding by Design: How Datasets Shape Video Models Omnivl: One foundation model for image-language and video-language tasks,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.670604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.670604Z digest=sha256:2356a4120368359052b68a7a97228be8978e6f9383cb72d755793ef600c30b2b

Observation dcf3a715-2126-4325-8124-ed369c523a1d · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,.

Video Understanding by Design: How Datasets Shape Video Models Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.673640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.673640Z digest=sha256:078707b05e67862ee664d8132b7dc83c1a6f320e0216e269ecc61e94511473cd

Observation 435b8b26-c420-4acd-a4e5-dba506fbf4f0 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

Video Understanding by Design: How Datasets Shape Video Models Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.676601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.676601Z digest=sha256:5d3fb995e475c6bba8b2fc6898d13fefff0235c90426ef623c78cfa62b777f2f

Observation 9b5490e0-2e3a-4634-999e-122cbdad0bbe · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video Understanding by Design: How Datasets Shape Video Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.679514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.679514Z digest=sha256:1deff3e22f6a7caacede89c33563864ed3aa7974006a06e5dff8d35aad01132a

Observation 5207923e-f815-4db3-9b59-2e52ba6f6ad1 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,.

Video Understanding by Design: How Datasets Shape Video Models Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.682433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.682433Z digest=sha256:872c9408928fc6ebda44d30e1220c593a6ce80a21fe444d70a873ebf0da89821

Observation 4dcda8f1-40f6-49e0-868c-1270c6a1440d · outbound

This paper cites Masked feature prediction for self-supervised visual pre-training,.

Video Understanding by Design: How Datasets Shape Video Models Masked feature prediction for self-supervised visual pre-training,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.685162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.685162Z digest=sha256:4b235ce5871686d76da685a12e769e2998503f81379c6b37d4ecb1a2debeac3f

Observation 77021111-4909-4dfc-8655-dbb89abc8b61 · outbound

This paper cites Transductive zero-shot action recog- nition by word-vector embedding,.

Video Understanding by Design: How Datasets Shape Video Models Transductive zero-shot action recog- nition by word-vector embedding,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.687849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.687849Z digest=sha256:96309c867eccacfb2f27dd1365b4cfc05a46a27229a1f855be5aaf533f9e7b0d

Observation 6219dd9c-0f4b-4d6b-bf6c-d006c4927ebd · outbound

This paper cites Out-of-distribution detection for generalized zero-shot action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Out-of-distribution detection for generalized zero-shot action recognition,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.690645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.690645Z digest=sha256:657124ae5bcf19d8acdf96096818e67add420ce01695a2d9d4ca16251f9c8e6b

Observation d7526746-074c-4203-a8ec-2f782c923469 · outbound

This paper cites Few-shot action recognition with permutation-invariant attention,.

Video Understanding by Design: How Datasets Shape Video Models Few-shot action recognition with permutation-invariant attention,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.693367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.693367Z digest=sha256:597147e7c8d78d6c288406c8cd302fb584eb8a58cb79f926335415c01701a163

Observation d81fe392-c880-4688-9e6e-3676c10a83e5 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Video Understanding by Design: How Datasets Shape Video Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.696663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.696663Z digest=sha256:67182ee6dbe5944d7cb44a16137da4b3c919fbb8ba79e88877e8de21ac1891c1

Observation 0853462e-8aa3-42fd-bbba-fb31e8679276 · outbound

This paper cites Temporal-relational crosstransformers for few-shot action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Temporal-relational crosstransformers for few-shot action recognition,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.699682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.699682Z digest=sha256:ddbffb084ae8dd4c1997c119dca0b292bec9005efc1ae03459ff482df5caab7b

Observation 77d5d7f8-c3a1-4acc-9e7a-a51fe6cd4de3 · outbound

This paper cites Temporal-viewpoint transportation plan for skeletal few-shot action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Temporal-viewpoint transportation plan for skeletal few-shot action recognition,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.702438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.702438Z digest=sha256:1fc89966d6bd8fcfd754ea42cbf945816461ac2219ae2094ab9d6fe8e05a6018

Observation d1dfd0bd-09fc-4b29-bac9-645a425db9ef · outbound

This paper cites Uncertainty-dtw for time series and sequences,.

Video Understanding by Design: How Datasets Shape Video Models Uncertainty-dtw for time series and sequences,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.705135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.705135Z digest=sha256:aa1927b508158ca10e2114bcc261cd66df6270fa28c596150ddfed81433d7f6a

Observation 525169ae-a5e4-403a-b3a6-d6cd8e947381 · outbound

This paper cites Reinforced Video Captioning with Entailment Rewards.

Video Understanding by Design: How Datasets Shape Video Models Reinforced Video Captioning with Entailment Rewards

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.707954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.707954Z digest=sha256:d1c63a3f5a8e37893b67910bc120554fc36ac0b8f5e8e8ff1a24fe140215a2a2

Observation ff198091-dbd3-4222-a246-9883e9874bfa · outbound

This paper cites Embodied question answering,.

Video Understanding by Design: How Datasets Shape Video Models Embodied question answering,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.710825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.710825Z digest=sha256:4d7b69a5d6602027cbe6cef6e647ee53024b1d38012ce86b01ab89d924111012

Observation 107ce85e-d625-49ce-a115-a097141dbf44 · outbound

This paper cites Merlot: Multimodal neural script knowledge models,.

Video Understanding by Design: How Datasets Shape Video Models Merlot: Multimodal neural script knowledge models,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.713506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.713506Z digest=sha256:2e7b61ee260fb28cfeb4e25ec9d9c24fdcb4b63f6c3ee1fcc3b6758e3b2b11a8

Observation cf8a4492-20c5-4de6-9030-5872e0641fe7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Video Understanding by Design: How Datasets Shape Video Models Flamingo: a visual language model for few-shot learning,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.716087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.716087Z digest=sha256:32042de253f303ff5f0f609ae7aea2f1fd977cf0ac727b8ff781bafcd90db4e9

Observation f2bbb6d3-55b3-4173-b271-9799c40e42b6 · outbound

This paper cites Human action recognition from various data modalities: A review,.

Video Understanding by Design: How Datasets Shape Video Models Human action recognition from various data modalities: A review,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.719093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.719093Z digest=sha256:4d66f94d27ec5f798625aaece69177f7ca2bf70a920ff3d80f268e5be10df89b

Observation acd93843-1245-4273-9ddf-fc8d3bb2bc92 · outbound

This paper cites Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives.

Video Understanding by Design: How Datasets Shape Video Models Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.721882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.721882Z digest=sha256:f72ac6142e6e97c42aa28bc2bf8101dc254e7874661b76381b5441064383f21a

Observation 0c106935-f0ae-487d-9b55-44be443ebbff · outbound

This paper cites Video question answering: A survey of the state-of-the-art,.

Video Understanding by Design: How Datasets Shape Video Models Video question answering: A survey of the state-of-the-art,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.724813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.724813Z digest=sha256:746c662b09ac84a705ce351e70b70d4887045ade0bac48a6959d7696765297bb

Observation 31bcdc0b-7e56-40b2-81a7-758c6c21f8e3 · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

Video Understanding by Design: How Datasets Shape Video Models A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.730222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.730222Z digest=sha256:d9d15c5bbf99b0fb455cdf915df1b040f27598a437bfc0a3d5e6759f29496825

Observation c2c92449-c826-482c-a6e2-524519aeb595 · outbound

This paper cites Human activity analysis: A review,.

Video Understanding by Design: How Datasets Shape Video Models Human activity analysis: A review,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.733204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.733204Z digest=sha256:7e76e94a864cfffed31202b264cb977fb2b07d18314271ef7f6f2413fbdfbe10

Observation fd60c122-08a2-4ac8-a417-484c94867174 · outbound

This paper cites Going deeper into action recognition: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Going deeper into action recognition: A survey,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.735946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.735946Z digest=sha256:b32787615dd12ef7592388a4292f99731c018ebdfd66b80fd115c0b42039b908

Observation 7484f816-5a30-4e7a-9785-c74dc3bc7b56 · outbound

This paper cites A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,.

Video Understanding by Design: How Datasets Shape Video Models A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.738840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.738840Z digest=sha256:80ecb365bd002cd0a6a1ab8a859d4116206223e0e708299e0cf51a32736aaec5

Observation b9ed3d35-625b-44b0-a307-ecee01b7523f · outbound

This paper cites Vision Transformers for Action Recognition: A Survey.

Video Understanding by Design: How Datasets Shape Video Models Vision Transformers for Action Recognition: A Survey

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.742070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.742070Z digest=sha256:1c6a7d99d27a990bce16b7c65bd0103ac5b85c2e3f547ad0da7b97edf22d9fcf

Observation 554aeb2d-04fa-4b0d-a5d9-afeedeadee1a · outbound

This paper cites Video transformers: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Video transformers: A survey,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.745553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.745553Z digest=sha256:50a0833b620442e3dcdba7bc9423ab0b94370a6cd3bd7c38e48d4095719e5d62

Observation 08c4997d-03ab-4165-8fc9-7f2f7dec14ee · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos,.

Video Understanding by Design: How Datasets Shape Video Models End-to-end learning of visual representations from uncurated instruc- tional videos,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.749025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.749025Z digest=sha256:8c0487606f0642b136477f1a31bb4f2a0f1a706c7fdbd648cf95fc6ff9fc16d0

Observation 874c92ce-cc18-49f2-8bf4-abb8e70106f2 · outbound

This paper cites Multimodal learning with transform- ers: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Multimodal learning with transform- ers: A survey,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.751971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.751971Z digest=sha256:663f8e2545955d23e5f6473a9453d5e1647cd245e609eeb323354b07655b743b

Observation 52a66120-939a-45d9-9441-3f6016525276 · outbound

This paper cites A survey on human activity recognition from videos,.

Video Understanding by Design: How Datasets Shape Video Models A survey on human activity recognition from videos,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.755142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.755142Z digest=sha256:4f8315eb1269fb1838fe8269d0fd3ff68525e4e58d9aaa0f64b6644886cd2d4a

Observation c39646f6-9189-41a6-8854-4eecc7651061 · outbound

This paper cites A comparative review of recent kinect-based action recognition algorithms,.

Video Understanding by Design: How Datasets Shape Video Models A comparative review of recent kinect-based action recognition algorithms,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.758367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.758367Z digest=sha256:bf2610c664084b31b3154ce6ef7d5328b1d2edd44653ce59b0e3d47a54560178

Observation 8054463e-bc0c-48d7-9725-c6400020047f · outbound

This paper cites Skeleton- based action recognition with shift graph convolutional network,.

Video Understanding by Design: How Datasets Shape Video Models Skeleton- based action recognition with shift graph convolutional network,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.761547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.761547Z digest=sha256:ff739e9b6fba9e5c25ede8d5f8aff2ed5e6f9db3746c229f816237bba2856c4d

Observation 25a293b3-e079-4a0d-a344-0898308c7d7e · outbound

This paper cites Graph convo- lutional neural network for human action recognition: A comprehensive survey,.

Video Understanding by Design: How Datasets Shape Video Models Graph convo- lutional neural network for human action recognition: A comprehensive survey,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.764684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.764684Z digest=sha256:664406564a7e297f37dc90ca7b8f3b727a82491eaf124b112ce69a20658c88f9

Observation 6fc55875-7abb-4e8e-82f8-0766c7cbc3cc · outbound

This paper cites A survey on deep learning for skeleton-based human animation,.

Video Understanding by Design: How Datasets Shape Video Models A survey on deep learning for skeleton-based human animation,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.767528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.767528Z digest=sha256:8dbe5e6abc3c3242a1ababab8bc51127ee090b5218ded36338eac0bbb501530f

Observation f2b88aff-6e0c-40e3-9551-cae65747091a · outbound

This paper cites A survey on 3d skeleton-based action recognition using learning method,.

Video Understanding by Design: How Datasets Shape Video Models A survey on 3d skeleton-based action recognition using learning method,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.770265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.770265Z digest=sha256:ea37c51f5196448ca7751211b3fbb03bd914c95f51c7de92d0e69b64f5e5ca4e

Observation 6139f1af-29f4-43bf-ba9b-efdd8eff4bad · outbound

This paper cites Self-supervised visual feature learning with deep neural networks: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Self-supervised visual feature learning with deep neural networks: A survey,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.772962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.772962Z digest=sha256:09c47fc316be2eaa35ca6dc50704aced027d81605e4f1209e5f7b325ba57f130

Observation 5fb5a8a2-c48f-44e3-a6a8-51fa303abafd · outbound

This paper cites Self- supervised learning: Generative or contrastive,.

Video Understanding by Design: How Datasets Shape Video Models Self- supervised learning: Generative or contrastive,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.775806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.775806Z digest=sha256:a82267eb7e18e16dea9e36358e699dc0ee91a199a7e717aae2063eedb996c7ff

Observation e0adc5e7-5e0c-4b80-a685-b9b809160857 · outbound

This paper cites Self-supervised representation learning: Introduction, advances, and challenges,.

Video Understanding by Design: How Datasets Shape Video Models Self-supervised representation learning: Introduction, advances, and challenges,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.778740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.778740Z digest=sha256:65559e0f76985476618339e68e2f7468eb6243a6384ce62a278aef698be44a71

Observation d67e78c3-3907-435c-ad8d-98a6ffdf97ae · outbound

This paper cites Deep generative models: Survey,.

Video Understanding by Design: How Datasets Shape Video Models Deep generative models: Survey,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.781602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.781602Z digest=sha256:f43f3c4586fac146f09c3cd70ed7d486aaa62b2c98b75181dfd29757327efc6d

Observation 08df6e10-aa14-4c6d-a590-a69f6797d422 · outbound

This paper cites A survey of multimodal deep generative models,.

Video Understanding by Design: How Datasets Shape Video Models A survey of multimodal deep generative models,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.784280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.784280Z digest=sha256:4ab60bf4a75509a3867ce5284ce7e17c74fa1746de2f08dcee12f02c12dacd16

Observation 51bc42a8-cf82-4859-8e68-c361b0f90c19 · outbound

This paper cites Sora as an agi world model? a complete survey on text-to-video generation,.

Video Understanding by Design: How Datasets Shape Video Models Sora as an agi world model? a complete survey on text-to-video generation,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.787325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.787325Z digest=sha256:96c6c4636546910564b6d71a47cf47f2b97b56455142ae8e2a9ec50f961c6891

Observation 6ff27288-f786-48c8-8049-70c983c38830 · outbound

This paper cites A survey on video diffusion models,.

Video Understanding by Design: How Datasets Shape Video Models A survey on video diffusion models,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.790189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.790189Z digest=sha256:cd0ab278a900b14da23077cae7a287d5a8ba412804e6438c860c2be7828faf82

Observation 348cf097-dd21-4269-8785-def3e3d856c4 · outbound

This paper cites Benchmarking a multimodal and multiview and interactive dataset for human action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Benchmarking a multimodal and multiview and interactive dataset for human action recognition,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.792842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.792842Z digest=sha256:f1a247d34d041f30efe46469fbd21fe01b2d3e65ad8c878ee7ad8f9d3352118c

Observation 7f2df4d5-0dd3-43a9-8481-1f67f0a90918 · outbound

This paper cites Video benchmarks of human action datasets: a review,.

Video Understanding by Design: How Datasets Shape Video Models Video benchmarks of human action datasets: a review,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.796479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.796479Z digest=sha256:60ae3aa9a96b2cff3ede68e31027b1650a5ae4dff3c31dafd33d244b5d0ab829

Observation 76cd93f5-0b11-400a-bf4f-7bdce5c6c16a · outbound

This paper cites A review of convolutional-neural-network- based action recognition,.

Video Understanding by Design: How Datasets Shape Video Models A review of convolutional-neural-network- based action recognition,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.799528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.799528Z digest=sha256:1307686eff4d4ba8726e75f02f35c19d90f2f69a05d4b3c3a97caf54c309591a

Observation a2eb788e-7db7-40e4-9911-5d1d70a88faa · outbound

This paper cites Benchmarking micro-action recognition: Dataset, methods, and applications,.

Video Understanding by Design: How Datasets Shape Video Models Benchmarking micro-action recognition: Dataset, methods, and applications,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.802273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.802273Z digest=sha256:e1e18b7068f61a8e2631c1579d25b62380545776f544517b0a5ad3963d0b93c3

Observation c8df63e0-75e7-40e0-859c-8218a3bcf36f · outbound

This paper cites An outlook into the future of egocentric vision,.

Video Understanding by Design: How Datasets Shape Video Models An outlook into the future of egocentric vision,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.805214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.805214Z digest=sha256:5386c5ba1ab6fe077c6f1d3922b5e26ce0ef4717ba4bea5ebb90e3e4e67502e4

Observation a8f6e50b-f237-40bc-b825-469bc8b99776 · outbound

This paper cites A survey of content-aware video analysis for sports,.

Video Understanding by Design: How Datasets Shape Video Models A survey of content-aware video analysis for sports,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.807877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.807877Z digest=sha256:f28ca348bc52b5b54548df9f3ef1ee9f0a34679e8fbbfbc65aaf816566725336

Observation b21731e8-c9c5-4fa2-9186-d7403273ec4b · outbound

This paper cites Video transcoding: an overview of various techniques and research issues,.

Video Understanding by Design: How Datasets Shape Video Models Video transcoding: an overview of various techniques and research issues,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.810764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.810764Z digest=sha256:c62e8195ab2ef43e7ea070d62959b1a3dc85975ca64910481c4b3b39ee8c9ed2

Observation d3c6fe3b-a3b8-43fc-a3db-87bc48f4fde5 · outbound

This paper cites Video description: A survey of methods, datasets, and evaluation metrics,.

Video Understanding by Design: How Datasets Shape Video Models Video description: A survey of methods, datasets, and evaluation metrics,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.813690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.813690Z digest=sha256:dcb72215a2debaefb5ff7d5fa9cd6636f454bd4769e3ad0e371999d46dfe8a1f

Observation df6379d3-78a0-4078-b712-6fb025e639e8 · outbound

This paper cites Generative multi- view human action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Generative multi- view human action recognition,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.816412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.816412Z digest=sha256:63877c054fc6c32a8e4f0da94da3ef54d608309b1f3a44a3aeafb75fef245c92

Observation 4eeb544d-2cbb-4849-968a-e1c744b4f1d1 · outbound

This paper cites Human action recognition and prediction: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Human action recognition and prediction: A survey,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.819066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.819066Z digest=sha256:10b034742da98ce88210cb325ee9e40fd28a49edfd6fd04ead4e596d5cb0f2d2

Observation 2cb1cf0d-6d3f-45a2-bfae-96ee2912e6d2 · outbound

This paper cites Video generative adversarial networks: a review,.

Video Understanding by Design: How Datasets Shape Video Models Video generative adversarial networks: a review,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.821755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.821755Z digest=sha256:1c0d731612c30cb1998d26781827ff007647772e7dd4322611e822fc0576cffb

Observation 7227d4e1-5845-42db-bd91-084c134eb6f5 · outbound

This paper cites Search-map-search: a frame selection paradigm for action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Search-map-search: a frame selection paradigm for action recognition,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.824426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.824426Z digest=sha256:8a794bea2be1c2c836ca4791a33b7928d76009ec0385d907e0d4c4bdc7933230

Observation 466a9ec3-f511-4ccc-9a82-7bea40bf394a · outbound

This paper cites Self-supervised learning for videos: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Self-supervised learning for videos: A survey,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.827055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.827055Z digest=sha256:8f0866208632d9b2a75f32ac74cc77fab4ee0ce39e562d0e85390b8b8b7d5803

Observation 83d239b4-12f3-4dd2-ad6b-717589ed63cb · outbound

This paper cites On space-time interest points,.

Video Understanding by Design: How Datasets Shape Video Models On space-time interest points,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.829719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.829719Z digest=sha256:2ed34e6e61a7498feb4e738ad00dea21cdad5e46fcef0551fd828332cdb22a19

Observation 7a54a205-b635-4054-a346-79755aa9367a · outbound

This paper cites Histograms of oriented gradients for human detection,.

Video Understanding by Design: How Datasets Shape Video Models Histograms of oriented gradients for human detection,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.832545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.832545Z digest=sha256:d680218d1e8d923c5c68ae2552360ccbb6f5b963f77581a60ca99af7928db8e4

Observation 16a61d7f-0eb9-4fd2-8d50-675aa9e0a063 · outbound

This paper cites Dense trajectories and motion boundary descriptors for action recognition,.

Video Understanding by Design: How Datasets Shape Video Models Dense trajectories and motion boundary descriptors for action recognition,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.835545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.835545Z digest=sha256:814a2c6aafb41fb13fde01be4f79118e1cd48dd26e54d003de0384a3047dd1ad

Observation 232b5f38-071b-48e5-bd29-c93db579818b · outbound

This paper cites Transformers in vision: A survey,.

Video Understanding by Design: How Datasets Shape Video Models Transformers in vision: A survey,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.838250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.838250Z digest=sha256:208fb528ffba61a5a7e44483be06fa14c6b1ac01934d13e964ace7a0c673c0ee

Observation 238925de-89ca-4b9a-8802-0201e4604e44 · outbound

This paper cites Deep Reinforcement Learning: An Overview.

Video Understanding by Design: How Datasets Shape Video Models Deep Reinforcement Learning: An Overview

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.840934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.840934Z digest=sha256:3ed51a142a7322d881e819ef9b58aea667628e250cb213a0886d783b07dd91e3

Observation ce44ac83-9d05-471c-887f-f45921c23157 · outbound

This paper cites Deep reinforcement learning: A brief survey,.

Video Understanding by Design: How Datasets Shape Video Models Deep reinforcement learning: A brief survey,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.843803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.843803Z digest=sha256:9f43a30c7d37339f12f18efbcfbde6251a6e29d73b7a49cf5cc51d6d93057961

Observation dd59018d-85cd-4943-bd36-770cc3f106ce · outbound

This paper cites Continual lifelong learning with neural networks: A review,.

Video Understanding by Design: How Datasets Shape Video Models Continual lifelong learning with neural networks: A review,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.846721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.846721Z digest=sha256:9af20996944a9ca759df70fee1cfc65b64405bcc8ce0a45c83959af12fad0e1a

Observation f13ee776-fed0-471b-9cd0-df84cc97350a · outbound

This paper cites Federated machine learning: Concept and applications,.

Video Understanding by Design: How Datasets Shape Video Models Federated machine learning: Concept and applications,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.849450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.849450Z digest=sha256:f5b4eb5809f0bef70843cfdded32d4e220204c4ccb8860dac95e53d861ca5b70

Pith citing papers

No inbound Pith citation observations are available.