Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:21.811848Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.16701.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:21.811848Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b72c9351-9ef0-465e-b48e-f1ce3130d5c8 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Action genome: Actions as compositions of spatio- temporal scene graphs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33a8dfb9-c426-4331-85c2-628f14f2c224 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition OPT: Open Pre-trained Transformer Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c50662-bf37-4081-9d4d-80586bc74098 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78e5ac41-9a02-42fd-9a94-117b65d18d08 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Chatgpt-4o, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20c30f0a-6199-46e8-b882-0cbeac9c7b31 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Gemini 2.0 flash, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a01dcf5b-733c-4f47-ad8c-b86ba066d818 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Qwen2.5- vl technical report, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 373aa4ee-a0f3-4b4f-af5b-ad4fc7eb3c53 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompting visual-language models for efficient video understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 239205f9-1874-448e-a453-9ab9238b2366 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Llms are good action recognizers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04db58d5-6d67-424b-bbdd-7b8eff7218e1 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9c25835-2ced-4e4d-8a4c-7fafab3cb869 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition M-llm based video frame selec- tion for efficient video understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0db42020-bc5d-4b84-bfe1-195443b6a256 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb149e37-2dad-424f-9e59-0b8fafce6447 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Two-stream con- volutional networks for action recognition in videos
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 259051f0-9365-45ca-b031-1b99d5ba3968 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 284c16f2-3140-4f03-9820-0cd072e8dd07 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning spatiotemporal features with 3d convolutional networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1b3f256-2266-4cea-bdc4-10175834fbc6 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Quo vadis, action recognition? A new model and the kinetics dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 786b4f0b-98ff-44c7-b9f0-b03b7d76232e · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Smeulders
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53831e79-1ec1-4d5f-b53d-b1ce24f67bbb · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Slowfast networks for video recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64d244a2-55d8-420c-8cd6-4fdb5da3b246 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44aab4ff-dcef-4ac7-9517-6380c2b044cc · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Tokenlearner: Adaptive space- time tokenization for videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3ec7672-fe3d-4eb1-97aa-ee2ba973dbd7 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ConceptNet 5.5: An Open Multilingual Graph of General Knowledge
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4e7caf-2a1f-4f04-87ff-69fc0c041152 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e20bce3-48c2-4f9a-a53b-ae7be4ed1c08 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 990f1d3b-7972-4649-bcdb-6f1381f3ae86 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition COMET: Com- monsense transformers for automatic knowledge graph con- struction
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c46a1d6-4e5e-4fe1-b9bc-faeb99f209f8 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb5cccf0-a671-4502-a497-e36f674092a3 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Conditional prompt learning for vision-language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f5971ea-0443-4a1d-95db-f1e2e582866b · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompt distribution learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b6802c2-d0dc-4f3b-a310-e6ab3adf4e31 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning transferable visual models from natural language supervision
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9fe3d8f-922b-4af2-a320-e0060a24cc92 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Denseclip: Language-guided dense prediction with context- aware prompting
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58ea54c6-2266-46b1-8fb9-d96decd84584 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Expanding language-image pretrained models for general video recognition
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e16178ab-908b-460c-8c71-61bebae55ae0 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning to prompt for continual learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e60e9bb0-1d7c-461e-a67e-361fec39616a · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual prompt tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 626eb1b9-c04f-45cf-84c3-95f5616aca36 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Videobert: A joint model for video and language representation learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7615d3c-ec66-474b-b69c-708751fcd5ac · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Language models are few- shot learners
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edc83934-8674-4252-8dd8-09a0bd6a8aa6 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Le, and Christo- pher D
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad77e2ab-4f84-44c4-9a81-51ecbd5c80a6 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Multimodal few-shot learn- ing with frozen language models.Proc
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79406379-5f3b-4823-aa3b-eb5e72351b8e · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition UNIFIEDQA: Crossing format boundaries with a single QA system
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e08c08b5-1488-4201-afdf-32b8fcefbaaf · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Few-shot text generation with natural language instructions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67d8d725-f251-4fa0-acce-d060c304a619 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Generating action-conditioned prompts for open-vocabulary video action recognition
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e59d9d7-4bd7-43b6-9229-22b0bd086f9c · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Kronecker mask and interpretive prompts are language-action video learners
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad2c47ed-e007-499d-bad7-4a94f141ceae · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual semantic role labeling for video understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d9974a3-f7ec-40c7-b23f-fb6a181a81b6 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning Transferable Visual Models From Natural Language Supervision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb55a17d-2ff0-4de0-9d46-7fc584a684b9 · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Flava: A foundational language and vision alignment model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83e83f57-78f2-4534-b8fb-0af7dffdfdea · outbound
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.