Pith. sign in

Paper Citation Record · LEDGER

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.16701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16701 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:21.811848Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b72c9351-9ef0-465e-b48e-f1ce3130d5c8 · outbound

This paper cites Action genome: Actions as compositions of spatio- temporal scene graphs.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Action genome: Actions as compositions of spatio- temporal scene graphs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.520530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:19.034906Z digest=sha256:e726deaae1c26e06335fab974f44f089fe80479c45d11f0c344a34354cee9a1c

Observation 33a8dfb9-c426-4331-85c2-628f14f2c224 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition OPT: Open Pre-trained Transformer Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:19.144627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:19.144627Z digest=sha256:38d138abb7d4effccd3c6013f121d6b45476f31e9f58cc6fd833a1f648ba4856

Observation f8c50662-bf37-4081-9d4d-80586bc74098 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:19.256942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:19.256942Z digest=sha256:82194122756e23d9f87b992885b4eea36a7cad1cf6b2ffce702e3384da709820

Observation 78e5ac41-9a02-42fd-9a94-117b65d18d08 · outbound

This paper cites Chatgpt-4o, 2024.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Chatgpt-4o, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.508520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:19.370286Z digest=sha256:acf176d8ea8f60b6bf1ae74d6236e7e6cb009c7b489d84bf47134d6d2f39255e

Observation 20c30f0a-6199-46e8-b882-0cbeac9c7b31 · outbound

This paper cites Gemini 2.0 flash, 2024.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Gemini 2.0 flash, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.497191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:19.505819Z digest=sha256:fb70abea06cb2e5e49c1224699e2e9dafec165ff227e89c693eb8b5df02de56e

Observation a01dcf5b-733c-4f47-ad8c-b86ba066d818 · outbound

This paper cites Qwen2.5- vl technical report, 2025.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Qwen2.5- vl technical report, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.484668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:19.665627Z digest=sha256:c610cf8134198b308f0906f6ed8422a897f089380084f52c5e63ad8dc8f3c5ff

Observation 373aa4ee-a0f3-4b4f-af5b-ad4fc7eb3c53 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompting visual-language models for efficient video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.465022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:19.780660Z digest=sha256:0fc48e24ad5d2db5505f26f8d5bbef70013e8535ba057427ed561daf70dbe3f7

Observation 239205f9-1874-448e-a453-9ab9238b2366 · outbound

This paper cites Llms are good action recognizers.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Llms are good action recognizers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.451449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:19.851409Z digest=sha256:bd37c04c8f61e823a9befc8f480c3c3939ef470483e41e57a1d46781c32f2a9e

Observation 04db58d5-6d67-424b-bbdd-7b8eff7218e1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.430934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.001362Z digest=sha256:56102ebfb74e5540c41700d8c6b11316ead70d85ac54023a039a07780c61bdd6

Observation d9c25835-2ced-4e4d-8a4c-7fafab3cb869 · outbound

This paper cites M-llm based video frame selec- tion for efficient video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition M-llm based video frame selec- tion for efficient video understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.418534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.132309Z digest=sha256:3ee9112d24255747bfa7399f30e8aa6cfa88c143437673a413694f5f58d3dd95

Observation 0db42020-bc5d-4b84-bfe1-195443b6a256 · outbound

This paper cites HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:20.273161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:20.273161Z digest=sha256:e4e767990268bc1d8aa9c1c5e096afb1a904c2be5c8505d34f627c57fe79632f

Observation cb149e37-2dad-424f-9e59-0b8fafce6447 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Two-stream con- volutional networks for action recognition in videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.385943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.354728Z digest=sha256:f09000145913ad572c8d6b9c692fb9bd45d57ea1fe092f01a6875162e590389c

Observation 259051f0-9365-45ca-b031-1b99d5ba3968 · outbound

This paper cites Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.374140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.427751Z digest=sha256:b5005473ec07a5774fa8f53f457dda4f7bf64d9c18ef48087f1dd491a407d4b9

Observation 284c16f2-3140-4f03-9820-0cd072e8dd07 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning spatiotemporal features with 3d convolutional networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.360639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.486930Z digest=sha256:9e5f8d129b21c8d8f2c1042e4b3b15fdae319e78fda1fdd9784963db641f9510

Observation b1b3f256-2266-4cea-bdc4-10175834fbc6 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Quo vadis, action recognition? A new model and the kinetics dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.347775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.570357Z digest=sha256:6f7536bbc607753fa298105ca70fd02635114e05a6db94815ba688af254bf998

Observation 786b4f0b-98ff-44c7-b9f0-b03b7d76232e · outbound

This paper cites Smeulders.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Smeulders

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.335524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.634188Z digest=sha256:c5a20a77737c07aaea085f989a830efb47f5d084335ae405fbe28ca2a124e8f8

Observation 53831e79-1ec1-4d5f-b53d-b1ce24f67bbb · outbound

This paper cites Slowfast networks for video recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Slowfast networks for video recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.322239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.732330Z digest=sha256:fee9deefa570969195c925ea3df2a795d534c0cc0bd3b2367c260f8ef9658372

Observation 64d244a2-55d8-420c-8cd6-4fdb5da3b246 · outbound

This paper cites an unresolved cited work.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:22.308568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.811147Z digest=sha256:9e25c1ec66c927389d7e734be02aae1507431fda83706df740a79e4b0c6dfd32

Observation 44aab4ff-dcef-4ac7-9517-6380c2b044cc · outbound

This paper cites Tokenlearner: Adaptive space- time tokenization for videos.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Tokenlearner: Adaptive space- time tokenization for videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.293076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:20.821125Z digest=sha256:e0d59604c2ee84d018f2e675f545c6374f85fcb0f9585e70b82c3d0609f968f4

Observation c3ec7672-fe3d-4eb1-97aa-ee2ba973dbd7 · outbound

This paper cites ConceptNet 5.5: An Open Multilingual Graph of General Knowledge.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ConceptNet 5.5: An Open Multilingual Graph of General Knowledge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:20.983666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:20.983666Z digest=sha256:62ca6cf033b1bd042367792d46064a0d82be3e4df3ce91e150803f546d241884

Observation 0a4e7caf-2a1f-4f04-87ff-69fc0c041152 · outbound

This paper cites ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.078067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.078067Z digest=sha256:a6772a03f0e2bf8d4b9faeab5df36e7e4e183c151b4c9337c113a7fac9014f5c

Observation 6e20bce3-48c2-4f9a-a53b-ae7be4ed1c08 · outbound

This paper cites We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.271807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.192655Z digest=sha256:718aa236a51fe17a03393814344339f761d84dc2f759026ff354308d671591da

Observation 990f1d3b-7972-4649-bcdb-6f1381f3ae86 · outbound

This paper cites COMET: Com- monsense transformers for automatic knowledge graph con- struction.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition COMET: Com- monsense transformers for automatic knowledge graph con- struction

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.233158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.287096Z digest=sha256:082d4ef04bd2e07287f1cb472246e2e14ef001e54c3a6899627a942942703c9b

Observation 2c46a1d6-4e5e-4fe1-b9bc-faeb99f209f8 · outbound

This paper cites Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.213346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.306892Z digest=sha256:44e415dd20f9bb4f56d3f858111069c5f2ca117e9978d73a3789c76d45a91070

Observation eb5cccf0-a671-4502-a497-e36f674092a3 · outbound

This paper cites Conditional prompt learning for vision-language models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Conditional prompt learning for vision-language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.194019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.444572Z digest=sha256:1bbc705910e3bf17c33b4e33ce2612ef2d9434630af2ba985a8896f6a4e49770

Observation 6f5971ea-0443-4a1d-95db-f1e2e582866b · outbound

This paper cites Prompt distribution learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompt distribution learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.173041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.544753Z digest=sha256:c7297c808ac9fce7aec0807c5b9d9e6f2f20e5f8118f67bb5e6822362be2860f

Observation 7b6802c2-d0dc-4f3b-a310-e6ab3adf4e31 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.574755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.574755Z digest=sha256:17060f3d537dc3c862eb436f17a1dfa039a1b182a1b9c166ed2ac6beafbe593c

Observation b9fe3d8f-922b-4af2-a320-e0060a24cc92 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Denseclip: Language-guided dense prediction with context- aware prompting

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.151063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.633977Z digest=sha256:521af075f7531b836699337e170fc60485b5d49d33055b88d7f9db0a3d765d0d

Observation 58ea54c6-2266-46b1-8fb9-d96decd84584 · outbound

This paper cites Expanding language-image pretrained models for general video recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Expanding language-image pretrained models for general video recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.137255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.742005Z digest=sha256:df43d16902577ed5c50d82231ef302cf05eccec006094067abe453317a6be4ff

Observation e16178ab-908b-460c-8c71-61bebae55ae0 · outbound

This paper cites Learning to prompt for continual learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning to prompt for continual learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.122393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.745915Z digest=sha256:b8848b1679ffd899cf8ecd157d2b13aa5edbb6893ec9519e9aaf99c770d4c735

Observation e60e9bb0-1d7c-461e-a67e-361fec39616a · outbound

This paper cites Visual prompt tuning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual prompt tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.109828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.749643Z digest=sha256:6ba5fbce31f2ed5a9d94349286b9e7545787822891eae6bf2491a8c526066b9c

Observation 626eb1b9-c04f-45cf-84c3-95f5616aca36 · outbound

This paper cites Videobert: A joint model for video and language representation learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Videobert: A joint model for video and language representation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.096230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.756955Z digest=sha256:1c42e0378309707333aa53111d4854078fb1d28fb2a1e1b53e9a6fe32b020cf4

Observation f7615d3c-ec66-474b-b69c-708751fcd5ac · outbound

This paper cites Language models are few- shot learners.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Language models are few- shot learners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.075878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.763188Z digest=sha256:6d4f871d8fd55157b084c538d8b329e562c0ea8ba13529775cf410b8214b9916

Observation edc83934-8674-4252-8dd8-09a0bd6a8aa6 · outbound

This paper cites Le, and Christo- pher D.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Le, and Christo- pher D

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.054885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.768055Z digest=sha256:13a4e26fa0a3a9d4d1456620ef0b4b2a8d6250231b78693d257d489b39c9f606

Observation ad77e2ab-4f84-44c4-9a81-51ecbd5c80a6 · outbound

This paper cites Multimodal few-shot learn- ing with frozen language models.Proc.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Multimodal few-shot learn- ing with frozen language models.Proc

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.042786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.772487Z digest=sha256:319a89a4b166c187c504288307639ee54a8d73c83095eb169ada24a2546f200a

Observation 79406379-5f3b-4823-aa3b-eb5e72351b8e · outbound

This paper cites UNIFIEDQA: Crossing format boundaries with a single QA system.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition UNIFIEDQA: Crossing format boundaries with a single QA system

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.019630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.777249Z digest=sha256:6d32ff12692757751ceab50945a135e86296b340b991259727ca40a00524d768

Observation e08c08b5-1488-4201-afdf-32b8fcefbaaf · outbound

This paper cites Few-shot text generation with natural language instructions.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Few-shot text generation with natural language instructions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.005297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.781863Z digest=sha256:2d464557b1a7ef6382d6cba24848054802ac4ed3ebc41c48d562918834cb51ec

Observation 67d8d725-f251-4fa0-acce-d060c304a619 · outbound

This paper cites Generating action-conditioned prompts for open-vocabulary video action recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Generating action-conditioned prompts for open-vocabulary video action recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.987818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.786252Z digest=sha256:5420373c0700e6b3e82e92ea7404fed99d3c9ea5fb42a8ecdfe7684a1a42d13a

Observation 2e59d9d7-4bd7-43b6-9229-22b0bd086f9c · outbound

This paper cites Kronecker mask and interpretive prompts are language-action video learners.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Kronecker mask and interpretive prompts are language-action video learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.973922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.790912Z digest=sha256:405ead31bc04ace4f137c822c1f069889007b0c6c9235d0c8a4c77003b1916e9

Observation ad2c47ed-e007-499d-bad7-4a94f141ceae · outbound

This paper cites Visual semantic role labeling for video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual semantic role labeling for video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.960619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.794614Z digest=sha256:6d29c686942d118e826c102ad80a2861407c3674a8013400252cf7bd54529db5

Observation 8d9974a3-f7ec-40c7-b23f-fb6a181a81b6 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning Transferable Visual Models From Natural Language Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.799058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.799058Z digest=sha256:9855bf9bc1d0533ec0df12243ac84404823b674284264b38321ce731aad49e12

Observation cb55a17d-2ff0-4de0-9d46-7fc584a684b9 · outbound

This paper cites Flava: A foundational language and vision alignment model.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Flava: A foundational language and vision alignment model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.947281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:42:21.807905Z digest=sha256:4284e37316468ef2b5afa29e2cfc539e46af816aa52a7fcd24a3cdb3a138d34e

Observation 83e83f57-78f2-4534-b8fb-0af7dffdfdea · outbound

This paper cites Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.811848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.811848Z digest=sha256:39eee6f8f0111dcd906dbf6ccf1488f7ed80cbec169903a5ca85ba40e1bb06cc

Pith citing papers

No inbound Pith citation observations are available.