Pith. sign in

Paper Citation Record · LEDGER

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.16701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16701 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:21.811848Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b72c9351-9ef0-465e-b48e-f1ce3130d5c8 · outbound

This paper cites Action genome: Actions as compositions of spatio- temporal scene graphs.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Action genome: Actions as compositions of spatio- temporal scene graphs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.520530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:19.034906Z digest=sha256:cf450c3173e2975cae6e44023cd184973feb8bec139470c8a18a3681dbb498e1

Observation 33a8dfb9-c426-4331-85c2-628f14f2c224 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition OPT: Open Pre-trained Transformer Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:19.144627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:19.144627Z digest=sha256:38d138abb7d4effccd3c6013f121d6b45476f31e9f58cc6fd833a1f648ba4856

Observation f8c50662-bf37-4081-9d4d-80586bc74098 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:19.256942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:19.256942Z digest=sha256:82194122756e23d9f87b992885b4eea36a7cad1cf6b2ffce702e3384da709820

Observation 78e5ac41-9a02-42fd-9a94-117b65d18d08 · outbound

This paper cites Chatgpt-4o, 2024.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Chatgpt-4o, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.508520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:19.370286Z digest=sha256:4e5799f55d28310da4388951feaf8e9440bed01e19bd7ad774827b6323fe0a5e

Observation 20c30f0a-6199-46e8-b882-0cbeac9c7b31 · outbound

This paper cites Gemini 2.0 flash, 2024.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Gemini 2.0 flash, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.497191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:19.505819Z digest=sha256:981415a4470ba7be47d5955a2489df59558098a8f03f644a7be4ed31018f7f61

Observation a01dcf5b-733c-4f47-ad8c-b86ba066d818 · outbound

This paper cites Qwen2.5- vl technical report, 2025.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Qwen2.5- vl technical report, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.484668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:19.665627Z digest=sha256:013245c6018b6989fdfdb17f1e31c58399bbf4842ce4b2510303d8aeece5ff0c

Observation 373aa4ee-a0f3-4b4f-af5b-ad4fc7eb3c53 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompting visual-language models for efficient video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.465022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:19.780660Z digest=sha256:5e27627c490c918f2a0fe4eceff1a0cfdf2227ac760220312c4d934c9d9819e4

Observation 239205f9-1874-448e-a453-9ab9238b2366 · outbound

This paper cites Llms are good action recognizers.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Llms are good action recognizers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.451449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:19.851409Z digest=sha256:51b1cb59ffe67212215b0308322a1a1f04f6ac55dea04af988b461c0ad2c5a3e

Observation 04db58d5-6d67-424b-bbdd-7b8eff7218e1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.430934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.001362Z digest=sha256:10f4196d6f1571ebe8dd1452f0f8d49f5dedb277b4059d37a2aa41b0ba2f72ef

Observation d9c25835-2ced-4e4d-8a4c-7fafab3cb869 · outbound

This paper cites M-llm based video frame selec- tion for efficient video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition M-llm based video frame selec- tion for efficient video understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.418534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.132309Z digest=sha256:0e6d9671f893686b1f5e8c57ef6ac23e201e9143217637e4f308bed1eb0fb5c4

Observation 0db42020-bc5d-4b84-bfe1-195443b6a256 · outbound

This paper cites HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:20.273161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:20.273161Z digest=sha256:e4e767990268bc1d8aa9c1c5e096afb1a904c2be5c8505d34f627c57fe79632f

Observation cb149e37-2dad-424f-9e59-0b8fafce6447 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Two-stream con- volutional networks for action recognition in videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.385943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.354728Z digest=sha256:1aea24da77c55e23bfa8b32eca5af629ad8880ff3437e56d0ff54cc01652ad56

Observation 259051f0-9365-45ca-b031-1b99d5ba3968 · outbound

This paper cites Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.374140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.427751Z digest=sha256:2c562dce876af62f6b00462f98bd694abff284459d9d70fdb8ee1ba0a1331037

Observation 284c16f2-3140-4f03-9820-0cd072e8dd07 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning spatiotemporal features with 3d convolutional networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.360639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.486930Z digest=sha256:e66e76cd8c2f479f17a03685f3b678072430a1b536b20460d9c6a71750ba9119

Observation b1b3f256-2266-4cea-bdc4-10175834fbc6 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Quo vadis, action recognition? A new model and the kinetics dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.347775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.570357Z digest=sha256:d4ad132a39b3b4c996b3fe9b7d3ec2f95a5119a7ac5294703e98dd7a3978106c

Observation 786b4f0b-98ff-44c7-b9f0-b03b7d76232e · outbound

This paper cites Smeulders.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Smeulders

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.335524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.634188Z digest=sha256:8bc54630c58d5be95db8c0b374fe5b3fd53713cd849a703ae30843872645d3f6

Observation 53831e79-1ec1-4d5f-b53d-b1ce24f67bbb · outbound

This paper cites Slowfast networks for video recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Slowfast networks for video recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.322239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.732330Z digest=sha256:32a84465eff65eb3fee7c943fd9daadce3d4a9825505dbd4e78b62a2344391e3

Observation 64d244a2-55d8-420c-8cd6-4fdb5da3b246 · outbound

This paper cites an unresolved cited work.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:22.308568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.811147Z digest=sha256:79cffb9680b5a6796b7aadb8f2d8c4ef5126d5839aa68db521e328d06e3e397a

Observation 44aab4ff-dcef-4ac7-9517-6380c2b044cc · outbound

This paper cites Tokenlearner: Adaptive space- time tokenization for videos.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Tokenlearner: Adaptive space- time tokenization for videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.293076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:20.821125Z digest=sha256:ee3ea958a6c531e6c752f6f079b5436f8242d137bb2665c6e0f4a1dd714a1d6b

Observation c3ec7672-fe3d-4eb1-97aa-ee2ba973dbd7 · outbound

This paper cites ConceptNet 5.5: An Open Multilingual Graph of General Knowledge.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ConceptNet 5.5: An Open Multilingual Graph of General Knowledge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:20.983666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:20.983666Z digest=sha256:62ca6cf033b1bd042367792d46064a0d82be3e4df3ce91e150803f546d241884

Observation 0a4e7caf-2a1f-4f04-87ff-69fc0c041152 · outbound

This paper cites ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.078067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.078067Z digest=sha256:a6772a03f0e2bf8d4b9faeab5df36e7e4e183c151b4c9337c113a7fac9014f5c

Observation 6e20bce3-48c2-4f9a-a53b-ae7be4ed1c08 · outbound

This paper cites We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.271807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.192655Z digest=sha256:731272a8a6327de526c0162c859a662025ba8f4e53158381d46d5e69ad02ec86

Observation 990f1d3b-7972-4649-bcdb-6f1381f3ae86 · outbound

This paper cites COMET: Com- monsense transformers for automatic knowledge graph con- struction.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition COMET: Com- monsense transformers for automatic knowledge graph con- struction

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.233158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.287096Z digest=sha256:f3f6be616c3f53d0ff7d3fbbd1680ddc0cf881a6fa28500d24eac415d5535a20

Observation 2c46a1d6-4e5e-4fe1-b9bc-faeb99f209f8 · outbound

This paper cites Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.213346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.306892Z digest=sha256:17368a0ca29aeb903aa3b83a8f062192a3907bf93243742571499d42e4670364

Observation eb5cccf0-a671-4502-a497-e36f674092a3 · outbound

This paper cites Conditional prompt learning for vision-language models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Conditional prompt learning for vision-language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.194019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.444572Z digest=sha256:674a75243b6020a6e0baf7a17a725e54758449595a03b084d66c1818522899e5

Observation 6f5971ea-0443-4a1d-95db-f1e2e582866b · outbound

This paper cites Prompt distribution learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompt distribution learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.173041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.544753Z digest=sha256:06075821b31690d41c505fce05100d5602ec0291917ca3f93d740071e275517c

Observation 7b6802c2-d0dc-4f3b-a310-e6ab3adf4e31 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.574755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.574755Z digest=sha256:17060f3d537dc3c862eb436f17a1dfa039a1b182a1b9c166ed2ac6beafbe593c

Observation b9fe3d8f-922b-4af2-a320-e0060a24cc92 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Denseclip: Language-guided dense prediction with context- aware prompting

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.151063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.633977Z digest=sha256:c5d5e7be635cd9a53ce4b3f02f94e478f7649c6e8a5c8175fdc7e717a5d00e0b

Observation 58ea54c6-2266-46b1-8fb9-d96decd84584 · outbound

This paper cites Expanding language-image pretrained models for general video recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Expanding language-image pretrained models for general video recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.137255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.742005Z digest=sha256:5dc5d0ca2bb09ca62b68bacec1a6a2a1cd662580ee07aa3c8c32a20d3cef5f60

Observation e16178ab-908b-460c-8c71-61bebae55ae0 · outbound

This paper cites Learning to prompt for continual learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning to prompt for continual learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.122393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.745915Z digest=sha256:2db07f5e8fc48814f69979fe774c2ed412521e5d7dbd2f8d462017c01c9f9865

Observation e60e9bb0-1d7c-461e-a67e-361fec39616a · outbound

This paper cites Visual prompt tuning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual prompt tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.109828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.749643Z digest=sha256:d0c28fdd181ef9cdbac3659081fb6f430c2509040aa521b7120fdd8b7a22e58c

Observation 626eb1b9-c04f-45cf-84c3-95f5616aca36 · outbound

This paper cites Videobert: A joint model for video and language representation learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Videobert: A joint model for video and language representation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.096230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.756955Z digest=sha256:86bfa486f40f71737a447151226ab34265a25c11e5117307df23149d5aecaa29

Observation f7615d3c-ec66-474b-b69c-708751fcd5ac · outbound

This paper cites Language models are few- shot learners.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Language models are few- shot learners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.075878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.763188Z digest=sha256:2612a6f29e60b72c2006ed9e28bc57d0b11450890ff657c97c1338c24acab923

Observation edc83934-8674-4252-8dd8-09a0bd6a8aa6 · outbound

This paper cites Le, and Christo- pher D.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Le, and Christo- pher D

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.054885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.768055Z digest=sha256:32e03b015f84939ba692e58a50d80d0e13bf0153af9d4275a6cd750cec66597b

Observation ad77e2ab-4f84-44c4-9a81-51ecbd5c80a6 · outbound

This paper cites Multimodal few-shot learn- ing with frozen language models.Proc.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Multimodal few-shot learn- ing with frozen language models.Proc

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.042786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.772487Z digest=sha256:31d47f17ddf6cd393106f5ce31eb2c65aff489ffebce8328ade22caf17774d9b

Observation 79406379-5f3b-4823-aa3b-eb5e72351b8e · outbound

This paper cites UNIFIEDQA: Crossing format boundaries with a single QA system.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition UNIFIEDQA: Crossing format boundaries with a single QA system

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.019630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.777249Z digest=sha256:8ab82c6af123503441dca113d1ce99c42a72aefa56e070c83dd059caf4520a27

Observation e08c08b5-1488-4201-afdf-32b8fcefbaaf · outbound

This paper cites Few-shot text generation with natural language instructions.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Few-shot text generation with natural language instructions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.005297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.781863Z digest=sha256:285bb13259acf8b97d1e8f11d8b0b7f3ecd838690a1027c4ffe462b4988c55e5

Observation 67d8d725-f251-4fa0-acce-d060c304a619 · outbound

This paper cites Generating action-conditioned prompts for open-vocabulary video action recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Generating action-conditioned prompts for open-vocabulary video action recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.987818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.786252Z digest=sha256:fed29b3fa1130289cef855c2b3ab1451b2da581260cb383cd78fb15b1489e122

Observation 2e59d9d7-4bd7-43b6-9229-22b0bd086f9c · outbound

This paper cites Kronecker mask and interpretive prompts are language-action video learners.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Kronecker mask and interpretive prompts are language-action video learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.973922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.790912Z digest=sha256:313e7ead5d5b11864f8cd549188e17fa714fa24a609ce85b9d335bb6e5efddcb

Observation ad2c47ed-e007-499d-bad7-4a94f141ceae · outbound

This paper cites Visual semantic role labeling for video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual semantic role labeling for video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.960619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.794614Z digest=sha256:997f637e29f9b8b92088a2b258a19fa7ec1dc3eb41262499d71c73b71f573220

Observation 8d9974a3-f7ec-40c7-b23f-fb6a181a81b6 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning Transferable Visual Models From Natural Language Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.799058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.799058Z digest=sha256:9855bf9bc1d0533ec0df12243ac84404823b674284264b38321ce731aad49e12

Observation cb55a17d-2ff0-4de0-9d46-7fc584a684b9 · outbound

This paper cites Flava: A foundational language and vision alignment model.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Flava: A foundational language and vision alignment model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.947281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:42:21.807905Z digest=sha256:6c89c536a331ece84003f7a0af2212d46bbea42897bb874d0c314d58acb810e1

Observation 83e83f57-78f2-4534-b8fb-0af7dffdfdea · outbound

This paper cites Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.811848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.811848Z digest=sha256:39eee6f8f0111dcd906dbf6ccf1488f7ed80cbec169903a5ca85ba40e1bb06cc

Pith citing papers

No inbound Pith citation observations are available.