Pith. sign in

Paper Citation Record · LEDGER

VidCtx: Context-aware Video Question Answering with Image Models

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2412.17415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17415 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.391770Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.808179Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36b59a45-b65a-4919-a5fd-3d4d50acf58b · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.251138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.251138Z digest=sha256:383e42dbf4994ccf876fa563c9703c8653165f8b9930257cb02e55b93f4f15e6

Observation a0b15ddd-9712-425a-8d9b-836f8593e592 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.256797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.256797Z digest=sha256:3cdfd5cae3522a537f37fc1082fc1af534c6590a8b9cb70fd083288a5903d3da

Observation d862ea53-59ad-4fec-a7e4-5741daf06e7c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VidCtx: Context-aware Video Question Answering with Image Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.261821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.261821Z digest=sha256:91c0752736f145079904f681938afa464ed19c4ddca31bf085befde9fa8e0160

Observation 976db303-747f-4c80-bb8d-17d02eaadee7 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VidCtx: Context-aware Video Question Answering with Image Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.267504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.267504Z digest=sha256:3cf392d2c1b72641f30e2b51167dfd23df144047a0da0529c3f4fd7322f1be22

Observation 2cec20e0-c2b2-46ca-926a-2cd54c61a4e0 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding bench- mark,.

VidCtx: Context-aware Video Question Answering with Image Models Mvbench: A comprehensive multi-modal video understanding bench- mark,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.860581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.272473Z digest=sha256:c243133af287d0bb4e83ddbc84064c42d9888bbdb4d38968a33d369fa6a2eb80

Observation 34d2f7df-a516-4953-bf2c-2d7e264b4e14 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VidCtx: Context-aware Video Question Answering with Image Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.276954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.276954Z digest=sha256:4a1d5175917a291c8442115da6a08bd9891899438a4899ef87f7d80a483f1041

Observation 7f5ca850-f005-485f-bcec-d2d08ccae1fd · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.282682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.282682Z digest=sha256:2120d98190b97eacf4158d3199269ca66a3f4ef1b683ae41eb7795e94315e7f2

Observation b73ffffa-6723-4754-8183-1b97d3d9654e · outbound

This paper cites Self- chained image-language model for video localization and question answering,.

VidCtx: Context-aware Video Question Answering with Image Models Self- chained image-language model for video localization and question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.846020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.287475Z digest=sha256:7087664a6e9ab46b3005e8487c0b9a212c12195828e3dcf380849d96c9a96563

Observation e782af15-4a88-435d-b6b3-b7598ae1a837 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

VidCtx: Context-aware Video Question Answering with Image Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.291907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.291907Z digest=sha256:0c2c9b632381eb3eda0bf09c9eedbef33ef8345f161774fca08213958fa136ab

Observation 92a4f946-5951-4242-9af4-05643838b515 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent,.

VidCtx: Context-aware Video Question Answering with Image Models Videoagent: Long-form video understanding with large language model as agent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.832421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.296639Z digest=sha256:f8c51afce75194261ef6a4375c6181b86b6d64e3e60b4a6ea7ae4c344d4d4abe

Observation 779c2b7a-300b-454d-b289-4758696ef7eb · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

VidCtx: Context-aware Video Question Answering with Image Models A Simple LLM Framework for Long-Range Video Question-Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.302286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.302286Z digest=sha256:e425899e6df7434e6b63e648b44785c29871aa99b2276bfec32cb195042e6a17

Observation fb9a8cb9-4666-4f03-af54-b49098b2ee9b · outbound

This paper cites Language Repository for Long Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Language Repository for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.307324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.307324Z digest=sha256:910313a5690e82e50b3aad86cbd91b48e10de043e52726dfcdead669cff47ba6

Observation daf1d424-5f21-4a3c-8b3d-d98dfa29eb43 · outbound

This paper cites Question-instructed visual de- scriptions for zero-shot video answering,.

VidCtx: Context-aware Video Question Answering with Image Models Question-instructed visual de- scriptions for zero-shot video answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.816975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.313074Z digest=sha256:e8316ab87fecc8a1ca2c07fe7d0230056e4e6827be6b3bd2b330f829e8ee79a5

Observation 84ac54d0-55df-4757-a43e-2305ddb96686 · outbound

This paper cites Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.317436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.317436Z digest=sha256:78bee879cbb6ed54f4beea5ee44933d71d5b4b6d1e04f743612e5a0583388718

Observation 6736c950-611b-4dab-ad9e-ced58d2443c2 · outbound

This paper cites Large language models can be easily distracted by irrelevant context,.

VidCtx: Context-aware Video Question Answering with Image Models Large language models can be easily distracted by irrelevant context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.802738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.322333Z digest=sha256:02e8f6fa60b41c9889f8841ab37209b64cad336c24a74ecb13144d52f20fa1ef

Observation db2d6455-f26a-4ea6-bc2a-23a4cd9236ee · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

VidCtx: Context-aware Video Question Answering with Image Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.787037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.326968Z digest=sha256:fa7899b91fd7f9f08767a4de2c26d5bd16b2cfedb2ddf736692f6c67321276f8

Observation 5e862572-2f5e-4d59-b473-7d4514489060 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.331502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.331502Z digest=sha256:b6e7f4bbfa540cc85c8b746bb58ff49bd383e5d298cc56ca66ea579de84ceaaf

Observation f1290a7d-a931-433b-a740-3d6bc81295ca · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,.

VidCtx: Context-aware Video Question Answering with Image Models Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.771872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.337272Z digest=sha256:6d270c553c3897a4255625a71db61fe7b4c5531234f5bb75f3eb411c08de1c18

Observation c3947363-3c7c-4dd3-b60e-3dc3edf60c6b · outbound

This paper cites Enhancing multimodal sentiment analysis via learning from large language model,.

VidCtx: Context-aware Video Question Answering with Image Models Enhancing multimodal sentiment analysis via learning from large language model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.756953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.341609Z digest=sha256:cbb8484b37b891b4914ea4f0dcb16c717884991ec81703eeaae0bb2ce8a2bd81

Observation bd9a91da-8403-4e3e-90cb-d3d9c73af29a · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition,.

VidCtx: Context-aware Video Question Answering with Image Models Video-of-thought: Step-by-step video reasoning from perception to cognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.742345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.346135Z digest=sha256:d523e423fa066289a64efcd8bc8e339c23ec0d82a596297ccc0e8c2270811ffd

Observation 3e5aa05f-549f-4f40-9081-d470ee0498b7 · outbound

This paper cites Vamos: Versatile Action Models for Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Vamos: Versatile Action Models for Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.350899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.350899Z digest=sha256:a7af09bc20ad0faaba2754c0a3b7bf9810a6dc6089b118f213bbd86adb73667d

Observation 81f5fac8-7b5a-489e-bab4-c6a35ee3a06c · outbound

This paper cites Large language models are zero-shot reasoners,.

VidCtx: Context-aware Video Question Answering with Image Models Large language models are zero-shot reasoners,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.727217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.356308Z digest=sha256:e36bc1df59230d14146b4bee1d8662bc07307472efb9a93b5969fa4fd3e7ea1d

Observation 7e21f624-5f9d-4657-9b4b-40809b13690e · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

VidCtx: Context-aware Video Question Answering with Image Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.710596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.360643Z digest=sha256:6c3cc30a61a8a05c5806683a25c824e36637a191d6d6a80b439c927cb646565d

Observation 09818ffb-dafd-4aeb-a383-1fd0d3bc1f0d · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal mod- els,.

VidCtx: Context-aware Video Question Answering with Image Models Compositional chain-of-thought prompting for large multimodal mod- els,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.691893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.364961Z digest=sha256:c6470cb79e5ecdc120048d9bd36ac9db7eb27ce34502f4fdb503de728b63c2c8

Observation fa23c586-feed-418c-91dd-c826e7f0678e · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

VidCtx: Context-aware Video Question Answering with Image Models Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.676088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.369044Z digest=sha256:eb2178fe3a7f25943b842d94820a524b852a37e90308305f6eb81a34947fbbab

Observation ed5e1194-f9b4-4603-a465-a88004e6276a · outbound

This paper cites Intentqa: Context- aware video intent reasoning,.

VidCtx: Context-aware Video Question Answering with Image Models Intentqa: Context- aware video intent reasoning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.660083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.374146Z digest=sha256:fcdc6ff608ca87c5698da640d88555bd11b4ebb07b160f56fd36bfb931d4b5a7

Observation 1dc03d09-aa3f-4fc7-b7d6-a78b825ce3d0 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

VidCtx: Context-aware Video Question Answering with Image Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.378423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.378423Z digest=sha256:7a562981ceaa190eeb19ca5cbcf377e1627bf9712841f29f30b8f36194eaca24

Observation 29a74aa0-e440-4351-acaa-3a00735895d7 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models,.

VidCtx: Context-aware Video Question Answering with Image Models Verbs in action: Improving verb understanding in video-language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.644900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.382519Z digest=sha256:c263a7ecc3db61a4882fc1a56d243f9325067f1d9f34c889154f362ed67876bf

Observation 527a5034-1726-4fac-984e-db07468774e6 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.387257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.387257Z digest=sha256:1d06201223c107d3c58371dd92cf8889b7bf819d8c931db52463484d4099034d

Observation 199033f3-a787-4436-aa41-9da7e36fc172 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

VidCtx: Context-aware Video Question Answering with Image Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.628488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:29:51.391770Z digest=sha256:93285b4bee4e175f044fac1a222db00a912a81ce1e44cd5732ea7aea2673d190

Pith citing papers

Observation 4dccfb87-3f24-4aae-bd55-8b45679640d5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning VidCtx: Context-aware Video Question Answering with Image Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.809467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:ddad99a5bc8b6a9436da20b69292b9c68ed7961d3fc42c6159330f354864b277