Pith. sign in

Paper Citation Record · LEDGER

Language Repository for Long Video Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.14622.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.14622 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.390322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:03.032393Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 203814b7-9f3a-4acc-9ebb-9953a9e5bd9f · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Language Repository for Long Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.569002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:b79cf0ae71ab924a9e9b7efb3a0157c757e27e8a26351f6f42472238c96fdfc3

Observation 129af2f7-9a6b-4b7e-bc82-8b43005c2458 · inbound

On the Consistency of Video Large Language Models in Temporal Comprehension cites this paper.

On the Consistency of Video Large Language Models in Temporal Comprehension Language Repository for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.731830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.731830Z digest=sha256:4566cfa319282a523923a5db8324ca752cbfd079a1a5e1a740d6d9320b63b40b

Observation 776b7648-a119-449b-8d8f-2d49fa3dc781 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Repository for Long Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.181876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.181876Z digest=sha256:7529d9138935f9a15611c82e01992f93ddfd0f49c9b7e1429e7d48491879606c

Observation e49bfb55-3144-4a2f-b33b-59b8d134b026 · inbound

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering cites this paper.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Language Repository for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.143816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.143816Z digest=sha256:b1498683de854e8fa39a8089dabfb8289fc4a04a672fd434e74675030d53e1a6

Observation fb9a8cb9-4666-4f03-af54-b49098b2ee9b · inbound

VidCtx: Context-aware Video Question Answering with Image Models cites this paper.

VidCtx: Context-aware Video Question Answering with Image Models Language Repository for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.307324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.307324Z digest=sha256:6de6c984bd71903aad1721ba3f8229f0f41ea520077cd12421e6fae416c1e04c

Observation 6a1416b0-dba8-4eb8-a22a-a93777d61210 · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment Language Repository for Long Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.182626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.182626Z digest=sha256:cc0b8907345cc7708097a7b26d5be5910077bbe9e051b56858feb7c711d84ffa

Observation bc972c75-5884-407d-9b57-524c466e2845 · inbound

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process cites this paper.

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process Language Repository for Long Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:56:37.533803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:56:37.533803Z digest=sha256:633ee901837e7c90e98befe342b1e4c3051c8fed7b097654f1979520ce911bd6

Observation 690f6774-ef24-4d7d-a30c-aeba5e17c8c8 · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding Language Repository for Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.390322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.390322Z digest=sha256:0083fdbecd50e602ab738f07b9920a1d4a13fc77983f320cb20b1edaf36d08e3

Observation a9fd451f-0f61-4df7-9b37-e04adb55aa37 · inbound

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning cites this paper.

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning Language Repository for Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:26:44.378060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:26:44.378060Z digest=sha256:814a2fae8a43087bbd63a575e1d715056a63f687d6dbcb30a38fd293d4798290

Observation 0731a994-2002-4fc0-9026-e61e684b35e5 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Language Repository for Long Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.744283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.744283Z digest=sha256:d64e7574289a008d0d49d93b0c8ff1f3caee6792608e6e101551186c6a2cb1dd

Observation 378dba24-81bd-4297-9f4f-974a00cd3418 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames Language Repository for Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.647384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.647384Z digest=sha256:11208aa539e875c054f503de7a0bc089a8df17aaf881ad7c7560786af53044aa

Observation c5da05aa-da0b-4824-b946-801df598799f · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Language Repository for Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:29.804755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:29.804755Z digest=sha256:56b8cc84c6584e38726413b1ea62309f265d8be24285114041fa20ff1a433d8d

Observation 02bd78e0-fb21-4a3a-aa91-739025b936dc · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Language Repository for Long Video Understanding

Reference 300

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:12.253371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:12.253371Z digest=sha256:f771512f2847a14b92d1c692857fdac240b786261974625a466bb7ac6c739ecd

Observation 93e3cbc7-3329-48b7-8fe3-9b30deb89d17 · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning Language Repository for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:10.793495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:10.793495Z digest=sha256:b8f4c7193fabdb645b8ba69e0e1ff56fc05eac1b79c7f16c4f02f113d2fb5fae

Observation fae420d3-4fca-4572-a022-1281112ef2c7 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Language Repository for Long Video Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.874415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:bf9d5882565d628fbc03d6f3c78180c8ccab609b168f1650bb23e60476fbf3cc

Observation 0b07c839-57e8-4a8f-ad5a-13c7299fd366 · inbound

Why Do Vision Language Models Struggle To Recognize Human Emotions? cites this paper.

Why Do Vision Language Models Struggle To Recognize Human Emotions? Language Repository for Long Video Understanding

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:05:08.711449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T11:02:51.836285Z digest=sha256:9945dad4a2b1ed51729c4b5a6ff2a8a7757c26049a29855f039530c7ca3aaccb

Observation d689b70b-4670-4ce6-9430-3a91d56b3fd9 · inbound

Why Do Vision Language Models Struggle To Recognize Human Emotions? cites this paper.

Why Do Vision Language Models Struggle To Recognize Human Emotions? Language Repository for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T16:11:28.626096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:11:28.626096Z digest=sha256:2ed5810656cef8ba322391662753c87f299c58d777ba5d41189e1414d94eda97

Observation 7dd8ee88-1ac9-4f25-8b22-94e9d820cc0c · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Language Repository for Long Video Understanding

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.033857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:a8292b3bfad0ce43848f37c757fde32cd63f8ef35743d09c6fd9994d7f096da5