Pith. sign in

Paper Citation Record · LEDGER

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.05592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05592 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.511658Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6def475b-324d-4410-88f9-f47b87a9bc99 · outbound

This paper cites PaLM 2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs PaLM 2 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.472240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.472240Z digest=sha256:baa32b046918237f583be1cc0bb255fdf3da6ae72f85d90701e842d5b6cad2a9

Observation ecaf5288-9df8-4e82-a509-3ba87408f1fc · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.482845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.482845Z digest=sha256:7e090d4e72fbc95a075b998d1de19ee42274b94198a9b2c8c39466a2d6c8ee7a

Observation 62a6bca5-905c-4c51-a7e7-0ad0095cd7bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.489448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.489448Z digest=sha256:3e7acfb2bd66bd2a3975795e9e9347f8e85cd2f7c71cb2a60fea87a8fc53a58c

Observation 0c75d6b4-4d1b-4ecc-80f8-c31288db3ffa · outbound

This paper cites Commonsense video question answering through video-grounded entailment tree reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Commonsense video question answering through video-grounded entailment tree reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.492986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.492986Z digest=sha256:0b0018c4ef07257115759744338447554044b7ae215677314a9524c3334c581d

Observation 2963d8cb-b5c7-4071-a8bd-c57a84f4359e · outbound

This paper cites Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.496025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.496025Z digest=sha256:37e501318e749fb37662656d0fa5462db990d89e7f5c932b64380782d71f97c7

Observation ee3767d9-c33d-43d2-bb66-3cda6ad7a77e · outbound

This paper cites Qwen2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Qwen2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.499254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.499254Z digest=sha256:00d4edaf6d43069e02d4f1895e47d54e08d7d5d640b134c5e9d53000638b063b

Observation f23307aa-93c6-45e4-91f0-af19404e5a18 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.502565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.502565Z digest=sha256:bfdda4ffb568b5943458d8d37c02740c7e831471532ca0e209022bd5bdfd545c

Observation 22419d38-e134-4199-bea9-b3855609dab1 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.505711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.505711Z digest=sha256:a6c9bd3334ca4227f7abf87f94796bb7b2a7d80051894885214d9c328bbadb6b

Observation 53c686ca-3b8d-43fd-be20-a2701b80165b · outbound

This paper cites Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.508728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.508728Z digest=sha256:d792b532fc8f5ce3d430a3b289660d1f5e27b90d11b3780c6058bfdca127030b

Observation dbf2bbd4-9c1f-40b3-8550-3b437f0bf446 · outbound

This paper cites 14 C.2 Ablation on the Global Representation Depth.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs 14 C.2 Ablation on the Global Representation Depth

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:58:12.973330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T05:58:12.511658Z digest=sha256:1f13a2fe37e82a915612494b94f01620ffa0d2b48470df84216c3d55ea05a83b

Observation b0d01f99-df2b-4f19-80e9-377357e4ccb3 · outbound

This paper cites Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.475954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.475954Z digest=sha256:84a760d0b16e43582087740722d0d2adb0089f9cd059daa6da8bef68bb890c37

Observation fffa4d07-097b-4e3e-aa76-2dfb45a3bd33 · outbound

This paper cites GPT-4o System Card.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.486291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.486291Z digest=sha256:75549f84b09b1b7b8d7dc0e2bb568e19f03e13bbda41dce3d290dba42ab70d2c

Observation 884443e7-ef25-4171-a095-1834aa17c2ea · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.479352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.479352Z digest=sha256:7f82cd579ff7268b0143598e35cf2e72d20494811e51091840912da527cc779d

Pith citing papers

No inbound Pith citation observations are available.