Pith. sign in

Paper Citation Record · LEDGER

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.05592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05592 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.511658Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6def475b-324d-4410-88f9-f47b87a9bc99 · outbound

This paper cites PaLM 2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs PaLM 2 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.472240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.472240Z digest=sha256:847f95e5074d6fb111d4224c7641c35cfe9f497f0a70928cc5f7b5dc67fc431e

Observation ecaf5288-9df8-4e82-a509-3ba87408f1fc · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.482845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.482845Z digest=sha256:2ff1ff29d5c8f874f9f84e48a84497a5e2b098cea8fe054e73435ba44979c4c4

Observation 62a6bca5-905c-4c51-a7e7-0ad0095cd7bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.489448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.489448Z digest=sha256:a5cea03aee05a8e04284de5bbdf95365ff992845740d410408c1868efbd631bd

Observation 0c75d6b4-4d1b-4ecc-80f8-c31288db3ffa · outbound

This paper cites Commonsense video question answering through video-grounded entailment tree reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Commonsense video question answering through video-grounded entailment tree reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.492986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.492986Z digest=sha256:97e263feef798e4fa20bcda0429c9c69c4a34e059b7f673d9a65d045b0377e64

Observation 2963d8cb-b5c7-4071-a8bd-c57a84f4359e · outbound

This paper cites Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.496025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.496025Z digest=sha256:1c0e3b27913e426c149656b61b220571ede19e9bc18ce048da75bb66d088bd54

Observation ee3767d9-c33d-43d2-bb66-3cda6ad7a77e · outbound

This paper cites Qwen2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Qwen2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.499254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.499254Z digest=sha256:d6ca35871a7e2a3daa4fc1e506baae11330cca83d783ed899eda7fe5800e0c81

Observation f23307aa-93c6-45e4-91f0-af19404e5a18 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.502565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.502565Z digest=sha256:c68e53280040b24d761aea17b1a25dca6f9b896eeb9d4d56ced9847e7d02671a

Observation 22419d38-e134-4199-bea9-b3855609dab1 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.505711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.505711Z digest=sha256:737d85028899a1818f38d4d8c1fde66b957b5557c21775e520c322abc22199ac

Observation 53c686ca-3b8d-43fd-be20-a2701b80165b · outbound

This paper cites Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.508728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.508728Z digest=sha256:f493c4837cabb6c2b9a8e04284d74d97334be22d18226f9741f99a96a3c1cc68

Observation dbf2bbd4-9c1f-40b3-8550-3b437f0bf446 · outbound

This paper cites 14 C.2 Ablation on the Global Representation Depth.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs 14 C.2 Ablation on the Global Representation Depth

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:58:12.973330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:58:12.511658Z digest=sha256:5b13ef0a31d062b4016ca1dfc66ef77b8233099b44f3246aab94fba42dcb113d

Observation b0d01f99-df2b-4f19-80e9-377357e4ccb3 · outbound

This paper cites Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.475954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.475954Z digest=sha256:0c0537ebee21e522c0adec5f974b48faff04c21aa630b3f88109f9b4cce5dbeb

Observation fffa4d07-097b-4e3e-aa76-2dfb45a3bd33 · outbound

This paper cites GPT-4o System Card.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.486291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.486291Z digest=sha256:a8033b141302f7302fe953202bf1dec313353839fad707908ca97aff9ed7bbc3

Observation 884443e7-ef25-4171-a095-1834aa17c2ea · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.479352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.479352Z digest=sha256:9c15fb01e34389aed90097b09464d2762bd73a95cb05776315a10d466fe31326

Pith citing papers

No inbound Pith citation observations are available.