Pith. sign in

Paper Citation Record · LEDGER

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.05592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05592 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.511658Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6def475b-324d-4410-88f9-f47b87a9bc99 · outbound

This paper cites PaLM 2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs PaLM 2 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.472240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.472240Z digest=sha256:916af715e93d0932a63df5eaa23190e5d004d700ae14b3789e5c071d9dbd4df9

Observation ecaf5288-9df8-4e82-a509-3ba87408f1fc · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.482845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.482845Z digest=sha256:581ee6ab0d56c4cd4650e76dbb3f95af5e689e70b5cc9bdccbce812b1bfd94de

Observation 62a6bca5-905c-4c51-a7e7-0ad0095cd7bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.489448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.489448Z digest=sha256:20bf5faf4ae9d1c7666a2f7b9d763283ce0bec65219c79c9d43bc7e4658fcbce

Observation 0c75d6b4-4d1b-4ecc-80f8-c31288db3ffa · outbound

This paper cites Commonsense video question answering through video-grounded entailment tree reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Commonsense video question answering through video-grounded entailment tree reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.492986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.492986Z digest=sha256:1f0ae37688499456f8fe548938d3b210fd02d4a13534ad638a6734ca92048adf

Observation 2963d8cb-b5c7-4071-a8bd-c57a84f4359e · outbound

This paper cites Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.496025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.496025Z digest=sha256:f2c293e27c160fd2ca19d77353a431c174c6c5f70f70057bcf3a6b929351370d

Observation ee3767d9-c33d-43d2-bb66-3cda6ad7a77e · outbound

This paper cites Qwen2 Technical Report.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Qwen2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.499254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.499254Z digest=sha256:ee59f143377d9c11cca5cf167fc2f4e63a0a73191dc7c74ab1ecbc9ceef9fa86

Observation f23307aa-93c6-45e4-91f0-af19404e5a18 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.502565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.502565Z digest=sha256:e28136fee4c045b164b069055854ca6268df41e3c813d5e2d91a214b6c527b00

Observation 22419d38-e134-4199-bea9-b3855609dab1 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.505711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.505711Z digest=sha256:8c48d53e1d2c9f980552f70e18489908b78f6aece722df3571d668b6530b3342

Observation 53c686ca-3b8d-43fd-be20-a2701b80165b · outbound

This paper cites Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Videolucy: Deep memory backtracking for long video understanding.arXiv:2510.12422,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.508728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.508728Z digest=sha256:9a29d5cbfe70566a71fa32358c1c155e4974f96c843605f8753a138efe485f45

Observation dbf2bbd4-9c1f-40b3-8550-3b437f0bf446 · outbound

This paper cites 14 C.2 Ablation on the Global Representation Depth.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs 14 C.2 Ablation on the Global Representation Depth

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:58:12.973330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:58:12.511658Z digest=sha256:2355d72c64b09c4d1e7c0eca5c8a421771913f9be11760074f1fbe309b2b6d69

Observation b0d01f99-df2b-4f19-80e9-377357e4ccb3 · outbound

This paper cites Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.475954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.475954Z digest=sha256:b8ef9fa757d062915aa430271950902a74640012c99502b5f77f345de7281cac

Observation fffa4d07-097b-4e3e-aa76-2dfb45a3bd33 · outbound

This paper cites GPT-4o System Card.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.486291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.486291Z digest=sha256:b08c929cdeef5c4f0b45392c05afeab1bcbf580d27ae88b02c2f3acb03d860c8

Observation 884443e7-ef25-4171-a095-1834aa17c2ea · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.479352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.479352Z digest=sha256:7010bd03eb92cf0f26a681e839774c0a3cfcee557c527a213fcfc61530ee4a5b

Pith citing papers

No inbound Pith citation observations are available.