Pith. sign in

Paper Citation Record · LEDGER

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2507.02946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02946 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:59:17.302898Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2df4ca94-0573-4ef3-8129-00b94da7f0f4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.678922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.678922Z digest=sha256:e11de2784f689969dfacd453e15a297e781088636ac6b597c6cd8039227522df

Observation 29312749-d544-4313-a39f-c76c0a07fece · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.758936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.758936Z digest=sha256:809d65207a2f14690b3fa15bc2ad1626b037defd1156b5851b19ef1f94b35f28

Observation b4e04e31-54a1-4856-ae77-3dd741cf5db7 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.855545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.855545Z digest=sha256:204c21de4e9260b9c4a46ad53c3fff26281d1f97a506f73b6f5da757c6511ff7

Observation f02e1bab-94e6-4fcd-92b6-0f79569b36ec · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.010886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.010886Z digest=sha256:ba88ee678041b58dde0cb4d0d772d926cf8d596b1a8433d1384777b144ed85e0

Observation 48c49d92-36e3-48a5-acaf-059faa0018a9 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.134001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.134001Z digest=sha256:fc102e7e3b146a3f46983adb14a88482995b777e8f87865ff5b599dee21f1123

Observation 6048adb1-94be-4ff2-b950-086bcb07d394 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Reasoning with Language Model is Planning with World Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.221161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.221161Z digest=sha256:c5f559c6a06321aefc5a56a9427b9cd207def92d9a6e3ba72338eefeaecbd01b

Observation 8faf21c6-a3bf-4ade-a893-cd6fe75dedf0 · outbound

This paper cites From Image to Video, what do we need in multimodal LLMs?.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding From Image to Video, what do we need in multimodal LLMs?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.278173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.278173Z digest=sha256:1621444be7c5cb83e09e13f578e83f6745f73f0a52670a233b420e6ecd181d65

Observation 110c9dea-d75d-4a06-9dd8-94baa733e875 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.444148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.444148Z digest=sha256:410a69933d07036d40fb9d336357ffdb14e43db0cfd94361c8e99718de447fae

Observation 08d20e2c-0330-46ff-9b67-fe558e1ebfba · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.488837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.488837Z digest=sha256:14eb20c606bf53477b4ff517d85f3953d37b7df23f8ea7840b25e5406c47d645

Observation a2ad6307-3313-41c9-887d-fdf025e90a0e · outbound

This paper cites PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.530036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.530036Z digest=sha256:0ca50ff6ebcf0c8e805153a1a2d3b2d3cae30d0c2814db58437b5d7554c73668

Observation b822c8c7-c8c3-4c25-8055-bba79cbf1794 · outbound

This paper cites Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.614884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.614884Z digest=sha256:b9cc461b0a67e78b10c6597b97775795612206dc387cad8ffaba1fbe8fff9944

Observation 7d870a68-30b1-48c1-adc5-d6dd3459f69a · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.696016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.696016Z digest=sha256:d28c181dd5df2ae6839085ca8e028e41ca214cfd7f06a3c51551ddea17289241

Observation ca388100-361a-4ccc-b685-43025ec203b4 · outbound

This paper cites ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.773771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.773771Z digest=sha256:4ae3831520ce5fb082e6a215e29d527235761e7db7e97f16b4f95e26e613f919

Observation a03ecc67-4006-4690-9bf6-bdae02f2ad6f · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.899236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.899236Z digest=sha256:d7a1c5dbdd2a965ac0fb282b12f64c467104bd43bc0909f609522947c880942f

Observation 91492b46-e8be-4b35-8952-a6a972a07059 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.981202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.981202Z digest=sha256:71abada6c0cda656bcc644ff153b6022e7a84f8240f5be63169a5e385d93554f

Observation 587ba512-870c-4a04-bbaa-48b8069a308f · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.153103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.153103Z digest=sha256:a23d8e568f3290ef70edde27a0ea563ffbb5b505c246cb6c8d4e2a1e66b8fefe

Observation 53a350c1-9421-43e5-a25b-1b0516cc88a1 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.302898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.302898Z digest=sha256:4749d2385f2ae37b327442bbf5de009ae6598b07548582e028309ef422e5718a

Observation ca2d00a3-b6ba-4815-aa27-272eeedd7da0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 1985

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.364927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.364927Z digest=sha256:32d3a16ab3bf70b61579efb98d6afe4264be88cb122042a7e2cb3a65dccd8ce1

Observation a3101767-ba83-4287-90e5-94ed395b9614 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.224605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.224605Z digest=sha256:87219f0eb800ecb0eda51c8ea91032cd02afe56cf82358320f871837664639d0

Observation ec3ffded-ef9e-43d7-ad38-a67c8cf17179 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.058810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.058810Z digest=sha256:39f143de5e7f12f17cc4244bcdb760a90eefeb0410ed95d1999df894a4cda199

Observation a3e1ddd9-6e69-43fd-9211-6e2a79ced978 · outbound

This paper cites GPT-4 Technical Report.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.515921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.515921Z digest=sha256:e1caeb22dde40bb22524d7e131026bb295907198c4f071adb2037409fd999109

Observation 07a3c23f-0848-4e87-9a93-13c1267286ba · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.950597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.950597Z digest=sha256:4d49fb6ddd14a4a5257d3ddfa69ca31491fe26838f91530acaa69de3fef26e8a

Observation 0ebdbcd1-1622-47c7-8529-22f36dd1d1dc · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.578220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.578220Z digest=sha256:5fc066b8b6ace3af9ab2bac90663c2b9e24b5ddb830db48f128af572821c82d2

Pith citing papers

No inbound Pith citation observations are available.