Pith. sign in

Paper Citation Record · LEDGER

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 3 inbound Pith citation observations for arXiv:2507.09491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09491 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:58:28.781308Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:30.774303Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T12:53:05.467023Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d34b630c-e942-4ef4-8a8a-0d5a9c2c5151 · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.085682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.085682Z digest=sha256:62325445de7608fb3b13006db711a6fb6aab3e6bd2f610a8a94d9fc8c80e5501

Observation 2ceb502e-6b49-4ab2-a2ca-701b69ed2a18 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.181609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.181609Z digest=sha256:59f930a3e2ac1219ace14163c17e81ecfbc12b6396f03bda4d9c87a0bd1b13e6

Observation b4de45de-7812-4775-9d3e-8ec6a1d423cb · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.320461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.320461Z digest=sha256:fc20340d2c1c9219bbcaa8e1098d6da8e958d9b3503686e985ee21d28daa436e

Observation 429e1c97-5d48-4974-bca6-377506aab120 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.390172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.390172Z digest=sha256:bf39a4b0525e72d0bade6817b9e97fb3643cc4d496a0c641e75853e20cc0ffd4

Observation 5a7676cf-84d0-49ea-9aba-f61972b1194c · outbound

This paper cites ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:58:29.172776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:58:27.448450Z digest=sha256:5b3e8be77545ec4cd8c0dbb3f2be5c7bb60e0db716614e246e2007ef41e68bd3

Observation 2e9fe91c-3242-4a7e-84de-3a088856acb2 · outbound

This paper cites VLM-Eval: A General Evaluation on Video Large Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? VLM-Eval: A General Evaluation on Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.497281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.497281Z digest=sha256:a8d4af83d7d723cd287d6e5cfa5581a8cecc0bf52f67ddf28b577aff60ae31e1

Observation 12d3c181-6b2b-48f5-9a19-2e74dab9d8b7 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.574010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.574010Z digest=sha256:0ffd3fc4bbbad5477860a66bc02fc0e07875e20e979b0a6eee76a587fe9cd08c

Observation 29e4b54c-e510-4e57-a49e-06d18a014e81 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.664978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.664978Z digest=sha256:8ff3d4a9f2e231d5a2213dcfcf1187244ddfaea5c8bf4fb1997b3ac677d5cdc2

Observation e3afeb3f-c265-4265-b6d6-c827bc88199f · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.742728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.742728Z digest=sha256:baac4639ac03c8b48f9c3bce0878ca22029e53bc4e9a46c50aa50703f256fd4e

Observation 6bae6f8d-83e3-4ed8-bd95-2e4eb9a51a04 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.869001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.869001Z digest=sha256:4e53f63c795997ed42650677e953d0dcf256a769e077d5f327ad83c0863ee876

Observation 6a1d268b-9279-4f76-b196-2ea10c60c1ac · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.048573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.048573Z digest=sha256:9a64dcd848d231bab563facddb96f6866acc152824883e9e85b0aa8d4285c2f5

Observation 0df38b9a-421b-4c41-bce0-512fefeb66ad · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.123501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.123501Z digest=sha256:dc8edc39d2c78c0bcbf4e0e4825656cba07b0c1a013156a5423f9392c32c40d4

Observation 9690be43-8fa6-4a16-8d82-8cfc3cee34b6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.210161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.210161Z digest=sha256:65f30bd3cf06367c0ba7955467bdb93f095880ad9deb31293f63f5568f7cd6a0

Observation 5bc9de2a-5ae8-4fba-a9dd-d6a1ad4f60d7 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.290683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.290683Z digest=sha256:141dde351ce3dd01cdffa6f9548a14397b85cd43166600f63cfbf4c6de57769c

Observation f38c72b6-ab00-48f1-9357-6ad9026aae26 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.410060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.410060Z digest=sha256:f147ffa4d120c31029865a0c9be51be6de2de15334181231cc83f41846ebb770

Observation c90621e8-d91b-4dfa-bc18-bf43b47f0302 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.607448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.607448Z digest=sha256:f4cd324bcd8b51c66e74bf45261406cf5606f47d7b7137cc75e19ddc0211ace9

Observation cee639af-b059-46b5-a330-3c58a5c54f4c · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Calibrated Self-Rewarding Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.693651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.693651Z digest=sha256:7e535e5b2c4e4a9fd26ce4b1bcc34d87c64e6c1ba100604f5ef91c2921cd5836

Observation 7185eb2e-ce12-4469-a11e-d8a45750ae99 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.781308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.781308Z digest=sha256:7c25643c31d8c68cc04eb3d8b4e0785a6099af5a8f364309cee1ee818193916a

Observation 51d4fdf8-37b1-40a6-99de-acddef1db62f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.922661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.922661Z digest=sha256:403b9d190e823cb1f46e01bf8ad0b650bc8088fac42a512509c2a898966423fe

Observation bbe643d1-35b2-4d2e-9e1a-e0b64b67fc02 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:26.999928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:26.999928Z digest=sha256:4ce6a3757ddb87d64a012619880f80c3f1f1a324ab2ee0ae7e076d674857116d

Observation 1ab8663f-6b5a-4105-a085-03e570d9dc04 · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:28.549081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:28.549081Z digest=sha256:cbfefdcb8be685a23465d6c65be98765c9ac34498dee91cb189badea88bd8280

Pith citing papers

Observation 237a465b-3f7e-41f7-992a-de9b69301f24 · inbound

Low-Cost Test-Time Adaptation for Robust Video Editing cites this paper.

Low-Cost Test-Time Adaptation for Robust Video Editing GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:30.774303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:30.774303Z digest=sha256:d115203c4a4cd3ba2f48011b7b78999196ed07c0fe1c3a436601a6f28b3416d1

Observation 68f4f59a-9d09-44bc-bba5-561a89b54990 · inbound

Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers cites this paper.

Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:53:05.471912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T12:53:04.416698Z digest=sha256:a8633880fe9f495c0309d806c96f9a8696e3ffae8353afc558e20e04fbfb1692

Observation f14a26fc-4da5-4428-b4a9-2c93fcae3bf0 · inbound

A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture cites this paper.

A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T12:53:15.118127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:53:15.118127Z digest=sha256:9e7f91ca89021d375a27a22ea39111059b0c4e539185830063725e8fd6768839