Pith. sign in

Paper Citation Record · LEDGER

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2505.24371.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24371 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:34.249618Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:31.083510Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:30:34.809388Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c65a519e-1029-4de4-bda0-8b02e1b64b5a · outbound

This paper cites Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:34.865426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.083510Z digest=sha256:6bf2fb944af530171edfef5cebee1f7504663eed7720174a411a5be6ed7c174b

Observation 352a4101-0848-419e-83ce-c1e6b08b7fb8 · outbound

This paper cites Further research [5] extends the image VLM to support multi-frame and video data modalities.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Further research [5] extends the image VLM to support multi-frame and video data modalities

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.424721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.146651Z digest=sha256:0a2ec4b56fccffbd6ade043d5fe3d465e5b0e7491812b79db53c50be089ab2c4

Observation d97db90e-1f28-4999-8b0c-b3b39bbfd3d3 · outbound

This paper cites The first phase is the tran- scription generation phase.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The first phase is the tran- scription generation phase

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.266635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.239277Z digest=sha256:ef96ac69931f6cc6d4bd3c0e17ff87f997e92114958d0aea901dc2f9c61f9462

Observation 74f755b7-84d5-41e9-a367-80dcf7a6f2de · outbound

This paper cites Baseline models We leveraged the LLaV A-1.6-7B VLM [23] to generate local and global transcriptions and the Llama-3.1-8B LLM [24] for the VideoQA task.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Baseline models We leveraged the LLaV A-1.6-7B VLM [23] to generate local and global transcriptions and the Llama-3.1-8B LLM [24] for the VideoQA task

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.179602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.336355Z digest=sha256:51427b3de1bd20fea79c433968ac7a1a84e85cf390ebd243e4fa96d405ea8d57

Observation 797c5fb9-80c5-4173-a6bf-a4b17cc92020 · outbound

This paper cites The system is inherently modular and built on open-source VLM and LLM.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The system is inherently modular and built on open-source VLM and LLM

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:38.058928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.423370Z digest=sha256:5b9d25d41fb4315d0719a28818f68b1120c49ef15d31475eac99e97d36913205

Observation 53e15e52-b447-4329-b1ab-f728b312f753 · outbound

This paper cites MOKA: Open- vocabulary robotic manipulation through mark-based visual prompting,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering MOKA: Open- vocabulary robotic manipulation through mark-based visual prompting,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.875398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.499625Z digest=sha256:a43970d2732b59c2e90e372d9980d512094d8cff1d50c6b697b93f4041c794c3

Observation 382d85ff-e26b-49d7-9343-02f316d43980 · outbound

This paper cites VLAAD: Vision and language assistant for au- tonomous driving,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VLAAD: Vision and language assistant for au- tonomous driving,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.669183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.568370Z digest=sha256:3d35539905ed708bf574618e5230a1f9cb2d615bb72eacb14c8352853186db8e

Observation 59c1cdc9-db3a-4ac2-9831-49adbccd3cbe · outbound

This paper cites Smart customer service in unmanned retail store enhanced by large language model,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Smart customer service in unmanned retail store enhanced by large language model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.490277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.641448Z digest=sha256:5ac57b531fcc333cb36ea1365077e91a8d3233b1eebf5e04c2d2ab75671e75f5

Observation c0d8171a-5c54-4539-9534-0024ca4953cf · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:31.714734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:31.714734Z digest=sha256:770203b634982f67ba3094ac2a3461b9fd38a4f7cf84aa90c42a7da013347b09

Observation 0b8ea0af-1122-47fb-9588-f63d5c0fe1f3 · outbound

This paper cites LLaV A-OneVision: Easy visual task transfer,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering LLaV A-OneVision: Easy visual task transfer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.325359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.782618Z digest=sha256:f7cc09aedf51cc0fbb893bd674a55d623b59e8c78de9cf18ecf221f06ffa3e28

Observation ac08881b-48a1-4659-a175-f4a53b64c0d8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:31.859767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:31.859767Z digest=sha256:305945a978cdda86ccbc88e37ab8c35c793123918b91dac0a28eb6f4668bcf9e

Observation 31d3d6c0-03e8-4cb2-803f-7631c67d5c9d · outbound

This paper cites A simple LLM framework for long- range video question-answering,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering A simple LLM framework for long- range video question-answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:37.158371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.933492Z digest=sha256:49d9c0213c8928c9e89c91e2327ec1709ff8ce9019edeff9999964480ce35612

Observation 0c71c004-3791-4536-a46f-72046256c69e · outbound

This paper cites GPT-4 Technical Report.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.012193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.012193Z digest=sha256:b3e0da21330d3bf4046a3ea4c2825b4d84df7734fa2fe19426ca7e8ce941a643

Observation cbe3c9c6-fc2b-408d-bb41-042cbc11112f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.096412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.096412Z digest=sha256:20a174cd358acf773414e078fdd0d603e92e07751e2198c81e9e39ba449318f7

Observation 5d6d3dc4-2dc4-41f1-ac3d-f8b7950efd09 · outbound

This paper cites Visual instruction tuning,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Visual instruction tuning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.202556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.202556Z digest=sha256:632476ab0268c277056842397dd1f316d1a07a2fa24f711310e711474702fc4c

Observation c49c3092-be31-4fe1-936b-8c2c65624229 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering An image grid can be worth a video: Zero-shot video question answering using a vlm,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.936620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:32.275288Z digest=sha256:3eb32d0780bc9dbbc8797e9fb16022af51d134d3498796923a2b6b3989f723d5

Observation c0414aa9-916d-4f00-9240-84c5999a2165 · outbound

This paper cites Self-chained image- language model for video localization and question answer- ing,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Self-chained image- language model for video localization and question answer- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.788952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:32.380318Z digest=sha256:4d806cc098b2a5e974a8b84182dbe3b9931ad3121bb656c06115d29722ee5cd7

Observation 3768e06e-5716-4740-8abc-34f9f9b9459e · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.456684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.456684Z digest=sha256:103c1891bba2730391aa17f391c53d1f83c6e50c965ba405665270fc99c976e0

Observation b2e5dfae-a11a-49e4-bcbc-1eedfffda2b2 · outbound

This paper cites GeoChat: Grounded large vision-language model for remote sensing,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering GeoChat: Grounded large vision-language model for remote sensing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.594439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:32.558947Z digest=sha256:3014b2495930a4122dc7760864e6c3cace051a8fab1ea59073af428a2b90451d

Observation 21b7a483-93aa-4aa1-8821-9ec4b97b7df8 · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:32.685636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:32.685636Z digest=sha256:e978c3162de4b3c3e0e504c8b94d30943a4d01186d56709fd9562e5b46cc2523

Observation 609b1966-0f06-4ac3-8f54-bb72cfdb3bb7 · outbound

This paper cites PIVOT: iterative visual prompting elicits actionable knowl- edge for vlms,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering PIVOT: iterative visual prompting elicits actionable knowl- edge for vlms,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.385702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:32.807866Z digest=sha256:b771a97cff81d08dfb436d5d48f57b052d905a56e6c485ae3a99edf251c4c8fa

Observation 4190cbcb-4f1e-4621-a6e7-d0333ec88c03 · outbound

This paper cites Open-Vocabulary Action Localization with Iterative Visual Prompting.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Open-Vocabulary Action Localization with Iterative Visual Prompting

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:34.604723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:32.904432Z digest=sha256:99f2a21f8d4449c31bcea8f840ae8350461a1dbb15bf1d3641bfecdb223b62ee

Observation 2579bd6f-136b-45e3-97c8-57fd2d7e1a2c · outbound

This paper cites Video graph trans- former for video question answering,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Video graph trans- former for video question answering,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:36.199823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:33.038946Z digest=sha256:5a76b41823d41dc6c0b9c823606a2c12d893930c79c134ef29e9086504d684cb

Observation a19f1bcc-2c7e-4e11-9527-3b1c0426aa55 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:33.191576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:33.191576Z digest=sha256:fbdb79dd1310d23e07ad439291be8667c906cc66563796510c54cb8f4ee21355

Observation 5a61dbaa-5afa-4bdb-af4a-d952a8e3dd2f · outbound

This paper cites MVBench: A comprehensive multi- modal video understanding benchmark,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering MVBench: A comprehensive multi- modal video understanding benchmark,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.967511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:33.328725Z digest=sha256:f46c730e158b6931a9e66e0d675f4ef6b5c489c04330e8df5408950a653b8529

Observation e5458f8c-7aa0-41de-b65b-35e2535e1d72 · outbound

This paper cites VISTA- LLAMA: Reducing hallucination in video language models via equal distance to visual tokens,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VISTA- LLAMA: Reducing hallucination in video language models via equal distance to visual tokens,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.767758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:33.421014Z digest=sha256:487164ee9331f333b3f0d56edbc87a1d423fa16b429efb32a1112268a0653cfb

Observation 3c87cb0a-7f1d-46f9-a926-58e4b0f43823 · outbound

This paper cites CogAgent: A visual lan- guage model for gui agents,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering CogAgent: A visual lan- guage model for gui agents,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.515474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:33.562294Z digest=sha256:ab11bb4f22313034f5522dd76d1e00954c4f0709ace760a5a0c7b6996f7f8e32

Observation a087d91c-8ca7-4179-9248-8d249309ce08 · outbound

This paper cites LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.344142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:33.653791Z digest=sha256:d98339cb213ae7e41738efb52fa9382db41f51c288781b95aa6498f14ba296f5

Observation 3bda2006-30c5-4bb7-a8ee-daf3b6c3f689 · outbound

This paper cites The Llama 3 Herd of Models.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering The Llama 3 Herd of Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:33.760176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:33.760176Z digest=sha256:b3d0bc2a118b620046b0f485e001bf3fb2a80c74abe15b8c2515aefa0e628607

Observation 16576960-2833-4818-a5b1-2a2608474e3c · outbound

This paper cites NExT-QA: Next phase of question-answering to explaining temporal actions,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering NExT-QA: Next phase of question-answering to explaining temporal actions,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.229409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:34.055988Z digest=sha256:724ff9e8ffeb250a9bf3e6c3279b39afb5d235ae2d2b6811eb14448348969ff3

Observation 22ba6029-c29b-4815-abb5-4aeecc966736 · outbound

This paper cites STAR: A benchmark for situated reasoning in real-world videos,.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering STAR: A benchmark for situated reasoning in real-world videos,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:35.085113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:34.249618Z digest=sha256:fb125a1693728b90b64eea7916ce95b8ddcca26f7e24ed8780ec6d7a643b99ef

Pith citing papers

Observation c65a519e-1029-4de4-bda0-8b02e1b64b5a · inbound

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering cites this paper.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:34.865426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:31.083510Z digest=sha256:6bf2fb944af530171edfef5cebee1f7504663eed7720174a411a5be6ed7c174b