Pith. sign in

Paper Citation Record · LEDGER

MM-VID: Advancing Video Understanding with GPT-4V(ision)

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2310.19773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19773 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:32:56.628654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.465524Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 495627fe-3d36-4200-a04e-5b3679e4326d · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.628654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.628654Z digest=sha256:78bcbce986f63cdedff24f58756b4d278c51663b82c6dcc964fe7dda1ce0cf03

Observation ebbb2ed5-4fcf-489a-bc72-066c46ea492a · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.209434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:79246b128d0773c3c74b749a2905fc3f309008720d6438d5a5de0ad811cff142

Observation a14c7dde-17b6-4d58-bd66-0a8d1c0f5471 · inbound

Making AI Drafts Count: A Quality Threshold in Audio Description Workflows cites this paper.

Making AI Drafts Count: A Quality Threshold in Audio Description Workflows MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:26:07.983894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:07:12.813384Z digest=sha256:8c99a05b60e02cef68ded57ff2775fb1f1997fc1035982defabd97b7e7107c41

Observation a29d900f-fa37-4fe9-903d-ee3b9f3b9e64 · inbound

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration cites this paper.

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:18:18.351189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:15:54.413960Z digest=sha256:5be7add3de5eef3d2c0ef32ea6418aefe7f0eb1bead6b963247c96d11c22a460

Observation e48ed5ef-bf5f-40ea-851b-10f4389b881f · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.177575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:61530dd68fdc0ae3d3f8827c097107ea8e35af43a7965e873b8ed48b03404787

Observation 0e7c4054-6bbd-4e1d-abc9-afb5768d06ba · inbound

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations cites this paper.

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.466891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:09:23.918802Z digest=sha256:d801c5019cb4d0390c966f9c961a2ee9a24aa4d191519a39d0ea823184d6c5bb

Observation 2ab7d186-2e3e-43ec-8421-433eb97269ab · inbound

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA cites this paper.

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:48:39.319110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T16:45:20.139230Z digest=sha256:03890301092a130b91b5577927c33a39412ce0c5d680af80d08ed3dd54aef6f1

Observation ebcd45b0-7ca6-4881-adda-8b9c5ec37d2d · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:56.204492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:56.204492Z digest=sha256:19d582e87997e8d55d75275f7576a7b6af21770f7dd147ee86bcd1ea8ea4c440