Pith. sign in

Paper Citation Record · LEDGER

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2412.12075.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12075 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:50.149539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:17.542022Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6fedff7f-aaec-4ec1-8767-dea3b12c325d · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.536691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:fbd7b0d8480248d8c6c5fdc220123125624d3b9eb6527128abdbc2de0e24bc3e

Observation e150a606-02fb-4a4e-8b6f-c23c42111155 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.149539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.149539Z digest=sha256:a8cfc0b0d2289df4d999384175462c34ab154498e14c9c766e48b13bfc5bc7b7

Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.529740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.529740Z digest=sha256:9265cb50fd5f4808069cba587b242e0d2904e620c5ea4c113017bd2b8b60b794

Observation e150c5ae-358f-4b63-a745-006588ae60ac · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:58.009101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:58.009101Z digest=sha256:6ffd8d29f8ac9a7f67017323d0997c0e22e403b3a034664a054f144cff2c5ac1

Observation d426b64f-a31c-4848-bfb2-48061f7ec977 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.855381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.855381Z digest=sha256:ddf2316cffbfe16315005f144f6f41030f4e3db86d5c0e782e2bc1e609d1eef7

Observation aa3c841a-826e-4063-b441-59d7aee0678b · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.805315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.805315Z digest=sha256:caf42b48744529f260d4daff7258260621542f1801aa298c43b8aa548f8af18c

Observation f0c1fecd-0d5e-46e2-96fd-1bf79aa69b5e · inbound

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs cites this paper.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.070388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.070388Z digest=sha256:a7f18590cd69e1ad0019313e1c87c483bd45dee6361073b272b338eb8d16d9c5

Observation 19e47eb9-e53e-4ef6-b09f-014ff4bff5ee · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.702112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.702112Z digest=sha256:e18ba06c2eee5efe86c10e605381b1c9e16d78cf29629d5f3d67e98d5a16d30e

Observation 60dad2f5-2f7a-48c4-8107-13874d6f42b4 · inbound

EMCompress: Video-LLMs with Endomorphic Multimodal Compression cites this paper.

EMCompress: Video-LLMs with Endomorphic Multimodal Compression CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.596950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T20:40:42.995833Z digest=sha256:8eb6d59515ed02d9c01d5070edea5df53a1d2e7191b9fa0e9578b58f79c961bc

Observation 5b5217f4-7191-472c-8ce7-66cc9f303bd9 · inbound

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding cites this paper.

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:20:22.941804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:19:36.366837Z digest=sha256:21990cb7816142ad7860cfe2f02dcff2d42ac08e910acccef11f41755285956a

Observation df201b4e-fc59-42cd-b341-1f002533f6fa · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:51.923578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:4283288524a0e9cdcb78c467edab750c83ba3c4c63ed459665b1a9dd0b77c6c4

Observation 8cb99da1-f860-46d7-964b-44a6e0bcaba9 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.453527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:6c89ef42df254f9897ea7b5ee17eddaf6cf2effbd86cd90fdf4488592bdd0f80

Observation 1e641f19-8de6-4e96-990a-0d80cdbc04a6 · inbound

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects cites this paper.

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:50.955671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:54:04.104227Z digest=sha256:316007671f3d34feead18c7dfeda6665c205ca4123b39f7b9e13940b5b4ae5e0

Observation 267929d3-e852-4c10-a1e2-0fc519398c0b · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.728687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:87807e47578c09036cdfd79a76addfb79799b238faaf256ebd11f4505cce8d48

Observation e610b167-f3da-415c-8fe0-d640e963b60a · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.414508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:44dbb1a5d5bd1878f5b9b16192fc5fc63225282ec7b1758dfd76a386e511e232

Observation 420a45c4-ccb8-4b28-ac77-fb6637033bd0 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:50818865f6e5bfc7420c86b6ad54e638c9bb42c4a4a7409e107bdfd8bbfccac9

Observation bb6e4c7b-c7d6-46af-9ab0-a9de1463aa6a · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:06:24.690741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:773f690d87a1519f2cd64747bb76698b13e9a984c723c4fc5afa74f97c01dce1

Observation c8e94c6d-64a2-4df4-bab7-3e695f1db87e · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:25.014007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:ef37a18c4e21143e4f7989a6d2bc6b6468cdf7c73c0ffcfbd2b11af2b64ea273

Observation 6aa79895-1600-4daf-9aec-21d29475123a · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.166760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:4fa97423392a855b87adaa7ce21616284ce8b64c3b4a7088a6b2631e1a6de98a

Observation 9c8f22a7-42e6-4db0-b5e9-7ccde862993c · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.545143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:1925e318f65228fb62a01fcac2ab756cc4673256708740a8da9b431d817b1d52

Observation f043ac01-df25-4a2e-86ed-07be12eda13a · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.333752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:7414ea8d3d5ddbd55607e7d253d2ded265a732e46fdf2cf65847d643b7dacf00

Observation 6d90e396-1fe9-4aa3-8d49-02c641377eef · inbound

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? cites this paper.

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.999172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T06:30:33.428489Z digest=sha256:19cc6f11e6154684e55b2d6d9adf9926e37b06a8efcec53a4d324474d1caa6ae

Observation 432ebe94-bbaf-4073-84f0-3a6c78567a71 · inbound

Incentivizing Vision Language Models to Search for Long Video Question Answering cites this paper.

Incentivizing Vision Language Models to Search for Long Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T05:50:16.895740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:50:16.895740Z digest=sha256:3413e4a99e955b432576d55e55725d069b3991ce6e4921a92978604620a66446

Observation a86fe1e0-42ac-4ff6-a65f-c6d62ca74575 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:13e5dde79b35b2bd18d16301b9e68f5d9d0db745b56e3c6c6e2c785160b306ab