Pith. sign in

Paper Citation Record · LEDGER

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2411.14794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14794 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:03.727067Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:27:44.037400Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3712b52b-dec3-41ff-9eb8-98ae17197a82 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:44.041073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:d7fd710c33d6fe11cc7fd802916491cddd8fc42b0c34eff970f63de68b2b0f11

Observation d5d8e53a-eb28-4c15-82dd-5b8c67579206 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.266805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:92b9d0699a037822cfd8afb4e1cde339ea3d1b26ac7be5ced042b8e046f3d6f7

Observation 4b8a6322-451b-4ed4-ad9c-d2caa6fce31d · inbound

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR cites this paper.

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:03.727067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:03.727067Z digest=sha256:668204a4b14913c6adad47765c342811d2fac1e3837668d65be3b14a9b87d1cf

Observation 2771756b-e8fa-4632-881e-95c545e86d39 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.372866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.372866Z digest=sha256:fbd2a9130e91c809de457e7bcc44df7d6edd662b353bd4220f111566cedbc9c9

Observation 7c58fbad-de91-4d02-b625-61d4030b97fa · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.852974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.852974Z digest=sha256:95caa38107b9d282ba402c90547ef2d9b00851dddb25cc4372b591e3fc201d09

Observation 7b07519b-5ef0-457f-b17f-6d27a24ad23e · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:54.039135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:54.039135Z digest=sha256:f6b9d4a954c63d3edc7d4624cd09ea845540f948412450a68600190cc91749cb

Observation 55638bfe-9742-4f24-9d1a-33ba2c263953 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.618589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.618589Z digest=sha256:16cad8170daf01744835d20ce7ae625839907365aa8f60c3f1d320c224f3f2ab

Observation 83fac34c-6574-4a22-a63a-a92e08f38943 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.372667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.372667Z digest=sha256:797d885977e469cdfdd8ad1b64febb483f6d6fc900574371b533b83856f41e34

Observation c9c3a023-4234-4918-89ae-f0b15cb0194e · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.642866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:a45f4866d8e424c58cfc3d0cf4202a990806e54d436cfb1250a7b70a7da4696f

Observation b8bc99ed-566d-4acf-9f27-d1f693e9db29 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.051371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:a0181963bfe3f261d4090145482b50ebe62681e8da258335452c8af225b457a6

Observation 258700ec-3943-4844-af85-1f3c86fba121 · inbound

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning cites this paper.

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T23:10:24.087228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:10:24.087228Z digest=sha256:7c6888b65553159799e200771b547eb7b9ca1a67d07c306db7e77ec52d1a12cb