Pith. sign in

Paper Citation Record · LEDGER

Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2312.10300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10300 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.807797Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.149993Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5eead944-f033-4430-97d5-61140c8c3ef9 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.807797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.807797Z digest=sha256:9fef63a69c845c76bde087694c45e4e2a318a39530b59433c36cffa935ae0f0f

Observation 3959a487-d5e7-4256-979c-3260fefde00a · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:59.302001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:59.302001Z digest=sha256:9271c8470e00d3a33c05582f20e3c87822ecdaed6d3b6b978faea92330dcb7c4

Observation e5aed2b2-69a5-42fe-87bd-b4dce7efe8ad · inbound

Comparing Learning Paradigms for Egocentric Video Summarization cites this paper.

Comparing Learning Paradigms for Egocentric Video Summarization Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.969901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.969901Z digest=sha256:bbf9c4aeaecd8d7cfac4ef78e0b08d184a5955b405be973a6904738f6b87cd20

Observation 6f53dc11-111a-4c9e-9ca0-ef9b3f4c1988 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.482506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.482506Z digest=sha256:f8b632a0592bad946a2553b306a9783cbbff00e4f06564eb0ced5eb7ae2f9804

Observation 6b864fc9-9403-4a05-a35b-e0262781c528 · inbound

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting cites this paper.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.824025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.824025Z digest=sha256:5e313eddede195c0a4c997a62499c65e6c4de6e98eb9b982ae0c651f37bd18c4

Observation 71c663e5-043f-4578-b719-2fe8944040cd · inbound

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding cites this paper.

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T18:39:29.914405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:39:29.914405Z digest=sha256:88c63272844b6f0ef73e7ce9ffbd6e850e720a86df0e5bddbbbee22161e2b058

Observation 072fa698-ddc5-4a3a-878a-c96338c1d383 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.847902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:5f67f21ed1888dd795cc9e993b019bce1845480d4fc7273bc1334dc820f209c2

Observation de5a588d-53e0-4229-a971-0412b48b708f · inbound

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production cites this paper.

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:13:29.933725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:10:31.245928Z digest=sha256:c10afdb35525183bdbbf29bead573f6ea39426051e959d02dbd67503a414dda6

Observation 1b527f97-cdc2-4c25-a66d-72bd07c71120 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:19.082732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:28:42.129881Z digest=sha256:42d29591feefb6c54ce02d061a8a85252fbcaa967723c5e70f93d2e5c49016be

Observation 78b7ed1d-47e0-4726-993f-b16a86d3fa47 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:15.120188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T00:50:10.509727Z digest=sha256:cb2ee2e3a4b1bb906d7dcaff361ec6f415c27dd0a43d61a758fefddaa88f9811

Observation afdedb8a-49e6-476f-9ab3-44a88148bfaf · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.681985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:f852d6afe4e2aaf5188911bb253bb3661d21bd78c3388b3d16c34a1163c1a509

Observation cabb546d-af1a-46f3-8abe-4ff690e6b441 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.151854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:886cd69ed89818eb7c79c5e2142582c6bc8121f4aa8b40ed0bf3230bc15f455e