Pith. sign in

Paper Citation Record · LEDGER

Movie2Story: A framework for understanding videos and telling stories in the form of novel text

As of 11 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2412.14965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14965 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:48:07.259130Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd9a2d0f-4f81-454b-86d7-ca7e8b50d817 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.102344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.102344Z digest=sha256:59e0b68e35747e519a585a395a45937d3d8e75e8f2f2c7aa51417c435f40965f

Observation 6629e209-1133-4450-b9eb-37a1247dfd21 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.106481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.106481Z digest=sha256:b6bdc222a33c9f37ce45b40712b4ba4b268e49f1b99274d4e9f87d8e71533011

Observation 665c0f6b-3883-4577-8704-1dba32b7c9c1 · outbound

This paper cites HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.109873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.109873Z digest=sha256:4e26b64ef30da4c4364cfa08d0a4a34650402b31a2220adf53d4adbf9477f931

Observation 5ddfea5d-024b-46a4-b01a-fb5db9d6ef34 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.113036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.113036Z digest=sha256:cfd9c650c9c93a52bc12f1c7c5fa3d581740396afe6d156832f8246c8dde6a20

Observation b93bcddd-58ff-49db-b644-c6af165c04fe · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.116192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.116192Z digest=sha256:03bccfb9d47d4d764cdedeaf2f011a590cc48d78f5dd233c33ed462c0eb43b33

Observation d7e4111d-1e26-4fea-a0c8-528596a46b7a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.119559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.119559Z digest=sha256:54a386868c2b013f5c0c3bab79a10264ba99c02ba4493b0e306ffebc7684072b

Observation f96e29d8-7ca7-4892-82c4-dc466d4740d6 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.122823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.122823Z digest=sha256:ba6c29154be6430022d2568946172b02c6deedc05a99cf3114baf4badcfe6f6f

Observation ef7a6140-d6ea-41cd-84e2-b95b63334213 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.125832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.125832Z digest=sha256:ddd7f94fadff938e1dba0b8ebb5c47521643149c5d97d5a4fdf1c82c4577158a

Observation de6d5932-b665-4bb6-92a6-5f3642cdbeba · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.128749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.128749Z digest=sha256:c6ddbcf48c6725388c5252de846005e05542ded256228cfeeacaa02d1242f72b

Observation 7ed0f38a-da12-4fa4-9290-714370f0ffd6 · outbound

This paper cites The Llama 3 Herd of Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.131658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.131658Z digest=sha256:40564adae0591645c7de3cbaa651727cd3b55b5d2bb6c7682a80dab1a027607a

Observation e19bf70a-4876-4679-aec0-124eb466a99b · outbound

This paper cites Mistral 7B.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.133826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.133826Z digest=sha256:1bb5d66692288c8c0fdc8316c4f8b0a9491dd2c63848e24e0fe372078a43be07

Observation 3f4e0708-b829-4151-9a3b-7db13f38056a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.136034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.136034Z digest=sha256:c3d7e8dd4332f35e6be1312fa1a715bdee9e17c0f766d9e79e1d6986d39f2c48

Observation 2c77b2f8-6403-4a6f-8bcc-1d2915e1130e · outbound

This paper cites pyannote.audio: neural building blocks for speaker diarization.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text pyannote.audio: neural building blocks for speaker diarization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.138051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.138051Z digest=sha256:9ecadf20f7850573ecb012416b7243749e3c14a3ab421af161ac85099f36bab1

Observation 9533073b-bed7-41ba-8f87-3d69b11d8777 · outbound

This paper cites Qwen Technical Report.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Qwen Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.140182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.140182Z digest=sha256:79e6e0e0aa01c5ff2fef63cf25645618029dfbef95468b4a463a285a1834a213

Observation 918f1167-b37f-4b81-88a2-be8fba7523a2 · outbound

This paper cites GPT-4o System Card.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.142835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.142835Z digest=sha256:7dbee9506348784da492b7bd0d6a4a5421b6b19b86162806c384f45dd9a00501

Observation e7f9a7a9-556d-458a-bfbb-362be4585dec · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.145680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.145680Z digest=sha256:95fc22c945d7faa6dc6fe38bae4b4836bd86ec981d240a34f2e8d8b346601b3e

Observation 1587efce-0860-4235-a2f9-c4753eca11b1 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.919248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.148685Z digest=sha256:59e997bb582ff8f53826b41df924ecc139afaaf627e8b2a3f815969446f4ea71

Observation 918c887c-7417-4ee4-8c50-3483b9cf65c3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.152211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.152211Z digest=sha256:71676dfdab4d61d91044865ea0f2195436f0820feed7eb96c577b3eb741ebc64

Observation 0d037024-1051-4a3c-a188-cd172d282873 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.155335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.155335Z digest=sha256:65851e5a077739ca6f9e2e819938c9642e73dd58796a63955d7728538a0c22fe

Observation b37fe7cd-2232-484e-bdc8-a2fada3b1a5d · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text ImageBind: One Embedding Space To Bind Them All

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.158316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.158316Z digest=sha256:47785695e14238823062ff1232476daeced0d59b84f4ce647ec478159ccb0adf

Observation 6f95db44-06c5-461b-b93d-2fb0f6537abb · outbound

This paper cites The "something something" video database for learning and evaluating visual common sense.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text The "something something" video database for learning and evaluating visual common sense

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.161317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.161317Z digest=sha256:91cb0415f070a18ec1d902ec0e4814738c42c31ebcb5f5ae077957c8115692f2

Observation 89da694d-c378-4495-bbaf-48d4a02e93df · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Language Is Not All You Need: Aligning Perception with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.164357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.164357Z digest=sha256:f3b9043667e4c87db953069582db1f166b367527b3c1b6438c43663ab3887adf

Observation 2ed91dd6-bca6-42cd-a65d-adfc38ce79bd · outbound

This paper cites The Kinetics Human Action Video Dataset.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text The Kinetics Human Action Video Dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.167213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.167213Z digest=sha256:ab848099ed302afccd56de57c44054403c07ded547c167a92e323a5859a55754

Observation 2507dc6f-705b-48cd-8927-98653dea6df3 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.169934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.169934Z digest=sha256:921d411001d03c70956952230b657fbf6be2c3c3c846e046945d045ac8b76f0a

Observation 57a20458-e932-4b76-93c1-4b894ad55745 · outbound

This paper cites NowYouSee Me: Context-Aware Automatic Audio Description.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text NowYouSee Me: Context-Aware Automatic Audio Description

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.172477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.172477Z digest=sha256:61a4cc4a4d7837e518d9626ec5c5880d0d083bfab5596a01c686de0fbdba5bc1

Observation 0d52e13b-9464-4109-8aff-0babfde8d622 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.175428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.175428Z digest=sha256:c8ef6564f0bb334c870b94292a06fd0456e0ddee0500b2951489a0d2ae985e01

Observation ddbc04f6-4c79-4f35-b983-679adfe59a15 · outbound

This paper cites Kankanhalli.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Kankanhalli

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.178368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.178368Z digest=sha256:b19e340d7e210b583ed68b889d2e7a821d0cbc8620c20938181d358d9ae81834

Observation fa077860-7905-48e7-87c5-3891165760c8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.180993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.180993Z digest=sha256:f2dce9de579f7a3e918f727e9e6ca44c838b5c7d20c42243f493064a748536ba

Observation e5eb1154-0f5e-4f3b-97a6-4c520c7c4bb5 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.183821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.183821Z digest=sha256:458d2cf669464b814fedefebb2dbc0ff919a7d68c32a90210e66ef11efb11860

Observation 234b8113-a46c-43d2-a947-c42230ea32d9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.186559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.186559Z digest=sha256:6f2fbed0a12aeeb3b47d868d937cf589b34b5e234f7a7bdbb0b421bf672e1c6b

Observation 8279db1b-2f6a-461b-b949-46621b2e2bb9 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.189316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.189316Z digest=sha256:9a6b3262928801cec5bc330bc5cfd704df564a4e56b90882b8d6b2e75a320add

Observation c7417c79-a758-4d7a-af3d-68c16309da6b · outbound

This paper cites Visual Instruction Tuning.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.191264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.191264Z digest=sha256:35abca1a2f354ef76819423cac580fe63e4c199e156841b912ea480e65948d3e

Observation cef6858c-35bc-46f4-8f41-19f74e14c841 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MMBench: Is Your Multi-modal Model an All-around Player?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.193389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.193389Z digest=sha256:efbe9adf223104cd8159d4b1b8464f8130f33902fa9c0aadd154da55b896aed2

Observation 96e9d748-0816-492a-909a-bb559c1a83cb · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.195569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.195569Z digest=sha256:01850f54e84c1b9e5ec1522f9fccff55726401bf82fc4e16869f5f04dd0d22d3

Observation 6ec3c927-8318-4f6e-9ceb-87ba9fa6ae6d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.197598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.197598Z digest=sha256:2eac95f6088594db52399d3a2ae4e1fee0b94ad047ee20629e93ea844f5b2fcd

Observation caed6f36-bebe-43b8-9854-1642dccd3482 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.904681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.199634Z digest=sha256:9fc158b74d5f2eb026bbe4f0324441ca8b49fa7acbbd564fa26fcd05ab02d3fa

Observation c9274df6-1dd0-4f2f-ab4c-c2e9eca1e918 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.201536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.201536Z digest=sha256:611a10e35d76a86367e9e52b7636e9aec85d9899fab902a41ded522666f903c6

Observation a9b6ac0b-271d-4707-9b02-efa9b456f11e · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.204032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.204032Z digest=sha256:b3939c2f4c691be8cbc113ed7378a223eb69e0533661cc4eaed07825039596b4

Observation ea0ada39-3cd6-43c0-a629-3b7cd5c75c09 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Robust Speech Recognition via Large-Scale Weak Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.207064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.207064Z digest=sha256:56ce55dd15ae08d4620c3a841335fdbf04fd364aa761b1cea761cf6bdf15fd91

Observation d6ecd06f-889b-4095-9958-d7e4bc40c222 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 40

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-11T11:48:07.471391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.209868Z digest=sha256:cd2d9d8b5354a5c33106fe0e9ff8594b5cb3f741c8eed497bc62ebdbce2d7db6

Observation 556e9dc9-89ad-4f62-bd1a-07b39e84537f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.212367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.212367Z digest=sha256:2c24e6a518ea3ce3af02806dab2e200b5223ae6cc19e7cc23291c4c6587f5173

Observation 7d677305-df6c-43d0-a5ed-c76897f128eb · outbound

This paper cites AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:48:07.367997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.215478Z digest=sha256:59250be8c29e0f19c1c1cbf49becfc5670f5c3e79c66726ad816ff836d242c6c

Observation 9cfbe0f1-63d6-407a-8981-1ef078e0e4a7 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.218794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.218794Z digest=sha256:ab614fc8469eb5320366ddebbc23b6933d750b1750805b3e65736548a41ffb47

Observation 0ce6ac81-ac32-4b5c-ac18-878f9ed0616d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.222604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.222604Z digest=sha256:aa6dd04b340f90b3060877c8b4e126973ed9daea1b73432976eb6033be69e0ed

Observation a182f44f-cc8e-4c4e-b4c9-72f1d94c21d1 · outbound

This paper cites NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.225318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.225318Z digest=sha256:57a60e5d8e7ccd88144e3e64cbfb78d8d0a08fcd94b587155a7e35163024deaa

Observation cebcaa12-1d22-47e2-9e6b-74f8d28985bb · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text FunQA: Towards Surprising Video Comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.228234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.228234Z digest=sha256:dc9411893be1e9fd7e10f362bb04db27c2e7ff25094d2093a59d14157fd8df38

Observation 22177569-1448-45b4-a04e-406c37531f4a · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.895738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.231041Z digest=sha256:3bfb4aae78af9e33926f2babf60dee7ea05af993053bb754cda2dd74573ff860

Observation 543570b2-49ae-4006-8152-d932ba4c8155 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.233568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.233568Z digest=sha256:e8d9e9736a2ed4bb5764f11b8adcd092fc4b877cb1cd5461a76157a547dc63e7

Observation 0e2f40d3-0671-4b39-95d9-7cfe102ba48b · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.236231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.236231Z digest=sha256:fb689134a99202ae5021e18e3e2a65b13fca8c3b91a0488c0c12f98812c12b8e

Observation 0db9eb1f-fce8-40db-ab3f-3c67398d39fb · outbound

This paper cites Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.238929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.238929Z digest=sha256:db270277a615f40352df8acdbc0ab7bc76044d135aee459e538f9882e29b7fe0

Observation ee578662-6381-4d09-a55b-038b0767eb72 · outbound

This paper cites What is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text What is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.241651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.241651Z digest=sha256:94f93ef182a09bcfd1559e3fde7c4ff7d6265a56e35daf5693af12a2a5c5e461

Observation 3856e7d6-087c-4f6f-99e5-df0dc02fc371 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.244506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.244506Z digest=sha256:991b309ca5afab22c0a6b527968ea109df15fa196762922fa2ee181a45e33b9f

Observation 260020f6-a8cf-48d9-98b7-faa274e780c4 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.887203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.247202Z digest=sha256:e7f6237b14304b80d3ad9cd4198fc6522f351e8731ab60457be2d232486cb190

Observation 5325b058-cfc7-444b-8ad9-06d1bc24f3c1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.249685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.249685Z digest=sha256:2af567b22da0dea1f9270ff9a73b60f9ca492fb99537d99a02772d172dcde26b

Observation 98f1a2a0-8a7a-4c6e-8f91-c50e1ee5aad4 · outbound

This paper cites Distilling Vision-Language Models on Millions of Videos.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Distilling Vision-Language Models on Millions of Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.251910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.251910Z digest=sha256:359d256b739bd59f780e4e81b3d72a3fd42128643b5c2e907459ff98f6f06ac7

Observation 9afde8d9-d360-4f23-b161-c02b8ba82897 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.254058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.254058Z digest=sha256:2e79bf28a47bbe4913ae371b06b08a18eda31cf736cfae4be6f4d2bbbc5ccfcf

Observation e4705e8a-b71d-4bba-a84e-73e898a89bb1 · outbound

This paper cites online" 'onlinestring :=.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text online" 'onlinestring :=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.256287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.256287Z digest=sha256:7bf0f161628c374e8345c8ce6f88382c3af0c915ab229ab5d126566b398721ef

Observation ce3b0de5-1f1a-4ee4-8842-fb92cebbe38d · outbound

This paper cites write newline.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.259130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.259130Z digest=sha256:822d4967e3df54edf670df765e733f2481eca6449ab1200017d2c6f4cf3f694f

Pith citing papers

No inbound Pith citation observations are available.