Pith. sign in

Paper Citation Record · LEDGER

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2507.17050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17050 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:21.032754Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:13:38.580739Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eba5d90d-1e6b-4e03-803b-235051e41324 · outbound

This paper cites Optimizing marketing strategy: a video analysis approach.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Optimizing marketing strategy: a video analysis approach

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.300597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.947571Z digest=sha256:280e018dd2fa35cc0efb56e227fa9faa4894b67125de5443aef5a380d19e25a2

Observation 3086b80a-7c54-4933-b054-130272eabcec · outbound

This paper cites METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.294481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.950304Z digest=sha256:41756d6fa7c650c61724344755ad27677e356fc879ef146e221a721a600f9be1

Observation f63121dd-c975-455f-a460-2afcb70cbe12 · outbound

This paper cites Mm- au:towards multimodal understanding of advertisement videos.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Mm- au:towards multimodal understanding of advertisement videos

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.288029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.952675Z digest=sha256:989667daad30202f2ddca047620fff8611f83dfe07b5d935688a5dfca660e2ae

Observation 117d45b5-b834-4480-99bb-a65dc24eff5f · outbound

This paper cites Blog: E-commerce product videos with examples, 2025.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Blog: E-commerce product videos with examples, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.278050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.957283Z digest=sha256:6be5079b1f8224f977bd0b1dc74afead710f1a8fbfb2f1d630d99e3608d38ca4

Observation 727b0d93-9bae-42ff-bee5-0021c6fa1fa8 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.959349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.959349Z digest=sha256:a8be248f20977de7aa60262b2332e65dd7956aac0376b3ce68e4405b93d2f5f6

Observation 2f87acd8-6b4d-4b05-8db7-23d0659e1b2e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.961859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.961859Z digest=sha256:4b0a786284145285e6a23709552640708fd4bf5269a432fe600b2df1b8c38ee4

Observation f5ad6c06-6171-41f0-a505-2ca5f57f794a · outbound

This paper cites Yolo-world: Real-time open- vocabulary object detection.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Yolo-world: Real-time open- vocabulary object detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.268146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.963956Z digest=sha256:83aa07395ea5f774333458d70c83e4eab28c2d118216a0876148459ffc5d1a4c

Observation b71f51b6-6e6b-4d98-ab16-c62e5e6ae583 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.966095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.966095Z digest=sha256:3b7c23f859dea0ab2e3410b17e293098bfe2cc480327c4fbeb80554e78f29f52

Observation 958d8f7b-fda2-4189-b8d7-be08f4085284 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.968013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.968013Z digest=sha256:d1f61ac324bb72b5c7607c274655e49eddd10f170e3581be5a328d3e3e017558

Observation 612a4483-0884-4de4-ba48-09f93e430382 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.969999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.969999Z digest=sha256:43f94d0d509ce7ac4f7b4997070b8be3a8a7ddb2dc162985f0a87d682c03f829

Observation 60b03cf5-3ed3-41be-a131-eb254f80051d · outbound

This paper cites Sketch, ground, and refine: Top-down dense video caption- ing.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Sketch, ground, and refine: Top-down dense video caption- ing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.254043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.972296Z digest=sha256:c0619be8bdb33fba6353607707c1045cc9995b669f3ac1bd32d4345290142420

Observation b451ef04-ecc6-4ebf-9417-15dd9c8e62e9 · outbound

This paper cites Video-mme: The first-ever compre- hensive evaluation benchmark of multi-modal llms in video analysis.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-mme: The first-ever compre- hensive evaluation benchmark of multi-modal llms in video analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.247827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.974389Z digest=sha256:6ecf83ce16b8bfa1863b08d851d78157f7c3bd478f4b395db2c299c5c82adcdb

Observation d38fc442-b07b-49da-ab2f-5ce2082f2202 · outbound

This paper cites The Llama 3 Herd of Models.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.976339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.976339Z digest=sha256:a7df58e536a01a119b20ed92cc4ec14b21d7e0983be75a1a71a8fd29976be463

Observation abe00bd2-dd7f-454b-aa97-67645fbf4b99 · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.241390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.978448Z digest=sha256:f9530367a573e8ec5e8779331024db57d2cea81446a5fc952948bfd5135e4db6

Observation 5adda0b6-68c2-4ad5-a57d-064b04934878 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Lita: Language instructed temporal-localization assistant

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.234775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.980333Z digest=sha256:841817df9fa6fdd0fccae3a2b77d014a2f7e72831d7be913cca63772c26637b5

Observation 7265ce93-0b36-4462-9633-cd40092fb2a0 · outbound

This paper cites Automatic understanding of image and video advertisements.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Automatic understanding of image and video advertisements

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.228575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.982297Z digest=sha256:3a4e56d388d0332a2a2140df163f06cf70c8024579d487bb56f16cf8661a15e9

Observation 2f3b36da-67d6-4e87-aaa5-961b35cebf3f · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.221739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.984436Z digest=sha256:8df2b01caa209781a8fafdaf1d7c49d60f68e57059bf6fc36e7441d172831096

Observation f95af62b-679a-4149-a605-ae30d0115040 · outbound

This paper cites Multi-modal dense video captioning.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Multi-modal dense video captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.215221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.986346Z digest=sha256:ced23d16453bdf8acc6c9ce5b97482bdce15b8c70649d23b0b3ee36492e47755

Observation d56f9c4f-cbfe-41c8-90e9-57f9c5b9c2de · outbound

This paper cites Video re- cap: Recursive captioning of hour-long videos.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video re- cap: Recursive captioning of hour-long videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.209047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.988207Z digest=sha256:f7efa525d9c8ce19991e2f4f81ead6bdcce4a3b2079b538e762e9b0a6865cb6e

Observation 053435b6-fb70-4682-bfb0-3a5909386f8b · outbound

This paper cites Drive targeted marketing in retail with video ana- lytics insights, 2022.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Drive targeted marketing in retail with video ana- lytics insights, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.202857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.990066Z digest=sha256:490525a9cf6572a3fe57abca8f25ef5aef265db2821e609e08893c48e81cc74c

Observation 16221969-ee3e-4131-ad09-1d1227d2a7ed · outbound

This paper cites Dense-captioning events in videos.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Dense-captioning events in videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.196759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.992028Z digest=sha256:96f6a8842d26a5c4bbba319bcd2ba39e1f1367c37259ca6daf79891b27729f65

Observation 2c9170b1-69de-47f5-8039-3b8022545839 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.993917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.993917Z digest=sha256:e5197978c524ce310f3846baf7eca6c61591100cc995677c2158ba2ae0db9422

Observation c4d10611-336c-473e-b65e-c957ca7e16e1 · outbound

This paper cites Development and challenges of object detection: A survey.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Development and challenges of object detection: A survey

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.186367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:20.995737Z digest=sha256:fab00217a0b26ac7c525c842abcef6ffdacf2e6140973ee02d7d8c5dc3dc4be4

Observation 4a99b57e-5102-42ef-87f4-15f547d55e79 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.997595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.997595Z digest=sha256:e66361ef5f4097a0f091a5b34e120cfaf2c1def33c0487bf9e85cd1e1ef09734

Observation a2b000a7-c7a4-4477-9912-9fb4940a4473 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:21.000083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:21.000083Z digest=sha256:e4016adbbd83527c22707813fe608dd37689b75a4c2ab789ecfad8d634f5a1e2

Observation efacaea5-5a1c-49bb-a7b8-7814f71bca7a · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:21.001977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:21.001977Z digest=sha256:e0fa02fc74cda4bf36e6c0acf5ac2731620da90770f32f3c2cb7f67e6147c978

Observation 5ca176c6-65d6-473b-9254-22bbb358aa98 · outbound

This paper cites Dense video captioning: A survey of techniques, datasets and eval- uation protocols.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Dense video captioning: A survey of techniques, datasets and eval- uation protocols

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.175346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.004396Z digest=sha256:4255ae3602f9c0e57e299b01b28704cff91eade6c55b8fbea1ab77b1ca4293e1

Observation 01df04a6-cf88-41fd-94cc-01ab46d00a6f · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.169008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.006348Z digest=sha256:45c1ce9d0e6fad3d6dd59689bff95ae6b0da2a18532aed37c0dbb7de4958b323

Observation 6e6dc0b6-d181-47ce-b5d2-eec2163aa5e3 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:21.008633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:21.008633Z digest=sha256:df5887059a251e198f84984bb3fbd70a1ceff8e1bf057c0d1398a7a9a5eba2b3

Observation 66e6dad2-3916-4c63-b8ca-151b68520dbd · outbound

This paper cites Video understand- ing with large language models: A survey, 2024.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video understand- ing with large language models: A survey, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.158816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.010558Z digest=sha256:6544c28353f1496dd1399f6dd53f69c946299ecf57d17afee49aab46bf8e2bf9

Observation 09ebf718-9af9-4545-a0fb-85318b5503e2 · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Llama: Open and efficient foundation lan- guage models, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.152894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.012655Z digest=sha256:140c6cbaad60919174f6ef37aaa323ee883489b72f5eb37bf0227be3b7ba4bab

Observation 430ded0d-8fda-4b06-a5c0-369f4fd5360e · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Cider: Consensus-based image description evalua- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.146944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.014579Z digest=sha256:6863ef4212d6cc389d922fbcb4983849edc343ef7de1250f2af514b59785d362

Observation c38a615f-9637-465c-956a-0d0a52776827 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.141023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.016571Z digest=sha256:845898bda2d35b10115a88f3b058aedc76d3d36f236792ded658cc4e5862ceed

Observation 4476f321-8a75-4258-8440-923db242c431 · outbound

This paper cites an unresolved cited work.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:21.133792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.018390Z digest=sha256:0ed40eb14d8b8a471909a4373336d568f83ed5dd073ce226da8d6a609b2dcf26

Observation 46b28dad-b922-49d2-972b-75cbb8de2139 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models End-to-end dense video captioning with parallel decoding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.125984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.020652Z digest=sha256:9f097c1fef4ace84ce26a34b0a2e8b20a4428afafd2667b832905b49d7a1fa08

Observation 0b53ed68-c29a-4201-99dd-3bd8ca64f195 · outbound

This paper cites an unresolved cited work.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:21.118602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.022609Z digest=sha256:3cbc184fb5d92108ca135e6b51c78c5c6ee3a7789eaaf303060523042482ae2c

Observation bfae3521-0115-4d23-939f-5ef65fa177af · outbound

This paper cites Advise: Symbolism and external knowledge for decoding advertisements.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Advise: Symbolism and external knowledge for decoding advertisements

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.111473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.024520Z digest=sha256:008340b8a35e54cbaacf89b494d3f48cf67d29d17b650b067a7a02f0f0ee4ff1

Observation c227290e-e616-4373-acf9-30ac7471b18c · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:21.026368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:21.026368Z digest=sha256:dbf04a510da85f77ebfe513eb27339a5e6ee24f2e5a971eff9668662f0e20c89

Observation 292164a1-f20b-40fb-be92-cb7a07e7acf3 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models BERTScore: Evaluating Text Generation with BERT

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:21.028640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:21.028640Z digest=sha256:6459cf7cc43ccf351de318f77ddab29802a0062f33aa6dee506372db9d9922ef

Observation 65f006a9-25be-4729-ba27-caed35ba339d · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Towards automatic learning of procedures from web instructional videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.104987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.030741Z digest=sha256:aa7b01bbae0d4497b12057a5b731640c6be3278c3c1a6cd9781a57882d26e417

Observation b76d694a-719c-482f-9df9-111cba72f2d2 · outbound

This paper cites End-to-end dense video captioning with masked transformer.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models End-to-end dense video captioning with masked transformer

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:21.098048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:01:21.032754Z digest=sha256:9c17664613f84bbdaaf728d60055372664aa9b61d60cadc890c98d4803b261ce

Observation 4bbfa2c7-ae93-4542-a974-454ca2e2bf22 · outbound

This paper cites an unresolved cited work.

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:20.955209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:20.955209Z digest=sha256:d51656564f33eeed8841de21088ad1c34386f13c728fe5e33d500569f27c596a

Pith citing papers

Observation 70d83a25-4c97-4471-8c44-fd8ede4213ab · inbound

SERUM: State Extraction and Refinement for User Modeling cites this paper.

SERUM: State Extraction and Refinement for User Modeling Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-03T12:13:38.580739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:13:38.580739Z digest=sha256:b21779bfeced01ae38d5cb46d82ecd8ca1a175d91e452ad4812751d11e278154