Pith. sign in

Paper Citation Record · LEDGER

VideoLLM Benchmarks and Evaluation: A Survey

As of 19 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2505.03829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03829 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:10:57.491105Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0457cecf-8f81-4e45-bb37-3fb973083e06 · outbound

This paper cites Frozen i n time: A joint video and image encoder for end-to-end retriev al,.

VideoLLM Benchmarks and Evaluation: A Survey Frozen i n time: A joint video and image encoder for end-to-end retriev al,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.703017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.208332Z digest=sha256:c7b68e0a8ee2fe9d4cefc9df01935de82fe0febd71131e8311c4f83d63cffa2e

Observation 08461a35-ad1d-4559-8d15-2fd6d89cff1a · outbound

This paper cites Scale invariant feature transform,.

VideoLLM Benchmarks and Evaluation: A Survey Scale invariant feature transform,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.691096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.212724Z digest=sha256:06b040451987100ede4fc408729a4984c23b6a34332b2c95c66224914539dfc6

Observation f47cd2e7-f2a0-4bea-82c1-7e31ea624775 · outbound

This paper cites Speeded-u p ro- bust features (surf),.

VideoLLM Benchmarks and Evaluation: A Survey Speeded-u p ro- bust features (surf),

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.678497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.216370Z digest=sha256:bf408810a9be3e8d96cd5389254e158d598caa60802fa0c3a02b05ed88c8af99

Observation d9163e71-2ed7-4c94-a706-223aa3952cf9 · outbound

This paper cites Histograms of oriented gradient s for human detection,.

VideoLLM Benchmarks and Evaluation: A Survey Histograms of oriented gradient s for human detection,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.665742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.220205Z digest=sha256:d1a0c529dc411f55113bd05e8c2b8e32d28ca721def277178b5731f2383b9d79

Observation c413d45f-92cf-49bd-8da6-6fbfd18877e0 · outbound

This paper cites Large-scale video classification with conv olu- tional neural networks,.

VideoLLM Benchmarks and Evaluation: A Survey Large-scale video classification with conv olu- tional neural networks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.651998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.223893Z digest=sha256:b53d65ab1a5c20d5332d6f7f8baebc4e2e73aebe07cdbff40af79b04c27fdf68

Observation 3be2e011-c426-4941-bea1-b8889489d7e1 · outbound

This paper cites Convoluti onal two-stream network fusion for video action recognition,.

VideoLLM Benchmarks and Evaluation: A Survey Convoluti onal two-stream network fusion for video action recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.640500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.227558Z digest=sha256:83c93e8af6b6c73941ba51ae5126fe2e1c3620b10b9e1c09e67e839c0f562e8c

Observation 7a15fdf0-f206-4377-9ab6-0c3dfe524e7c · outbound

This paper cites Videobert: A joint model for video and language representa - tion learning,.

VideoLLM Benchmarks and Evaluation: A Survey Videobert: A joint model for video and language representa - tion learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.628095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.231096Z digest=sha256:4869211242ee92e9a30ad0f05e6db2e032c20f8001ef49d4a40b1de725807d4d

Observation cc9c0b22-75d0-42a1-9682-af326be4f67b · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervi sed video pre-training,.

VideoLLM Benchmarks and Evaluation: A Survey Videomae: Masked autoencoders are data-efficient learners for self-supervi sed video pre-training,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.615844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.235248Z digest=sha256:1bf8029851172954b9f107a6ef0654353dd02a1da910ca79e0e504c6bc9456a5

Observation b3cbbc0d-6946-45bb-b1f7-40f56455398b · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoLLM Benchmarks and Evaluation: A Survey Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.238581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.238581Z digest=sha256:e1ea170ee98a6f5086fb406db66149420c1ed0befc15879ee38d125848fcf1f4

Observation 41a4c077-2842-4e92-b2c1-95f2903988ef · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoLLM Benchmarks and Evaluation: A Survey VideoChat: Chat-Centric Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.242459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.242459Z digest=sha256:4a020bddc7f8b20c975c0652484f0e07658502ffc75e2629797611a6228b15bf

Observation 30905b2d-f513-41b6-a5b6-a3fb1bc79228 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VideoLLM Benchmarks and Evaluation: A Survey Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.246180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.246180Z digest=sha256:32cbcb3031faa44c0aa20827dfd5386923d1181fbae00c980c552e15714b6af0

Observation 7ffd5089-4f02-4cc0-9353-71b6d8632fed · outbound

This paper cites Video understanding with large language models: A survey,.

VideoLLM Benchmarks and Evaluation: A Survey Video understanding with large language models: A survey,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.250137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.250137Z digest=sha256:c61828a89d16acd7d0dac52adb7edfe72310a3208a3bf1d7d2816324d19b2224

Observation 9a52db01-2d9f-4b38-850f-5d805656a4a6 · outbound

This paper cites Video question answering via gradually refined attention o ver appearance and motion,.

VideoLLM Benchmarks and Evaluation: A Survey Video question answering via gradually refined attention o ver appearance and motion,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.602266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.253389Z digest=sha256:3067f247b92afd7fa7af14a0bb53dfbe7f6fb8b5d475b21a8af54cb14a79f5ff

Observation 0c4efde5-f176-4d14-9ce3-6a915556c71a · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering,.

VideoLLM Benchmarks and Evaluation: A Survey Tgif-qa: Toward spatio-temporal reasoning in visual question answering,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.590327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.256649Z digest=sha256:d82a3a6a6cc0ce61ded9ac237e6cf15cbae94fe3700ea523e0bd514c20d8f8bd

Observation 0af542e8-025b-4f0c-9876-975dac7eafbd · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering,.

VideoLLM Benchmarks and Evaluation: A Survey Activitynet-qa: A dataset for understanding complex web videos via question answering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.577932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.259980Z digest=sha256:9c85ea50f23a9ca5902117d36da429cf3fd435439e84d8412e1635e3c936b0f5

Observation 989723b7-7e4a-43e2-8960-6215ec44ad74 · outbound

This paper cites Tvqa: Localized , compositional video question answering,.

VideoLLM Benchmarks and Evaluation: A Survey Tvqa: Localized , compositional video question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.564643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.263310Z digest=sha256:fcbaad30f12e223941a4833df23e7b7e9fcb7411467e24642b4f4978d0e8567b

Observation a7584a48-4dff-44f8-b212-2f08c75fafcc · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

VideoLLM Benchmarks and Evaluation: A Survey Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.552326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.266599Z digest=sha256:c834e3f88c38a8f829c9d6b7b9c2a554a0bcc37afd49fb1e001fb7ed283fa8e2

Observation c5e18990-fa29-4d29-abe7-3ec904d7cc23 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VideoLLM Benchmarks and Evaluation: A Survey Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.270051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.270051Z digest=sha256:16ae4171b70279dc3066653189952e7bae9451afcc1910cb9787fe7ffb40cbf2

Observation 7827d522-bce2-400e-8df6-f4151b52fb7f · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VideoLLM Benchmarks and Evaluation: A Survey VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.273744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.273744Z digest=sha256:c66f912c2bbde6d85171105b864625d8cf2532324a54ae883291b269b1d589cc

Observation c93d0d5f-e557-431d-a3fc-a6eb6398a25a · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

VideoLLM Benchmarks and Evaluation: A Survey CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.278096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.278096Z digest=sha256:354cccbb193935eb06b44435a65cc3af63ca58b21d36b1a22537f6cb26051ac3

Observation 2f64fe3f-2e34-4a4d-a9d6-34680788f4f0 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,.

VideoLLM Benchmarks and Evaluation: A Survey Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.281659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.281659Z digest=sha256:ff03da0d9564c9c55189ba05014b032233303d63688a64c0ae3c703d161ac987

Observation e5c30a6e-22c3-471e-9472-771585ae9e31 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VideoLLM Benchmarks and Evaluation: A Survey TempCompass: Do Video LLMs Really Understand Videos?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.285229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.285229Z digest=sha256:92aa3cab5e3454a7653ded6fea93bc3dadfc013653775e5c4576483b7957ebeb

Observation 9af3a1e1-5d5b-44c1-aca3-b62ad2d63219 · outbound

This paper cites Her o: Hierarchical encoder for video+ language omni-representa tion pre-training,.

VideoLLM Benchmarks and Evaluation: A Survey Her o: Hierarchical encoder for video+ language omni-representa tion pre-training,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.539725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.288735Z digest=sha256:19b90726ae69f41fb22432a6eba708cc4bb2e53fe90b3b27d12059ebdfe0c873

Observation a38e6837-e11d-4e71-910b-0a895c294139 · outbound

This paper cites Star: A benchmark for situated reasonin g in real-world videos,.

VideoLLM Benchmarks and Evaluation: A Survey Star: A benchmark for situated reasonin g in real-world videos,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.527959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.292147Z digest=sha256:84c60d74ccacea8a53d92e25e87b6df77e91a57fefc5e9b629a1d0fd1d1bd3e8

Observation 2942772e-a2f3-4211-b5d8-6c2a222f5199 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

VideoLLM Benchmarks and Evaluation: A Survey EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.295597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.295597Z digest=sha256:453e61f85e4c884cd212b582089cbe4115dee7b832508c150ce29c68387b9fcd

Observation 53072aab-6734-4f68-9aa4-25c5fdc9f789 · outbound

This paper cites Autoeval-video : An automatic benchmark for assessing large vision language mo d- els in open-ended video question answering,.

VideoLLM Benchmarks and Evaluation: A Survey Autoeval-video : An automatic benchmark for assessing large vision language mo d- els in open-ended video question answering,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.515745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.299508Z digest=sha256:71abf27cacccc0bfd707c942dbd3ae05a602606f9725e6f4bcb5d3dd612c9578

Observation cdde5b75-7588-402e-a959-a055637c308a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoLLM Benchmarks and Evaluation: A Survey Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.302910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.302910Z digest=sha256:a410b2934f124d4909735d0e23a7529ed75c71cbb798cc8c68cdb49fc6e68485

Observation 1ffb5038-5031-4671-9b6f-8c6c32ed87b6 · outbound

This paper cites Sok-bench: A situated video reasoning bench - mark with aligned open-world knowledge,.

VideoLLM Benchmarks and Evaluation: A Survey Sok-bench: A situated video reasoning bench - mark with aligned open-world knowledge,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.502546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.306568Z digest=sha256:ebd50e1160a6f9f24032cff5eca414b560b19570569504f1220de21a680f32e2

Observation 96461d7b-a7ae-465d-9552-ea619cb64dde · outbound

This paper cites Long Story Short: Story-level Video Understanding from 20K Short Films.

VideoLLM Benchmarks and Evaluation: A Survey Long Story Short: Story-level Video Understanding from 20K Short Films

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.310180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.310180Z digest=sha256:8d18be8eeeee10834f7cd34f6985f4e594cbf3bd64eb6c757244764c530f37ce

Observation e7847b22-1578-4fe9-9df3-5e32153a7437 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VideoLLM Benchmarks and Evaluation: A Survey MLVU: Benchmarking Multi-task Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.313889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.313889Z digest=sha256:922c2d10231b347f0a9917d297e2f59edd30f7d01c23f3ed65ff7eaaace7f140

Observation bb73934b-8e76-4fc1-ad34-f7ff3b48399b · outbound

This paper cites MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos.

VideoLLM Benchmarks and Evaluation: A Survey MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.317505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.317505Z digest=sha256:4dc07b3c347ca28d3ad6e776d954813dde1beb1ef1a77453941bf010ac6f0ff9

Observation 26e35818-4686-46ba-bc00-b8fa21d25181 · outbound

This paper cites VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment.

VideoLLM Benchmarks and Evaluation: A Survey VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.321005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.321005Z digest=sha256:42ee90312b205239d53189d0544f40f2192522c997213554b5408bca04bfd5ec

Observation 3f093539-0a41-4e41-8255-28fbf86fa408 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal action s,.

VideoLLM Benchmarks and Evaluation: A Survey Next-qa: Next phase of question-answering to explaining temporal action s,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.490657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.325091Z digest=sha256:babeead79343e469566ff5b0dda979529c6eb2b74e286eba803dcb006ded1799

Observation 4ff75dd1-8b15-4161-9509-23de8554df3d · outbound

This paper cites Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model.

VideoLLM Benchmarks and Evaluation: A Survey Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.328774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.328774Z digest=sha256:fd0fd09ef0017d3a8ed87b2d3b1fba5b0f51bd533e8010d2ee9cb3d1ba636c5e

Observation a0758e65-dd60-454b-98ff-1f843bd22dc0 · outbound

This paper cites Cider : Consensus-based image description evaluation,.

VideoLLM Benchmarks and Evaluation: A Survey Cider : Consensus-based image description evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.479156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.332438Z digest=sha256:551fab5a40c466ff9f86b09ac4b101f5b9171dd303c5419677555475a98b7b53

Observation b2f1fe92-9012-40ad-9d96-a9bd2b7d64bb · outbound

This paper cites Meteor: An automatic metric f or mt evaluation with improved correlation with human judgments ,.

VideoLLM Benchmarks and Evaluation: A Survey Meteor: An automatic metric f or mt evaluation with improved correlation with human judgments ,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.466039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.336309Z digest=sha256:ebebc6f393db22a8b0679e29821a3586077abda0d3010dd4e6cb228316a71116

Observation 8aaee166-238b-4fa2-97d3-7415a0172be5 · outbound

This paper cites Rouge: A package for automatic evaluation o f summaries,.

VideoLLM Benchmarks and Evaluation: A Survey Rouge: A package for automatic evaluation o f summaries,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.453481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.340100Z digest=sha256:09febe83a3bc4ce8de467aeed89e200fb65317cf688fa482165c12d9591a0a18

Observation 10154545-9f00-4c87-ac65-3385c473647a · outbound

This paper cites Spi ce: Semantic propositional image caption evaluation,.

VideoLLM Benchmarks and Evaluation: A Survey Spi ce: Semantic propositional image caption evaluation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.439700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.343546Z digest=sha256:981cf9bed2039307a48a07488165fa8750891b9962f7eab2ff63ea58e21d9047

Observation fd49f749-fda1-418d-b0aa-5a11e99e2bac · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VideoLLM Benchmarks and Evaluation: A Survey PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.348050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.348050Z digest=sha256:431fd4bd001f225fff763b4a6cf9056616009ae80ac45c7ef733b733dd7352a6

Observation c8bdc7b7-d4f6-42e8-a55c-a5f99ed9d973 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vl m,.

VideoLLM Benchmarks and Evaluation: A Survey An image grid can be worth a video: Zero-shot video question answering using a vl m,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.426279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.351876Z digest=sha256:f046a8974bead69304155f716fbdc262284e9b39cb0482c376c06ca870637e73

Observation 1551d931-4d42-4c77-a1fa-faee1b5e021d · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

VideoLLM Benchmarks and Evaluation: A Survey ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.360169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.360169Z digest=sha256:b6074ee7d543d612319e4062640261493ded44971a47182f7e655f130f4ecc68

Observation a25d3613-5ba0-47bf-8b3b-1bf2ccbb4ac0 · outbound

This paper cites Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search.

VideoLLM Benchmarks and Evaluation: A Survey Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.363973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.363973Z digest=sha256:d352d1ab77a8a1aa054123e72af12290aa8d3d404689c972f66adfd4a27ae21e

Observation 97e10699-b042-47b8-bdf9-ceb03bcaac73 · outbound

This paper cites Prompting visual-language models for efficient video understanding,.

VideoLLM Benchmarks and Evaluation: A Survey Prompting visual-language models for efficient video understanding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.401414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.367703Z digest=sha256:fd779fcc6d3d659bd53f81366abbb51e73263d17fa78fd744d11d3b67edc949c

Observation 2c694bcf-fb46-454a-a33d-b2ea7fe61535 · outbound

This paper cites Vid2seq: Large-scale pr e- training of a visual language model for dense video captioni ng,.

VideoLLM Benchmarks and Evaluation: A Survey Vid2seq: Large-scale pr e- training of a visual language model for dense video captioni ng,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.387048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.370935Z digest=sha256:225b1d900798e8cfb8152ef86e5d7adfde380133b1b26eb86a69a923ad856b80

Observation 7a2d283b-0723-4264-bda1-06c3f4fa4d38 · outbound

This paper cites Lea rning video representations from large language models,.

VideoLLM Benchmarks and Evaluation: A Survey Lea rning video representations from large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.375575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.374152Z digest=sha256:fd8fa5460ceb25b47d8397325a895665cebe79043d7682f7fcf0df80272dcb87

Observation dd2bf198-3261-463e-9b15-da47f779f920 · outbound

This paper cites Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment.

VideoLLM Benchmarks and Evaluation: A Survey Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.377738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.377738Z digest=sha256:73728e63c1174d1a70d9e161d7f069d8b26e3ddbc17f34c3380104c28039eb3f

Observation cff9b54f-1fd1-4f1a-b1c0-b7d5f4637841 · outbound

This paper cites DiffusionRet: Generative Text-Video Retrieval with Diffusion Model.

VideoLLM Benchmarks and Evaluation: A Survey DiffusionRet: Generative Text-Video Retrieval with Diffusion Model

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:10:57.621559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.381159Z digest=sha256:0507ab6d04b0541a35aeb6572c8deec4c987e12c7ea8584a6eb0e6575eaf1784

Observation 632fe7bf-dc07-41d9-bc93-a9f8f9c48c9d · outbound

This paper cites Egovlpv2: Egocentric vid eo- language pre-training with fusion in the backbone,.

VideoLLM Benchmarks and Evaluation: A Survey Egovlpv2: Egocentric vid eo- language pre-training with fusion in the backbone,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.354358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.384947Z digest=sha256:1927cd6ee4f494cd3365a32331cf9174a9ca3c3a4276dd6f59fe794bfff68566

Observation a5c85418-434c-423a-b6db-6a8a8141fbf2 · outbound

This paper cites Large Language Models in Education: Vision and Opportunities.

VideoLLM Benchmarks and Evaluation: A Survey Large Language Models in Education: Vision and Opportunities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.389050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.389050Z digest=sha256:f620add5c9a25866af71923d004369ee0117611f7709ea9ed4888be2a7c6828f

Observation f6de4c2e-beb7-459c-bb93-36834d7552ca · outbound

This paper cites A Survey on Deep Multi-modal Learning for Body Language Recognition and Generation.

VideoLLM Benchmarks and Evaluation: A Survey A Survey on Deep Multi-modal Learning for Body Language Recognition and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.392691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.392691Z digest=sha256:1d370608ae84e21da2eb56a476a3a6edc2a7543348cfd5176a3ff8ce7e480b7a

Observation 8c8c58cb-9c19-452d-b9ac-4f52344f1e7a · outbound

This paper cites Machine translation from signed to spoken lan- guages: State of the art and challenges,.

VideoLLM Benchmarks and Evaluation: A Survey Machine translation from signed to spoken lan- guages: State of the art and challenges,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.339055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.396186Z digest=sha256:4da506ba1e3acbd17434479be68bb2bf71f57b6aa9916d0cc9076221e48562c8

Observation 6a583ddb-c0bb-45c5-af0f-16802d2cc2e7 · outbound

This paper cites Generating video game quests from storie s,.

VideoLLM Benchmarks and Evaluation: A Survey Generating video game quests from storie s,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.325720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.399555Z digest=sha256:266e2a06c65ea3b66d68184a46f41bfe10f740d2da7f1e31ce949fabd4a319a0

Observation 2b6d144f-71e8-4bca-baf9-e27bb6d3d6d5 · outbound

This paper cites Text generation for quests in multiplayer r ole- playing video games,.

VideoLLM Benchmarks and Evaluation: A Survey Text generation for quests in multiplayer r ole- playing video games,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.313403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.402915Z digest=sha256:290a9ea5b9e9f80946637cc19a1e7cae871f28448c00337953ab0c7dd74bfa91

Observation fa00f981-94d5-49b8-bf36-108fb955f490 · outbound

This paper cites The role of artificial intelligence and robotic solution technologies in metaverse design,.

VideoLLM Benchmarks and Evaluation: A Survey The role of artificial intelligence and robotic solution technologies in metaverse design,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.301198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.406452Z digest=sha256:19beff4d9c1842206f6f6681ea3f4c831c623d7c0ae8f8ac0e109810c1b271b1

Observation 38d1517d-d868-411a-93f5-6e9df3ab346d · outbound

This paper cites Jung and M.

VideoLLM Benchmarks and Evaluation: A Survey Jung and M

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.289191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.409962Z digest=sha256:0737c212fa367219ab6199162d23a53e2a44636600fda29bc043cf4143e9cbf6

Observation 03b6e04c-908e-4e08-8cd0-44e0d4ecfe61 · outbound

This paper cites PromptFix: You Prompt and We Fix the Photo.

VideoLLM Benchmarks and Evaluation: A Survey PromptFix: You Prompt and We Fix the Photo

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.414062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.414062Z digest=sha256:9bdc134d4612aa3874a5626d038b08aaa241abd01b6a60f4c4d5b022b1536732

Observation c134139c-ac6a-4450-98f7-d5590369b4a2 · outbound

This paper cites Egocentric audio - visual object localization,.

VideoLLM Benchmarks and Evaluation: A Survey Egocentric audio - visual object localization,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.277305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.418116Z digest=sha256:9dd90bb17e5f47fc9eea139ccccea4c1f411d3d88d30560c407f83a2c9da62f2

Observation abc99f2a-1a97-4173-b087-98960b9c1f2a · outbound

This paper cites Misar: A multimodal instructional system with augmented reality,.

VideoLLM Benchmarks and Evaluation: A Survey Misar: A multimodal instructional system with augmented reality,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.264739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.421836Z digest=sha256:6f835e5d63373ef7699a8a42de11220e61192ad11ec2ba34c219a3b30d244c5a

Observation 123e926d-2a69-4728-9ee5-95ae57077aa7 · outbound

This paper cites Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning,.

VideoLLM Benchmarks and Evaluation: A Survey Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.252034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.426486Z digest=sha256:5a51277d854b8541589913118c1e7458fe24af15fe84efab854ae663168220f9

Observation 84041543-f715-41ec-bf3a-d5a94720c217 · outbound

This paper cites The role of chatgpt, generative language models, and artificial intelligence in medical education: a con- versation with chatgpt and a call for papers,.

VideoLLM Benchmarks and Evaluation: A Survey The role of chatgpt, generative language models, and artificial intelligence in medical education: a con- versation with chatgpt and a call for papers,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.239305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.430221Z digest=sha256:f4f0538188ab326a7109aa6727bd6beff5e507945a2abba86fee508cbf9f3b73

Observation b699bdf1-e3ec-4f01-b360-88270bb966f3 · outbound

This paper cites Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesi s,.

VideoLLM Benchmarks and Evaluation: A Survey Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesi s,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.225749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.434122Z digest=sha256:39a1c97bd60690a8ad4a7fe349ed64cde7678313120780eac62461c36a6c70c1

Observation f21a76c5-1eef-4af4-8019-227a4c0556d3 · outbound

This paper cites Disco: Disentangled implicit content and rhythm learning for diverse co-speech gestures synthesis,.

VideoLLM Benchmarks and Evaluation: A Survey Disco: Disentangled implicit content and rhythm learning for diverse co-speech gestures synthesis,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.212875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.438175Z digest=sha256:f8a085d5a9a14d51112187e09b058242fdedfcfbbba912ff24d80a712495debf

Observation fb50e9a4-e2e1-4316-bc52-5881bca592d5 · outbound

This paper cites Emage: Towards uni- fied holistic co-speech gesture generation via expressive m asked audio gesture modeling,.

VideoLLM Benchmarks and Evaluation: A Survey Emage: Towards uni- fied holistic co-speech gesture generation via expressive m asked audio gesture modeling,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.198635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.442033Z digest=sha256:8158a1225c7904531e54e453f7acb9830ba7e45b3622eda5ab4e35b284c87272

Observation 087c594f-de19-4643-8fa2-ac170830ed03 · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

VideoLLM Benchmarks and Evaluation: A Survey LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.445714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.445714Z digest=sha256:5d6e5060b511b8d974e2ef488d6004172e4b65b213c5e3e96e4267fa5fd4f7f9

Observation 06414b93-953f-4053-9c7d-8f9bec932fbf · outbound

This paper cites Chatgpt for cybersecurity: practical applications, challenges, a nd future directions,.

VideoLLM Benchmarks and Evaluation: A Survey Chatgpt for cybersecurity: practical applications, challenges, a nd future directions,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.185033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.450162Z digest=sha256:b365b4c143d915b9696001469139714506ae1c566fbb7cc55d477a1ba811ca92

Observation 0c1196e5-1cc2-4c52-bc9e-6e5129f94096 · outbound

This paper cites Modelling language for cyber security incide nt handling for critical infrastructures,.

VideoLLM Benchmarks and Evaluation: A Survey Modelling language for cyber security incide nt handling for critical infrastructures,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.171783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.454103Z digest=sha256:a780584163938bd04216c4c83008eb4f3d75ee2992cf86ebfae48375752ba0f8

Observation d9c89dbe-87da-40c5-965a-458ee2a764e3 · outbound

This paper cites Socratic video un- derstanding on unmanned aerial vehicles,.

VideoLLM Benchmarks and Evaluation: A Survey Socratic video un- derstanding on unmanned aerial vehicles,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.159088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.457788Z digest=sha256:a2d3ba9f5b702a98fe54c8ad43ad868399aaa5a4e7599a1018c65d62494ed71e

Observation 337893e5-e600-47d2-901a-5c5eb9b5bce4 · outbound

This paper cites Lanobert: System log anomal y detection based on bert masked language model,.

VideoLLM Benchmarks and Evaluation: A Survey Lanobert: System log anomal y detection based on bert masked language model,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.146125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.461676Z digest=sha256:5a3ce49b5bf064655af59a21c5902975e94cac4a49d6123769f6a139343a5ff6

Observation 9e006d68-503a-4a62-846d-b3740f65e80e · outbound

This paper cites Logfit : Log anomaly detection using fine-tuned language models,.

VideoLLM Benchmarks and Evaluation: A Survey Logfit : Log anomaly detection using fine-tuned language models,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.132101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.465491Z digest=sha256:f5b740074314d098825f1ae2721a4d3ca448f3135a2fe223eb9aa09e4902f37d

Observation 51c40cc3-ebe4-498a-bd32-9c912e6b9c82 · outbound

This paper cites GraphGPT: Graph Instruction Tuning for Large Language Models.

VideoLLM Benchmarks and Evaluation: A Survey GraphGPT: Graph Instruction Tuning for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.469272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.469272Z digest=sha256:a57aeb5b6a3d60c1383688306fd0fe12be7805fa04cd03ff426e2ae89c388abe

Observation 16a4d9ce-1d6a-4f39-92a9-8840ec38568d · outbound

This paper cites Drive as you speak: Enabling human-like interaction with large lan- guage models in autonomous vehicles,.

VideoLLM Benchmarks and Evaluation: A Survey Drive as you speak: Enabling human-like interaction with large lan- guage models in autonomous vehicles,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.119702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.473276Z digest=sha256:929ad49d940917a113d4f97662c3d8cec863516803abd3780f425dda0acb3b89

Observation 873ee204-4992-4a7a-922b-a62edde51234 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

VideoLLM Benchmarks and Evaluation: A Survey Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.477036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.477036Z digest=sha256:5882e9aac7923af3f9cc4ce9f998a83c42f8107e876b95153d5827d4ce6e48c5

Observation 16568131-7681-421b-bfb1-1c9a22c477e4 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

VideoLLM Benchmarks and Evaluation: A Survey LISA: Reasoning Segmentation via Large Language Model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.480975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.480975Z digest=sha256:44d33b63eb3c07968265add839b0fd04ba7ef913f466cc2efeaefa22cf03d798

Observation f2741f88-e54a-40dc-9562-72d60aef8840 · outbound

This paper cites How retail video analytics enhances cust omer experience,.

VideoLLM Benchmarks and Evaluation: A Survey How retail video analytics enhances cust omer experience,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.106588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.484416Z digest=sha256:efa774b05219284f6c06a17f26fa924708e6284f53bc6961431257742f27aa59

Observation 720ed770-7920-462c-9723-62462dbb55b4 · outbound

This paper cites How video analytics is transform ing the luxury retail experience,.

VideoLLM Benchmarks and Evaluation: A Survey How video analytics is transform ing the luxury retail experience,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.093348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.487821Z digest=sha256:e7c33742198b279720dc9ac72069e9310270111a16681f1cfb11738396863b83

Observation 83ca2291-8e1f-48e6-aa14-dad7416d4a75 · outbound

This paper cites Visual search and its growing influence on e-commerce,.

VideoLLM Benchmarks and Evaluation: A Survey Visual search and its growing influence on e-commerce,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.079736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.491105Z digest=sha256:7c15064c2a171f4a9e10a49b2b0faa08cddd023dc4dadc9e2095f566c3bc1919

Observation 34ed7320-45fb-4f0c-8dae-03e4858cd505 · outbound

This paper cites Available: https://arxiv.org/abs/2403.

VideoLLM Benchmarks and Evaluation: A Survey Available: https://arxiv.org/abs/2403

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:10:58.414170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:10:57.356168Z digest=sha256:6e34ddd5214c4f970f60a2f306384c1ef6fcb57faa9ab92b8812b57c0ce3d003

Pith citing papers

No inbound Pith citation observations are available.