Pith. sign in

Paper Citation Record · LEDGER

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2505.23693.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23693 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:50.619699Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T14:51:33.951351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:20.413167Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e62cd0b2-3720-49ec-8dea-00813429eb42 · outbound

This paper cites URL: " 'urlintro :=.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.741569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:43.741569Z digest=sha256:951e69d448e646abf75f25f5c3f0213213797b0f74073216203f1802bc2947d3

Observation e3ec7d4b-9381-4207-8114-bb0c96a8d985 · outbound

This paper cites write newline.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.846912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:43.846912Z digest=sha256:f04331a719e327637c766b7881329ab1a01f0c5f6c3d0caad3b3a2061a45cfdf

Observation 3d689502-d095-41c7-a03b-c2b128bf3b9f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:43.962180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:43.962180Z digest=sha256:211b34784450f429769d4430202bd2cd1b6fb37ee8c01371651d14b7ba882dea

Observation 533b0339-a776-4d3c-8466-f712a75eaeb6 · outbound

This paper cites Qwen2.5-VL Technical Report.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.067256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.067256Z digest=sha256:68626268a61d80097649b4ac4567e7a6bbd1e0c818dd69d4de90aaa0bfd5c2d8

Observation abf7a0ea-213f-4d0d-87c1-6d9500376319 · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.170543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.170543Z digest=sha256:9ef7c0ded9399b2ee8ce8d670ee4e668a2681ec9c55693fba7370c98cc320f8d

Observation 939add0e-cde7-4d97-a8f7-4d8a88594e64 · outbound

This paper cites Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:52.476729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:44.316344Z digest=sha256:a616a2255fe4acfb940b08565cfa15b87489191c51633557f9ba49c7b7c8304f

Observation 5872b0d7-1f29-43f8-805d-5331c4a711d8 · outbound

This paper cites ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.438501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.438501Z digest=sha256:b79d06b23a1fe5391787e6d2f86dc0212c1d258d32b98267188e42ad95c86eb1

Observation 48ded090-d928-45ea-8b60-f677fc3e38e5 · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.538343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.538343Z digest=sha256:c745092a47d41b377dcf049ede705df9ad5d4bde6bf7e1f44ec9e1c9aadd713e

Observation 931fd778-7db9-4656-9d6d-7386d5d7549f · outbound

This paper cites The Llama 3 Herd of Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.679138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.679138Z digest=sha256:b99ea289e156a19d15d4c6e3a7575589a6fca2fdb85e73a6f558e9847e21daec

Observation 3000991f-59c4-4821-afbc-381b73476d7a · outbound

This paper cites AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:52.259109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:44.801085Z digest=sha256:40b6947a6806d7b2fdf0fbfa675f4b4eb27103624de013414729da49997522d9

Observation 2b773435-f45d-44d3-a2a5-28ef34363b23 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.920605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.920605Z digest=sha256:e8d7f15627320518f9de2a08d4961becb7fecd28e57782520a035914e96bf3bf

Observation 6db424d1-fead-4382-9b1b-eaac116a4e15 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.015168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.015168Z digest=sha256:e8b5bc62f56e9f135edad9fabce5362712ab130951af857a7a38fc09f91c83ac

Observation c70e6167-669e-4b4c-a7cd-9d330fad562b · outbound

This paper cites LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.113785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.113785Z digest=sha256:fd381ccd9b9fa3394d859b4bd4080d7fa0bd4a8cef9c34717efa4db2168e8e31

Observation 2c6f8ff8-ad4f-4412-a6ae-31560ec03131 · outbound

This paper cites Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.265067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.265067Z digest=sha256:ad801d3f918b143c95409cfc7095f2140cc026d1f536a2a920ad34e8e23f8efb

Observation e3340e9e-7745-407d-bb4a-6f16a6251130 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.379677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.379677Z digest=sha256:244ca5b0460d078ff6fb6296e1ba1a46da03b04e14c07a0a14f3f7a22a6f849d

Observation 957dce73-1c6d-42dc-be04-396c265a6aa9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.475594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.475594Z digest=sha256:3ce2cf65703826651810dec40711e41582d3072ccd446b18f3ef2256d7de2e11

Observation 32257275-7bc1-4355-b316-0d74ce9e0869 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:55.433502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:45.622864Z digest=sha256:9aebae7e119dd1573743498fd269babed797c195f2d61202e3acc7b5e821e707

Observation 8a0adae8-f90e-48c0-a83b-bca2917ef604 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.740034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.740034Z digest=sha256:9923e8fdde756f7d435c7d84e9e8d8de492203dd2f30b005cd3d22d45295ff53

Observation 768a3bfd-c353-44fc-8570-88a19169d001 · outbound

This paper cites MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.857292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.857292Z digest=sha256:0fc188dc7e736cdfb55caaa58cc91b6be7addefce6d36386b885500dee95b194

Observation ddb0f8d9-ba86-4eaa-bdb9-be3a7b772af0 · outbound

This paper cites Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.930805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.930805Z digest=sha256:e93a841160a376ae08a1b3915990365763964f3328c7508d24be68cdb428d82d

Observation bdd63013-ba0d-4a87-a072-49ed107d3cec · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:55.208825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:46.024858Z digest=sha256:212b04c125eec584d6a766ca4407f7e956d26f439ca9fa77407478ec636de5c1

Observation 8046bc62-217f-40d9-83c9-4fe1946366d6 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.117617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.117617Z digest=sha256:0d197f17be9f3c549146bf7cc09baef47566987441c6b9ad29d954330c594519

Observation a4d8b984-834c-40af-b9a4-eb29c5ee3c72 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.176451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.176451Z digest=sha256:1761c942e195c4cef31ef0039924dac2bc459f1dac7505d0333e6caf3e4e377d

Observation e423f4ba-9fef-4bd5-ac1a-47c6a316cf4e · outbound

This paper cites Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.236243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.236243Z digest=sha256:479943f176ab1a3efc7cfc08ca837bc3371881c1f574c85cd272b2867023bdbf

Observation f38dbbaa-c200-4890-8354-1c0c167ceb48 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.299347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.299347Z digest=sha256:b199049e779047976a210edb43ecafda9c781ad555c3dddbcaca1c645a61ef37

Observation 370c580d-0a04-4306-8fb4-1f918c5537d6 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.417128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.417128Z digest=sha256:a82b0e6ea0937b144f0d6d8241fe2702e9bc1924611cc43987c720a2cc45f69e

Observation cfb42be7-2600-43d1-bc3f-806577ce11ec · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.512570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.512570Z digest=sha256:8c0fbfc7698e2f4712a1b3f8398b2379d9beee5ca356a392c8cfd1a5ba9fa4b0

Observation 5c3f3fc5-4b05-4dee-a396-5883fe2c7ab8 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:54.892978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:46.638279Z digest=sha256:77edbccf4c21d4f7175ae720b79449e809ade38a4b307888b6fb083d91400db9

Observation 33b32872-3a47-4d49-9582-00ed23d0e82a · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.762139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.762139Z digest=sha256:7625068bbe34e58b6c34e20dfd83b8425dae15d7a52f8458b1d4bc795a5e77dc

Observation f4e62474-1e58-4261-b56c-11dbe3e6481f · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.866771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.866771Z digest=sha256:467cbe87d02b6cc6b989dbeed12a1553183c94b7e4e5d4ca02bed7ea9f8a58c9

Observation 25a57ab5-a095-4b2e-9d00-571c75d43ef8 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.934791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.934791Z digest=sha256:6cfd694fbd87f815d84545b70fd7168a3f1cdda41ad0649b41f8f57f2d92a88f

Observation 8c73fc07-fd34-401a-aeef-2da79402d7d8 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.060863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.060863Z digest=sha256:af23c77b4f96698178b3bb872c8530d6ccaf129284364a439e9248393e21150b

Observation 9d3921b0-67c6-460d-a225-b6484f7a61fd · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:54.647025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:47.143106Z digest=sha256:a25ac92007571f70675ceaf981c5be7fecc1392639304f2e36e9143d0848909e

Observation da49f215-98c6-4595-9b16-c524a000d953 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:54.420036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:47.209808Z digest=sha256:0a77c2dc12af48dd8245ba0a56aa241305a6acaf23f5b8de6cc3baf72543e9b6

Observation c8a1f739-ef73-45c0-ab0a-7aa86a36405f · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos TempCompass: Do Video LLMs Really Understand Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.315409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.315409Z digest=sha256:b33001bf0c8a07841744f31ac8b9eacf416aa7d59bcbc1ff31183e38f36e6bac

Observation e30ebd95-91dd-4632-90e8-d9560d30dc93 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.443705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.443705Z digest=sha256:0eb9a63d64bb1c109384b56cb7b612e0125ab87852b665fc8c057e48eea31588

Observation 39f30c25-7fce-4277-af82-32f295dca67d · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.596974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.596974Z digest=sha256:a9c1761d5db739aba7e76cf9268080286d7893ff57053afd1ca438903613c113

Observation 061ae997-6995-458a-a16a-dfbbb2036e48 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:54.169749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:47.719458Z digest=sha256:53cd95a2d789bf30c4b53136a99ef9739c64b2516fe1c9edd89d6034e063b61c

Observation e79e1b7e-d904-4d9e-bac3-d687bd900045 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.827826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.827826Z digest=sha256:9129dc452c0849bcb2025d60a422b0ffe5387fc0622e46231928194a017b3f2c

Observation c56c653a-e8bf-4d33-b1b4-46ccdbe8278c · outbound

This paper cites GPT-4 Technical Report.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos GPT-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.934620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.934620Z digest=sha256:16289d6b3a4cdb45c526b46cbb37937f292a1a07912efe1b94d5739e3912a035

Observation 8fded728-4d52-4a7c-a247-77c5221686b7 · outbound

This paper cites Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:51.709687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:48.041458Z digest=sha256:3fd998ae476043dc73caa4295f3c28ddad8629c41f14bf375ff66f3ede80a5fc

Observation c5774b23-4dd8-4aa4-87d0-f249c70fb3ba · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:53.980201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:48.154875Z digest=sha256:8764a56815b6aec4cf6ad2386c0c04db29afb7d83e2ab72f835828df6348a199

Observation 51d896f8-66df-4027-9c6e-e48beb66ab3b · outbound

This paper cites TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.261199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:48.261199Z digest=sha256:d820227d847314672176841b2e14c29c497e8af8ab5e7d6bca0985664b0b31db

Observation 1d1b9025-dfa8-4623-ab78-e404bc7a46e3 · outbound

This paper cites 2020-2024.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 2020-2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:53.658856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:48.386238Z digest=sha256:5aa7db028f93e3ba67907a404ddb26426f971f9808f2b3ef6806b12e149fd2f3

Observation ed4ca545-399d-4284-9814-4c6876c2381b · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.502389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:48.502389Z digest=sha256:53b6de6ebfe134288da58f86ad5f1965b835ecea11f74eb94ab02af52069ab13

Observation 9692c2c6-a8e4-4a3b-9d0d-22b165b52e3c · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:53.451589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:48.635380Z digest=sha256:3ed978124e5c690d1292391ea2e9de50a421dba4262fe13ded3ce83782c35a12

Observation bd778fa0-cc31-4ff2-af7b-a704eecad666 · outbound

This paper cites Camera-Based HRV Prediction for Remote Learning Environments.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Camera-Based HRV Prediction for Remote Learning Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.781522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:48.781522Z digest=sha256:5be36dfed14773901a339a125cdd120f3c99c3ab46808f825921412842e1b98a

Observation abc5d636-b7bd-4bfe-af42-a838d218fe02 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.866290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:48.866290Z digest=sha256:0133efe10c99c3ed421e1c62f78da10eb56d5d82f7be8a4d4863ac8ed57c394e

Observation 7e315a08-adef-4e27-bc2d-17cfe2d36839 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos CogVLM: Visual Expert for Pretrained Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:48.950196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:48.950196Z digest=sha256:6a6f752ae874a483d5cb0858c102f0bf1f6394e93f37716a3c5afeea9818b128

Observation d1939357-4ace-45e9-b8e8-8547e57afecb · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.003746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:49.003746Z digest=sha256:f8b0c048a616d4a07c440570eace501740b46cb572721debf86e191fba8b88f7

Observation 01ef8c20-4f38-45d4-ac7f-d997d7992b77 · outbound

This paper cites GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:51.382028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:49.068917Z digest=sha256:cbfd639e195a11a45932ac39a3925a80fdae64c84fc7ae023ec489fe01be4649

Observation f020b70c-87f6-44bb-9d4e-21811f81ec82 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.132169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:49.132169Z digest=sha256:3c450e0d9500d52188932d1eecdbc9125012271c6e5324f0159147bd335036e0

Observation c1df389f-4589-48c7-a8a1-380212b6e94c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.200194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:49.200194Z digest=sha256:8dfc3b81e4f9b7617f86efadf3a9250967883a13e70216126a4286ca6e852838

Observation f57660b8-ef58-4f99-8b06-e5b49585ffe6 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-07T12:45:50.863614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:49.270104Z digest=sha256:f1532eae8eda368594c32561f8d5f1fbf73d3dd2a18e672e2fd67c308dee1dad

Observation 1b6c111f-edbb-44fb-a5df-e83d32f5b28f · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.396161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:49.396161Z digest=sha256:3b580ac15e7802673c4bd57b17f93aa0763215ddfe3224716409fbc1e28214f7

Observation e71d5117-a2c8-4d50-8e69-71e178bc275b · outbound

This paper cites Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:51.193106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:49.527744Z digest=sha256:e9744fc5ae03617ab8e786763283e217aabfff2172def3fb70b15b84cc9d3284

Observation 8b5b7977-9ef9-40da-b4e2-b95023332ad6 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.660056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:49.660056Z digest=sha256:48caeabc609f33160cd027017098e5ec690edddce085b9c1ffd4f3e16b2af812

Observation 3f443b40-92ff-4dd4-ad46-dac628dbf559 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:49.782603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:49.782603Z digest=sha256:fe87a0e141ad15d72c85442a495873a70972588baee6695952ce092884009060

Observation 0a2db99e-b946-47d9-8fae-57c1c058f341 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:53.191528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:49.921766Z digest=sha256:7f7bd44c50dfe7c6290b53a2db1ea656ac0366abc259be1c09288c0c4e25589c

Observation 41b00560-3f8b-45e5-a9b5-fce0cc131fb6 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:52.913128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:50.050824Z digest=sha256:aea1da7c70452a996a179c21514faa545bef6331612d9bbb6711ebd85093a4cd

Observation 59f62b73-7618-40f5-89b7-ea8e69ca370f · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:50.222550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:50.222550Z digest=sha256:66b23e041f8a31a513489616a470282b3285eddd16c9236edc69b6963ddb3479

Observation d798a551-6738-4141-8842-eda6e81f2314 · outbound

This paper cites an unresolved cited work.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:52.674815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:45:50.343930Z digest=sha256:3ff7b7818e2adbade09744bbfea426907625d4462b5e9a059b6efcc7f350e512

Observation 64e81e8d-5414-4023-a441-3494ccd95145 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MLVU: Benchmarking Multi-task Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:50.619699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:50.619699Z digest=sha256:137d5ebc180f8951b1e190cca66594ed74efe1dc737a022c0d8383bdc2562575

Pith citing papers

Observation eb3d2990-9dde-45f9-a425-9dcaa0fb0205 · inbound

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation cites this paper.

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.415369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:51:33.951351Z digest=sha256:52c30c56d51e81aaa683505ebfa8c944f4656f1a142926ebe64bbacc055a07c1