Pith. sign in

Paper Citation Record · LEDGER

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model

As of 23 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2506.04715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04715 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:40:44.329318Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a07b24ca-6e2a-4635-858e-02e5d81925f1 · outbound

This paper cites GPT-4 Technical Report.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.114739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.114739Z digest=sha256:5085790d44ae8a1ce5f0faa3ee3d9606f533cbd56bcd9ba50ab9dc2b80054a57

Observation d4d8a437-10dc-4032-a342-f2c0d627b82a · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model PaliGemma: A versatile 3B VLM for transfer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.119678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.119678Z digest=sha256:3fd04354fd571cb9ec6701c1385a6d7c2a6b2e6742e4d25e57fa4534db01b776

Observation a9ff83c0-0d01-4117-a313-954579a3319b · outbound

This paper cites Video generation models as world simulators.OpenAI Blog, 1:8, 2024.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Video generation models as world simulators.OpenAI Blog, 1:8, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.965886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.124898Z digest=sha256:cdf23bb8910d9cce519a079c2b7a642e7cbb1955f8f69b92a2446283387f4e5d

Observation 83054b39-1872-43d2-8751-5e7388374fbd · outbound

This paper cites Gaia: Rethinking action quality assessment for ai-generated videos.Advances in Neural Information Processing Systems, 37:40111–40144, 2024.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Gaia: Rethinking action quality assessment for ai-generated videos.Advances in Neural Information Processing Systems, 37:40111–40144, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.955665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.129180Z digest=sha256:df5139fce96e21a97bd3627e2e291a57d421123f83148b5321104be8638eadc7

Observation cde9eb70-cf40-4b76-a9b0-568517b381e5 · outbound

This paper cites FineVQ: Fine-Grained User Generated Content Video Quality Assessment.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model FineVQ: Fine-Grained User Generated Content Video Quality Assessment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.132814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.132814Z digest=sha256:94f41050921a4870dc18cb6594abc924e4023ad33bef0f3bb10d706baf01898b

Observation 70292eb2-ec04-43d3-98fd-bddd339d4244 · outbound

This paper cites Slowfast networks for video recognition.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Slowfast networks for video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.944062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.136726Z digest=sha256:4496ae93031659b3575e57c944a3079c798bd6acd9457a8c41d7f8ab9745d2af

Observation f107fd90-7332-4310-944a-47692c976595 · outbound

This paper cites LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.140384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.140384Z digest=sha256:a1fa31ec7a4fac58fda7365bf959e1f484ccaf352a4c7dc363433c2a66e4da0b

Observation f27c5f0d-be89-4559-9385-25e53a47fc21 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.144122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.144122Z digest=sha256:b31944123eb251ee85d5d099940ed2478210dc70a371818c045228f500a893bf

Observation fe7a4433-8d72-4cca-8287-d3516d4b07e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.148464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.148464Z digest=sha256:7bc449014e23721528cfd96e1fc4ade043ee89db22beabeefd443415538faa8c

Observation f15939d0-1231-49d1-9fdd-eb51e5b1d537 · outbound

This paper cites Deep residual learning for image recognition.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Deep residual learning for image recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.153003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.153003Z digest=sha256:f8fc7b36cfadc778c43795d037cf34f2fdb965b7048de13c4c84d32dc83dedfe

Observation 29cdfd3c-1b3f-48eb-a91c-d9b440d1da04 · outbound

This paper cites VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.156703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.156703Z digest=sha256:1e0062b083b5b4ec1e2fd5d49d9e22449b0188b0612a0939400862cc553c0242

Observation 51e2a8d0-1b30-411c-8bc5-695271c9e2cc · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.160413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.160413Z digest=sha256:7bf549f2ac3de9d5a6bf0fc63967ea120f558495933d381608a36cf1defbf79d

Observation 0acd423a-8bc7-45bd-82bb-d9ac9d4f0e62 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.926878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.164913Z digest=sha256:4f5c7711123f98cf418a3a4f73cf3c00b2659b43e3ac2983ed44cbef365ac7ca

Observation b8ec7eb7-bf5b-4406-ba70-9c726a0943e9 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.168646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.168646Z digest=sha256:26da7461dc861fde0a1827bb5ea2c83b639e20c258e03bc5215784bbcec8e16a

Observation 60577905-30eb-4f96-9fbe-9621bceebc64 · outbound

This paper cites T2vbench: Benchmarking temporal dynamics for text-to- video generation.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model T2vbench: Benchmarking temporal dynamics for text-to- video generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.916843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.172419Z digest=sha256:ad8c55a03a5ce1f238bb68ba25b6898ba80ba69dafe498df8b3d88af322a4ab4

Observation 7a2f419a-51e4-492a-b327-16ce279e1381 · outbound

This paper cites VQA$^2$: Visual Question Answering for Video Quality Assessment.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model VQA$^2$: Visual Question Answering for Video Quality Assessment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.175524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.175524Z digest=sha256:ba2887d2d74dc1e8eff1974846187892a5c2797d361099686ca40e03cc03ea80

Observation 07b76e28-bad1-4919-8c28-a32924d4c17f · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.179074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.179074Z digest=sha256:98b0c3ce5a56d00a999a527042463fc6848a5f2717c1b11d3150c1605e948f21

Observation 28be2090-27a2-493a-bb34-c42e933bb7af · outbound

This paper cites The Kinetics Human Action Video Dataset.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model The Kinetics Human Action Video Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.182568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.182568Z digest=sha256:c77ca31425af378ff7dc00c7deaae0f7a455a7f3f6a27b54a074d9920cf054e8

Observation f2479dc1-eedd-4002-b112-57effcd4eaec · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.186888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.186888Z digest=sha256:7f7df768cdcc7f373e6c3a6184cce72a5bbfbf770493944511657c8404104e1c

Observation 3546953f-b69f-4dfe-ae73-3a59cb2b9460 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video qual- ity assessment.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Subjective-aligned dataset and metric for text-to-video qual- ity assessment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.906519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.191010Z digest=sha256:19c16dbf3b5c4c6d5f3327f02dfddb9641062828eaa4d42e178b0bf34c86e3ee

Observation bf598dba-e315-46f0-9637-a073a898f36c · outbound

This paper cites an unresolved cited work.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:44.896064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.194367Z digest=sha256:c20236c66e58b70aad18e355e75d5eb539e22ec7edb9e9322a6665087b4e2d1f

Observation 86903391-b6d1-4232-96c8-5b9b5d6993f0 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.886501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.197760Z digest=sha256:91463b9764163d7ca01d8f7019370449e5a95a15ef7b9d0f712380314c6d1f16

Observation d2a30562-a971-4681-bb31-80e0f58f61c2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.201863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.201863Z digest=sha256:c4f0f655465eef054decec362d19d4e4a9e1400f18c761593019c5e7ad3048d2

Observation 0a2c6da0-da65-4cc0-a606-15fdd29e4e80 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Unmasked teacher: Towards training-efficient video foundation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.870663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.205993Z digest=sha256:9fe817ec53cfb484415e39a739bc24dcb62b436ab0bf00cda94ea283fdaa048d

Observation 52c7be92-9738-4ca7-9001-f73a2614b322 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Evaluating text-to-visual generation with image-to-text gen- eration

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.860199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.209106Z digest=sha256:bff19fd2af185d96b454bc9a9a5aa805e21511023f17fa972acabb57acaed1a6

Observation 50455e1e-506f-456d-af73-905e59297120 · outbound

This paper cites Evalcrafter: Benchmarking and eval- uating large video generation models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Evalcrafter: Benchmarking and eval- uating large video generation models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.849788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.212430Z digest=sha256:1f103c85debff64735bf0c9ae885bb6be2a5af7661a548297a80ea4a7293a7fb

Observation a43245b0-3b1d-4c7b-bba0-18ceef46d07d · outbound

This paper cites Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.Advances in Neural Information Process- ing Systems, 36, 2024.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.Advances in Neural Information Process- ing Systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.840009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.215724Z digest=sha256:9bbb8065c8ff90b712aef6bf9b76d38fb485f7a8df6589afd361e417ecaff221

Observation 6d2ec5a0-aad9-4d02-927a-461ea1eabbd5 · outbound

This paper cites A convnet for the 2020s.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model A convnet for the 2020s

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.219006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.219006Z digest=sha256:099e03d5af639255588bc2b2ef40307a313a132fc991e21b9debded908f058ea

Observation 68e3b078-b8b3-463e-8df2-e46017142904 · outbound

This paper cites Video swin transformer.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Video swin transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.222932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.222932Z digest=sha256:96e7191fb6ff63104c3355b06b43447eaca07fb5a2c2a0e69db15d923a4b8094

Observation 7fcb6ded-57e1-4926-93ef-bd9d10289684 · outbound

This paper cites Aigc- vqa: A holistic perception metric for aigc video quality assessment.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Aigc- vqa: A holistic perception metric for aigc video quality assessment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.226111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.226111Z digest=sha256:9b6fb5537a95f4a4a092286c3c89fa1bb2615597d69da85614c0c001be07d83c

Observation 33b35cc4-aae1-44b0-a66c-a5948ce78f64 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.229202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.229202Z digest=sha256:4578454ed88008dc0aedbf4eeeb663bf69120b5734871b5c7c320e5406bbf0d4

Observation 2bca5590-1785-4951-950f-004c454deaa0 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.233840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.233840Z digest=sha256:5ace0be96706f56001e5f9c28f4a2a46f43f006192221d553567760e226c3a26

Observation 888614ba-6a5d-42af-a596-16d8e859fe87 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.237208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.237208Z digest=sha256:8824d88114ac40954571fef39d87f7fea3dc07347a66f73c0b9b32644c50169f

Observation ac16df1a-4f5a-4c46-9528-58febc3545f8 · outbound

This paper cites T2veval: Benchmark dataset and objective evaluation method for t2v-generated videos, 2025.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model T2veval: Benchmark dataset and objective evaluation method for t2v-generated videos, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.811727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.240669Z digest=sha256:8fd36595369e2d8d78cb22301717c454d0d7bd3a06da4e2a2f7687aaf1b62633

Observation de7c3ed7-5a84-453a-ac76-80d65a94535e · outbound

This paper cites Comprehensive subjective and objective evaluation method for text-generated video.arXiv preprint arXiv:2501.08545, 2025.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Comprehensive subjective and objective evaluation method for text-generated video.arXiv preprint arXiv:2501.08545, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.244030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.244030Z digest=sha256:51c95e5ee692172faee2ceaf97b40923a623b75c0e86efdab162886711f65715

Observation 63dfae0d-6ccf-4033-98da-c0ba011c92b3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.247250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.247250Z digest=sha256:d879af5685ff031e6aac5b8551c57d049697a1979ae39272d9370edd529c4d31

Observation c5f906f3-4809-4d6c-bff4-c63f59be90d3 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.250446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.250446Z digest=sha256:14e5612c2fa2d8de9a5e0a6d329c7381a46e43b1b06ef695b3c3f8a1233544e1

Observation 48e0d44d-f7df-49e8-840a-43702c398847 · outbound

This paper cites A deep learning based no-reference quality assessment model for ugc videos.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model A deep learning based no-reference quality assessment model for ugc videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.794543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.254002Z digest=sha256:c39bb1612dc5cd2dae405c4412924c85bbd6fec99ca70d10d92d50287146f99d

Observation 66e8f99e-4d92-4084-908c-0a625085a732 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.257047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.257047Z digest=sha256:d3d991791399c1e9ff5fb97326df8797156a4507d45abf7abee405f429c733dd

Observation 3401f142-855d-4df7-b9ee-8fa22fea7c2f · outbound

This paper cites AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.261429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.261429Z digest=sha256:d99bb207ada0e3a9c8abb4b804f9d46d6021d543a5b5c81f0ab2fed0d5e87ea2

Observation 1add9b73-f19b-4630-bd41-a83353015b22 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Videomae v2: Scaling video masked autoencoders with dual masking

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.265763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.265763Z digest=sha256:581bef2bcee28fbad4f2063279471fabfd323c70ef71ba3268f6258ed3b351c3

Observation e871c747-e2c3-4241-9229-d221f6bd513f · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.268995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.268995Z digest=sha256:84b2c4c28f029f88482ae08aa9809d6576371eb09cd828e0487ea1a784c1ad39

Observation aa25a1f6-2f0c-4953-900c-79ed1d994ee9 · outbound

This paper cites An ensemble approach to short-form video quality assess- ment using multimodal llm.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model An ensemble approach to short-form video quality assess- ment using multimodal llm

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.778286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.272673Z digest=sha256:da3312d969569c874236b77c8d593c71f007933d07045a8453ff99acaf616133

Observation e23de760-b9c2-480a-878d-6b13cfd6ef92 · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.276578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.276578Z digest=sha256:f07f612e95015380fb43ed5cca2ba9db6f6d4b8f27f13f5e2543e4ffec275fac

Observation cb7d9e39-8d5e-4bc2-b803-cd00513ada45 · outbound

This paper cites Fast- vqa: Efficient end-to-end video quality assessment with frag- ment sampling.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Fast- vqa: Efficient end-to-end video quality assessment with frag- ment sampling

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.767817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.280372Z digest=sha256:fb2e282a69caf6ceb02bdc46485967073ee0f3681a9d05f6c85b87b56c2ca50c

Observation dfa46c75-9ae7-41ab-9d46-a60028823639 · outbound

This paper cites Exploring video quality assessment on user gener- ated contents from aesthetic and technical perspectives.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Exploring video quality assessment on user gener- ated contents from aesthetic and technical perspectives

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.757489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.283644Z digest=sha256:f8a5a24678f5b54f17cde1f78dfe7d5c058fa1e8b18e9ec2b3834ee83c63cfce

Observation b5b8cca9-1ef2-4526-bf60-e1b449fd3dea · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.287561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.287561Z digest=sha256:6d68b05501a541d8a5d221a36465e20933f6e11cbd9096049a553c439253c96c

Observation ef889838-8e80-4eb4-b263-6611978d116c · outbound

This paper cites Q-instruct: Improving low-level visual abilities for multi-modality foundation models.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Q-instruct: Improving low-level visual abilities for multi-modality foundation models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.291522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.291522Z digest=sha256:388e0846aa41946848941b9b88ceb2bfe8e4f1f29b1531fcaa0b0839ef51f8d5

Observation 97ca7cfb-dd89-4015-892b-afcd42b9f595 · outbound

This paper cites Grit: A gener- ative region-to-text transformer for object understanding.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Grit: A gener- ative region-to-text transformer for object understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.739499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.294798Z digest=sha256:470e78305a7eb89db1691965a987a5e4609fe163aec348d89ac4016c595b53da

Observation 7a8f4aa1-f205-479f-a0f0-c7e1f99203cd · outbound

This paper cites Ntire 2025 xgc quality assessment challenge: Methods and results.CVPR Workshop, 2025.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Ntire 2025 xgc quality assessment challenge: Methods and results.CVPR Workshop, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.726400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.298447Z digest=sha256:3c2af9c2b53adf96980ef2287220bbc119871f831b956771a576e93c89caf9e3

Observation f0a91cb3-f3a7-4f3b-b077-ae9e4c66857c · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.301982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.301982Z digest=sha256:1decca684e74cdf1c466dbe6415ea8408d8a3ef90db0d2bf3fffd122d7f01ab8

Observation b4ac5347-e102-48d8-8855-c34593cfa4f2 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36, 2024.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.708029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.305309Z digest=sha256:669a5ec49adbef5f7367bc1046998586f19c693330c3e29814524ac7eb86490e

Observation 1cb4c45e-3c98-40cd-809e-609daf22a195 · outbound

This paper cites Qwen2.5 Technical Report.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Qwen2.5 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.309647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.309647Z digest=sha256:23e521d74d8bc9a0dc62d850d1c2bda9610aa8ddb4c65578687b478cef86e0b7

Observation b071af26-8d4c-4a8a-aa99-61b0b8eaa9f5 · outbound

This paper cites Chronomagic-bench: A bench- mark for metamorphic evaluation of text-to-time-lapse video generation.Advances in Neural Information Processing Sys- tems, 37:21236–21270, 2024.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Chronomagic-bench: A bench- mark for metamorphic evaluation of text-to-time-lapse video generation.Advances in Neural Information Processing Sys- tems, 37:21236–21270, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.696715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.313355Z digest=sha256:9498caa2a137afed67a9a39d886d5d45051245c863ec09f3b53ea51a4d40c6de

Observation 9ffb76a0-0386-44f1-bab9-528dcac89c26 · outbound

This paper cites Sigmoid loss for language image pre-training.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Sigmoid loss for language image pre-training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.316789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.316789Z digest=sha256:27e700006899666e8b15d6ae37c82e74223b2bad10163ed9ad5fadd7d42ac130

Observation 86fbed20-9b37-4249-8fee-373b6ee76418 · outbound

This paper cites Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:40:44.365547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.320792Z digest=sha256:59ac90c544bb444d8e2d27e5feb995850489c60d992a0fc94519795b7a288612

Observation 03188bb0-db5f-41f1-a1fd-eb982ab514e2 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.679396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.324872Z digest=sha256:d30f8b29c213acaf0f52cb1bd3ac799dd6db14c1c0c70f21de492395d67e01c4

Observation f0b9a4e7-9957-4256-9693-8c4568e6499d · outbound

This paper cites Open-sora: Democratizing efficient video production for all, march 2024.URL https://github.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model Open-sora: Democratizing efficient video production for all, march 2024.URL https://github

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:40:44.668091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:40:44.329318Z digest=sha256:3bdf862e34d4e0addf6f449fe73cdc4a70db6d714d12703e44fc39d0826a9afb

Pith citing papers

No inbound Pith citation observations are available.