REVIEW 3 cited by
Measuring the Quality of Text-to-Video Model Outputs: Metrics and Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Evaluating the quality of videos generated from text-to-video (T2V) models is important if they are to produce plausible outputs that convince a viewer of their authenticity. We examine some of the metrics used in this area and highlight their limitations. The paper presents a dataset of more than 1,000 generated videos from 5 very recent T2V models on which some of those commonly used quality metrics are applied. We also include extensive human quality evaluations on those videos, allowing the relative strengths and weaknesses of metrics, including human assessment, to be compared. The contribution is an assessment of commonly used quality metrics, and a comparison of their performances and the performance of human evaluations on an open dataset of T2V videos. Our conclusion is that naturalness and semantic matching with the text prompt used to generate the T2V output are important but there is no single measure to capture these subtleties in assessing T2V model output.
Forward citations
Cited by 3 Pith papers
-
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
GeneVA is the first large-scale benchmark with human-annotated bounding boxes and text descriptions for artifacts in text-to-video generation.
-
LEHA-CVQAD: Dataset To Enable Generalized Video Quality Assessment of Compression Artifacts
LEHA-CVQAD is a 6,240-clip compressed video dataset with fused MOS and pairwise labels, a hidden test set, and a new Rate-Distortion Alignment Error metric.
-
NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results
All 19 valid entries in the NTIRE 2025 XGC quality assessment challenge outperformed their track baselines at predicting human quality scores for user-generated video, AI-generated video, and talking heads.
Discussion (0). Continue with ORCID to comment.