Pith. sign in

REVIEW 3 cited by

Measuring the Quality of Text-to-Video Model Outputs: Metrics and Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08009 v1 pith:GHANMQOP submitted 2023-09-14 cs.CV cs.MM

classification cs.CVcs.MM
keywords metricsqualityusedvideosdatasethumanassessmentcommonly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating the quality of videos generated from text-to-video (T2V) models is important if they are to produce plausible outputs that convince a viewer of their authenticity. We examine some of the metrics used in this area and highlight their limitations. The paper presents a dataset of more than 1,000 generated videos from 5 very recent T2V models on which some of those commonly used quality metrics are applied. We also include extensive human quality evaluations on those videos, allowing the relative strengths and weaknesses of metrics, including human assessment, to be compared. The contribution is an assessment of commonly used quality metrics, and a comparison of their performances and the performance of human evaluations on an open dataset of T2V videos. Our conclusion is that naturalness and semantic matching with the text prompt used to generate the T2V output are important but there is no single measure to capture these subtleties in assessing T2V model output.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts

    cs.CV 2025-09 conditional novelty 7.0 of 10

    GeneVA is the first large-scale benchmark with human-annotated bounding boxes and text descriptions for artifacts in text-to-video generation.

  2. LEHA-CVQAD: Dataset To Enable Generalized Video Quality Assessment of Compression Artifacts

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LEHA-CVQAD is a 6,240-clip compressed video dataset with fused MOS and pairwise labels, a hidden test set, and a new Rate-Distortion Alignment Error metric.

  3. NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results

    cs.CV 2025-06 conditional novelty 4.0 of 10

    All 19 valid entries in the NTIRE 2025 XGC quality assessment challenge outperformed their track baselines at predicting human quality scores for user-generated video, AI-generated video, and talking heads.

Pith tools