Pith. sign in

REVIEW 4 cited by

Towards A Better Metric for Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07781 v1 pith:2EQ6T3ZW submitted 2024-01-15 cs.CV

classification cs.CV
keywords videotext-to-videometricsvideosgenerationmetricbettercriteria
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit impressive capabilities, crafting visually stunning videos. Nonetheless, evaluating such videos poses significant challenges. Current research predominantly employs automated metrics such as FVD, IS, and CLIP Score. However, these metrics provide an incomplete analysis, particularly in the temporal assessment of video content, thus rendering them unreliable indicators of true video quality. Furthermore, while user studies have the potential to reflect human perception accurately, they are hampered by their time-intensive and laborious nature, with outcomes that are often tainted by subjective bias. In this paper, we investigate the limitations inherent in existing metrics and introduce a novel evaluation pipeline, the Text-to-Video Score (T2VScore). This metric integrates two pivotal criteria: (1) Text-Video Alignment, which scrutinizes the fidelity of the video in representing the given text description, and (2) Video Quality, which evaluates the video's overall production caliber with a mixture of experts. Moreover, to evaluate the proposed metrics and facilitate future improvements on them, we present the TVGE dataset, collecting human judgements of 2,543 text-to-video generated videos on the two criteria. Experiments on the TVGE dataset demonstrate the superiority of the proposed T2VScore on offering a better metric for text-to-video generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

    cs.CV 2026-06 unverdicted novelty 6.5 of 10

    CrashTwin recovers metric-scale crash dynamics from monocular rollouts and shows that strong visual scores routinely mask large momentum, energy, and identity violations in world models.

  2. ParticleGen: A Multi-Agent System for Particle Effects Generation

    cs.GR 2026-08 conditional novelty 6.0 of 10

    A multi-agent LLM framework synthesizes editable Unreal Engine 5 Niagara particle systems from text prompts and improves them in a closed loop using rendered-video feedback and diagnostic retrieval.

  3. FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A unified reference-free faithfulness metric for video-to-text and text-to-video that uses fact decomposition, semantic dependency graphs, and VideoQA models.

  4. Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.

Pith tools