Pith. sign in

REVIEW 1 cited by

A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06421 v1 pith:3SA7BIIW submitted 2024-03-11 cs.CV

classification cs.CV
keywords headtalkinghumanmetricsaigcaudio-drivendevelopmentevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code and data will be made available at https://github.com/zwx8981/ADTH-QA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space

    cs.CV 2024-11 conditional novelty 5.0 of 10

    LES-Talker defines emotions as 41-dimensional vectors over facial action units and uses them to edit talking-head videos with continuous emotion levels and per-muscle control.

Pith tools