Pith. sign in

REVIEW 4 cited by

DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.03988 v1 pith:BUNYYZNN submitted 2023-03-07 cs.CV

classification cs.CV
keywords dinetdubbingdeformationfacefeaturevisuallymapspart
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For few-shot learning, it is still a critical challenge to realize photo-realistic face visually dubbing on high-resolution videos. Previous works fail to generate high-fidelity dubbing results. To address the above problem, this paper proposes a Deformation Inpainting Network (DINet) for high-resolution face visually dubbing. Different from previous works relying on multiple up-sample layers to directly generate pixels from latent embeddings, DINet performs spatial deformation on feature maps of reference images to better preserve high-frequency textural details. Specifically, DINet consists of one deformation part and one inpainting part. In the first part, five reference facial images adaptively perform spatial deformation to create deformed feature maps encoding mouth shapes at each frame, in order to align with the input driving audio and also the head poses of the input source images. In the second part, to produce face visually dubbing, a feature decoder is responsible for adaptively incorporating mouth movements from the deformed feature maps and other attributes (i.e., head pose and upper facial expression) from the source feature maps together. Finally, DINet achieves face visually dubbing with rich textural details. We conduct qualitative and quantitative comparisons to validate our DINet on high-resolution videos. The experimental results show that our method outperforms state-of-the-art works.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new large dataset and the FSCD model improve automated quality scoring of AI-generated talking-head videos, beating 15 baselines in correlation with human ratings.

  2. GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.

  3. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  4. NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results

    cs.CV 2025-06 conditional novelty 4.0 of 10

    All 19 valid entries in the NTIRE 2025 XGC quality assessment challenge outperformed their track baselines at predicting human quality scores for user-generated video, AI-generated video, and talking heads.

Pith tools