REVIEW 7 cited by
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the \textbf{M}otion-priors \textbf{C}onditional \textbf{D}iffusion \textbf{M}odel (\textbf{MCDM}), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also release the \textbf{TalkingFace-Wild} dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation. Code, models, and datasets will be publicly available.
Forward citations
Cited by 7 Pith papers
-
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.
-
PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
PQ-DAF uses pose-conditioned diffusion generation plus CogVLM filtering to augment few-shot driver distraction training data, and reports large accuracy gains that are compromised by a non-standard train/test protocol.
-
Hybrid Compact Least-Squares and Central Weighted Essentially Non-Oscillatory Schemes for Hyperbolic Conservation Laws on Structured Curvilinear Grids
No verifiable result: the abstract and body address unrelated topics, so the claimed CLS-CWENO schemes appear without derivation, experiments, or benchmarks.
-
FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control
A single caption can drive pose generation, person-image synthesis, and relighting through a three-stage FashionPose pipeline, with reported text-to-pose gains on DF-PASS that are undermined by inconsistent tables.
-
DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On
DiffFit synthesizes virtual try-on images by separately warping the garment geometry and then refining texture with a conditional diffusion model, reporting SOTA metrics on VITON-HD and DressCode but with inconsistent...
-
O2Former:Direction-Aware and Multi-Scale Query Enhancement for SAR Ship Instance Segmentation
O2Former adds a multi-scale query generator and an orientation-aware module to Mask2Former and reports improved SAR ship instance segmentation on SSDD and HRSID.
-
YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
A YOLOv5 variant with BiFPN, directional detail enhancement, and two attention fusion modules reports state-of-the-art mAP on GC10-DET and DAGM2007.
Discussion (0). Sign in to comment.