Pith. sign in

REVIEW 10 cited by

Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12368 v1 pith:7353FRLU submitted 2022-11-22 cs.CV

classification cs.CV
keywords talkinggridportraitaudio-spatialdynamicefficientmodulenerf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage. In this paper, we propose an efficient NeRF-based framework that enables real-time synthesizing of talking portraits and faster convergence by leveraging the recent success of grid-based NeRF. Our key insight is to decompose the inherently high-dimensional talking portrait representation into three low-dimensional feature grids. Specifically, a Decomposed Audio-spatial Encoding Module models the dynamic head with a 3D spatial grid and a 2D audio grid. The torso is handled with another 2D grid in a lightweight Pseudo-3D Deformable Module. Both modules focus on efficiency under the premise of good rendering quality. Extensive experiments demonstrate that our method can generate realistic and audio-lips synchronized talking portrait videos, while also being highly efficient compared to previous methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  2. HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 6.0 of 10

    HM-Talker merges explicit facial-structure cues with implicit audio-driven motion features via stochastic feature pairing to improve talking-head realism and lip-sync.

  3. Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MF-Talk, a mask-free and identity-reference-free three-stage pipeline, improves visual quality and identity preservation in talking-face generation while remaining competitive on lip-sync.

  4. D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    D3-Talker splits face deformation into a general, identity-agnostic branch driven by a facial motion prior and a personalized branch driven by audio, improving few-shot talking head lip sync.

  5. M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    M2DAO-Talker decouples rigid, facial, and oral motion in 3D Gaussian Splatting and uses alternating optimization to reach state-of-the-art scores on small talking-head benchmarks.

  6. Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A shared global Gaussian field plus identity embeddings lets a 3D talking head model adapt to new speakers with a few seconds of footage while improving quality over prior per-identity models.

  7. GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.

  8. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  9. EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A 3D Gaussian-splatting talking head model that conditions facial emotion on continuous valence and arousal while keeping lip movements synchronized to input audio.

  10. DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis

    cs.CV 2024-12 conditional novelty 4.0 of 10

    DEGSTalk combines per-embedding Gaussian deformations with a hair-preserving fusion step to synthesize talking faces that beat prior baselines on PSNR, LPIPS, LMD, and SSIM.

Pith tools