Pith. sign in

REVIEW 7 cited by

A Generalist FaceX via Learning Unified Facial Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00551 v1 pith:V2VUO3WB submitted 2023-12-31 cs.CV

classification cs.CV
keywords facialfacextaskseditingrepresentationunifieddecomposinggeneralist
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously. To achieve this goal, we initially formulate a unified facial representation for a broad spectrum of facial editing tasks, which macroscopically decomposes a face into fundamental identity, intra-personal variation, and environmental factors. Based on this, we introduce Facial Omni-Representation Decomposing (FORD) for seamless manipulation of various facial components, microscopically decomposing the core aspects of most facial editing tasks. Furthermore, by leveraging the prior of a pretrained StableDiffusion (SD) to enhance generation quality and accelerate training, we design Facial Omni-Representation Steering (FORS) to first assemble unified facial representations and then effectively steer the SD-aware generation process by the efficient Facial Representation Controller (FRC). %Without any additional features, Our versatile FaceX achieves competitive performance compared to elaborate task-specific models on popular facial editing tasks. Full codes and models will be available at https://github.com/diffusion-facex/FaceX.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation

    cs.CV 2025-02 conditional novelty 7.0 of 10

    A conditional diffusion model generates image-conditioned vector sketches in under a second by denoising stroke coordinates, trained on a synthetic dataset produced by a new SDS-based optimizer with depth ControlNet.

  2. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  3. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DiffSensei combines an SDXL diffusion generator with a multimodal LLM adapter and masked attention to generate manga pages with multiple characters whose poses and expressions follow panel captions.

  4. ControlFace: Harnessing Facial Parametric Control for Face Rigging

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ControlFace performs zero-shot face rigging from 3DMM renderings using a dual-branch U-Net, a control mixer module, and reference control guidance, and reports the best average DECA re-inference error on FFHQ baselines.

  5. ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ConsistentAvatar aligns a Fourier high-frequency detail map through a diffusion model and uses it, with normals and emotion text, to condition talking-head avatar generation, reducing temporal and expression inconsistency.

  6. Dense-Face: Personalized Face Generation Model via Dense Annotation Prediction

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Dense-Face is a personalized face generation model that adds a pose-controllable adapter and dense face annotation prediction to Stable Diffusion, improving identity preservation and text alignment.

  7. Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

    cs.CV 2024-11 conditional novelty 4.0 of 10

    The paper introduces BioDeepAV, a benchmark of real and fake talking-face videos, and reports that state-of-the-art deepfake detectors drop sharply when tested on deepfakes from unseen generators.

Pith tools