Pith. sign in

REVIEW 2 cited by

Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13640 v1 pith:MFVCOTNC submitted 2024-04-21 cs.MM cs.CVeess.IV

classification cs.MMcs.CVeess.IV
keywords facetemporalvideopgtformerrestorationblindpre-alignmentartifacts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple complex degradations are coupled in low-quality video faces in the real world. Therefore, blind video face restoration is a highly challenging ill-posed problem, requiring not only hallucinating high-fidelity details but also enhancing temporal coherence across diverse pose variations. Restoring each frame independently in a naive manner inevitably introduces temporal incoherence and artifacts from pose changes and keypoint localization errors. To address this, we propose the first blind video face restoration approach with a novel parsing-guided temporal-coherent transformer (PGTFormer) without pre-alignment. PGTFormer leverages semantic parsing guidance to select optimal face priors for generating temporally coherent artifact-free results. Specifically, we pre-train a temporal-spatial vector quantized auto-encoder on high-quality video face datasets to extract expressive context-rich priors. Then, the temporal parse-guided codebook predictor (TPCP) restores faces in different poses based on face parsing context cues without performing face pre-alignment. This strategy reduces artifacts and mitigates jitter caused by cumulative errors from face pre-alignment. Finally, the temporal fidelity regulator (TFR) enhances fidelity through temporal feature interaction and improves video temporal consistency. Extensive experiments on face videos show that our method outperforms previous face restoration baselines. The code will be released on \href{https://github.com/kepengxu/PGTFormer}{https://github.com/kepengxu/PGTFormer}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Persistent Free Volume Governs (Anti)plasticization in Chitosan-Water Mixtures

    cond-mat.soft 2026-04 unverdicted novelty 5.0 of 10

    Dynamically accessible free volume, enabled by connected water-accessible regions, is proposed to govern antiplasticization then plasticization of elastic properties in chitosan–water mixtures.

  2. Disentangling Granularity: An Implicit Inductive Bias in Factorized VAEs

    cs.LG 2025-05 reject novelty 3.0 of 10

    A V-shaped pattern in the training objective of factorized VAEs is attributed to a tunable 'disentangling granularity', but the effect is confounded because coarser granularity removes penalty terms by construction.

Pith tools