Pith. sign in

REVIEW 4 major objections 6 minor 153 references

Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This survey argues that generative face video coding, which transmits facial motion parameters and synthesizes frames with a generative model, beats the VVC standard by more than 70 percent on perceptual quality at ultra-low bitrates.

desk verdict Useful benchmark and subjective database for GFVC, but the 'far surpassing VVC' headline rests on perceptual proxies that the paper itself shows correlate with human opinion only around 0.85, and no subjective VVC comparison is included. read the letter →

arxiv 2506.07369 v1 pith:BQLKQ73Q submitted 2025-06-09 cs.CV

classification cs.CV
keywords generativefacevideocodingultra-lowbitratecompressiondeepmodelsperceptualqualityassessmentsubjectivedatabasestandardizationsupplementalenhancementinformationmodellightweighting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper surveys and benchmarks Generative Face Video Coding (GFVC), a compression paradigm that sends only a compact description of facial motion — a few keypoints, landmarks, or a small matrix of parameters per frame — and lets a deep generative model synthesize the video frames at the receiver from a conventionally coded reference frame. The central claim is that on perceptual quality metrics, GFVC methods outperform the latest Versatile Video Coding (VVC) standard by wide margins at ultra-low bitrates, with the best basic method (DAC) reporting more than 70 percent savings in BD-rate, the standard measure of average bitrate reduction at matched quality. To back this, the paper builds a 1,056-sequence face-video database with human opinion scores, finds that perceptual metrics like DISTS and LPIPS track human judgment far better than PSNR and SSIM for generative content, and reports that GFVC parameters are being standardized as SEI messages — a standard bitstream mechanism for auxiliary data — in VVC, HEVC, and AVC. The stakes are practical: if the claim holds, high-fidelity face communication becomes possible on severely bandwidth-constrained networks where conventional codecs visibly collapse.

What carries the argument

The load-bearing object is the compact motion representation: facial dynamics are reduced to a few dozen numbers per inter frame — ten 2D keypoints with their Jacobians (60 parameters) for FOMM, ten keypoints alone (20 parameters) for DAC, a 4×4 temporal-evolution matrix (16 parameters) for CFTE — which are then quantized, inter-predicted, and entropy coded, while a face-animation network at the decoder synthesizes each frame from the coded reference frame and these parameters. The system is trained end-to-end against the rate-distortion objective $D + \lambda R$, and the benchmark shows the whole-class advantage over VVC is carried by the cheapness of this representation: the least costly parameter set (DAC) yields the largest savings, and the costliest (TPS at 100 parameters per frame) the smallest, on both resolutions.

What would settle it

Re-run the benchmark's rate-quality comparison on the paper's own human opinion scores instead of DISTS and LPIPS: pick the VVC and GFVC reconstructions at matched bitrates from the 1,056-sequence database and compare MOS directly. If human raters no longer favor the GFVC clips at the advertised 70-percent-saving operating points, or favor them by much less, the central claim fails, and the paper's Fig. 7 already shows where such disagreements are likely.

Watch

Extended reading notes

Core claim

The paper's central claim is that representing a face video by its motion rather than its pixels resets the rate-distortion tradeoff for face content. In the GFVC pipeline, a reference frame is coded with a conventional intra codec, an analysis model extracts a compact description of facial dynamics for each inter frame, and a generative synthesis model reconstructs the frames from the reference and the motion parameters, with the whole system trained end-to-end against a rate-distortion loss. Under common test conditions at 256×256 and 512×512, the keypoint-based codec DAC saves 66 to 71 percent BD-rate over VVC on DISTS and LPIPS, CFTE and FV2V save roughly 50 to 66 percent, and the multi-reference extension MRDAC saves more than 30 percent; on PSNR and SSIM the same methods fall far behind VVC, which the paper attributes to perceptual rather than pixel-level optimization. A subjective study of 1,056 GFVC-compressed sequences with 20 raters shows perceptual metrics correlating with human opinion at about 0.80 to 0.86, against roughly 0.5 for PSNR and SSIM, with visible residual disagreements the paper reports openly.

Load-bearing premise

Every headline savings number is computed with DISTS and LPIPS standing in for human quality judgment, and the paper's own subjective data show those metrics agree with human opinion scores only about 80 to 85 percent of the time, with visible disagreement on some GFVC codecs.

Editorial extensions

If this is right

  • At bitrates near 5 kbps, VVC reconstructions of faces break into blocking artifacts that are barely recognizable, while the benchmark's GFVC methods still deliver visually pleasant reconstructions, a gap reflected in BD-rate savings of 66 to 71 percent for the best basic method.
  • Because the next versions of VVC, HEVC, and AVC will carry GFVC parameters as SEI messages, the technology can be layered onto existing standardized bitstreams without normative changes to the hybrid codecs.
  • Perceptual metrics such as DISTS, LPIPS, TOPIQ, and FVD, rather than PSNR and SSIM, are the appropriate tools for evaluating and optimizing generative face compression, based on their roughly 0.80 to 0.86 correlation with human opinion scores on the new database.
  • Lightweight optimization of CFTE — depthwise separable convolutions plus BatchNorm-scale channel pruning — cuts parameters by about 90 percent and compute by about 85 percent with no noticeable rate-distortion loss, indicating GFVC can run on resource-constrained devices.
  • If the field follows the benchmark's cost-versus-savings trend, the least expensive motion representations, not the most expressive ones, set the practical frontier for ultra-low-bitrate face communication.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Re-running the comparison directly on the paper's own human opinion scores, rather than on DISTS and LPIPS, could shrink or reverse the advertised margin: the paper's Fig. 7 shows objective and subjective scores visibly disagreeing on several GFVC codecs, notably RDAC, whose temporal flickering humans penalize.
  • The analysis-synthesis pattern is portable: once compact representations exist for bodies, scenes, or general video, the benchmark's per-parameter cost trend predicts where generative coding will pay off next, a direction the paper itself gestures toward.
  • The MOS database sets a concrete target for quality-assessment research: a metric that exceeds the roughly 0.85 agreement ceiling of DISTS and LPIPS on generative content would change how GFVC comparisons are reported and how GFVC systems are optimized.
  • Interoperability via parameter translation is what makes the SEI-based standard practically usable across mismatched encoder and decoder models; without a standardized translator, the SEI approach effectively locks each link to a single generative model.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper surveys Generative Face Video Coding (GFVC), reviews compact-representation and optimization-strategy methods, benchmarks five basic and three optimized GFVC algorithms against VVC under JVET test conditions, reports BD-rate savings on perceptual metrics, constructs a GFVC-compressed subjective database with MOS ratings, evaluates 18 objective quality metrics against those MOSs, summarizes JVET standardization activities, and proposes a lightweight CFTE implementation via depthwise separable convolutions and channel pruning.

Significance. The manuscript provides a useful consolidation of a fast-moving area: it documents over sixty JVET contributions, gives a reasonably detailed benchmark setup, and contributes a new 1056-sequence GFVC subjective database with MOS from 20 subjects. The objective-metric correlation analysis (Table 5) is a valuable resource for subsequent GFVC quality assessment work. The strongest quantitative claim, however, is not supported as stated: the advertised large BD-rate savings over VVC are computed exclusively with Rate-DISTS and Rate-LPIPS, while the same paper reports that GFVC lags VVC substantially on PSNR and SSIM, and the new subjective database contains no VVC-compressed content against which human preference could be directly compared. These issues are fixable, but they are load-bearing for the abstract's 'far surpassing VVC' conclusion.

major comments (4)
  1. [Abstract; Section 3.2.3] The abstract and conclusion claim that GFVC can enable high-fidelity face video communication 'far surpassing' VVC, but the benchmark evidence in Section 3.2.3 is metric-specific: Tables 2-4 report BD-rate savings only for Rate-DISTS and Rate-LPIPS, and the accompanying text and Fig. 4 state that on Rate-PSNR and Rate-SSIM the same GFVC algorithms 'fall far behind the VVC anchor and show obvious saturation.' The superiority claim should be explicitly qualified to perceptual-quality metrics, and the pixel-level tradeoff should be stated in the abstract and conclusion rather than only in the body.
  2. [Section 4.1; Section 4.2; Fig. 7] The subjective database constructed in Section 4.1 contains only GFVC-compressed sequences (FOMM, CFTE, FV2V, DAC, TPS, HDAC, RDAC, MRDAC) and does not include any VVC-compressed sequence. Consequently, no human MOS directly compares GFVC against VVC, even though the paper's headline claim is exactly that comparison. This matters because Table 5 shows that the best objective metrics (DISTS, TOPIQ, LPIPS, FVD) correlate with MOS only at PLCC approximately 0.85-0.86, and Fig. 7 documents visible disagreements between objective and subjective scores, with RDAC samples appearing in contradictory regions. The >70% BD-rate margin is therefore a margin on an imperfect proxy and is not calibrated against human judgments for the VVC-versus-GFVC comparison. Please add a VVC condition to the subjective experiment or explicitly present the savings as proxy-based.
  3. [Section 3.2.2(c); Table 4] There is an internal inconsistency in the benchmark description. Section 3.2.2(c) states that the three extended GFVC algorithms (HDAC, RDAC, MRDAC) 'are implemented on 256×256 resolution,' and Fig. 5 is labeled 'on 256×256 resolution,' yet Table 4 is titled as reporting 'average BD-rate savings over the VVC anchor on 512×512 resolution.' The manuscript should clarify which resolution Table 4 actually reports, regenerate the table if it was mislabeled, and ensure the caption and text agree.
  4. [Section 3.2.3; Tables 2-4] The BD-rate savings are reported as single scalar values with no confidence intervals, per-sequence spread, or significance testing. Given that the classes contain only 15-18 sequences each and that BD-rate values can be sensitive to the anchor QP set and to the choice of fitted RD curve, the reader cannot assess whether the differences among GFVC methods, or the margins over VVC, are statistically distinguishable. Please provide per-sequence results or a measure of dispersion, at least for the headline DAC/CFTE/FV2V comparisons.
minor comments (6)
  1. [Section 3.2.4; Figs. 3, 4, 5, 6] The TPS method is cited as [65] in the captions of Figs. 3, 4, 5, 6 and in Section 3.2.4, but reference [65] is the SpyNet optical-flow paper; the thin-plate spline motion model is [44]. Please correct the citation numbering throughout.
  2. [Section 5.2; Table 7] The low-complexity claim is supported by parameter counts and kMACs/pixel, but no measured inference latency, memory footprint, or end-to-end runtime is reported. Since the contribution list describes a 'low-complexity GFVC system,' please add at least one runtime measurement on a representative device, or soften the claim.
  3. [Section 5.2.1; Eq. (7)] The pruning ratio r and the sparsity regularization weight lambda are described as experimentally determined, but the manuscript does not report the chosen values or describe the sensitivity of the 90% parameter reduction to these hyperparameters. Reporting the actual values would improve reproducibility.
  4. [Section 3.2.1] The training data is described as 512x512, while Classes A and B are evaluated at 256x256. The description of how the multi-resolution models [64] adapt to the lower resolution is not included, so the reader cannot fully assess whether the 256x256 results are obtained under comparable training conditions.
  5. [Section 5.1.1] The text uses both 'AhG' and 'AHG' (e.g., 'Ad hoc Group' and 'AHG16'); please standardize this terminology.
  6. [Section 3.2.3] The sentence 'TABLE 2 and TABLE 3. shows the BD-rate saving...' contains a grammatical error and should be reworded, e.g., 'Tables 2 and 3 show the BD-rate savings...'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's central claims rest on independent benchmark measurements against the external VVC anchor and on human subjective ratings, not on definitional or fitted reductions.

full rationale

The paper's central quantitative claim that GFVC can 'far surpass' VVC is supported by BD-rate savings computed against the external VTM 22.2 VVC anchor (Sections 3.2.1-3.2.3). These are direct measurements of rate and quality, not outputs of a model fitted to the claimed conclusion. The subjective database in Section 4.1 is built from human MOS ratings of GFVC-compressed sequences, and the correlation analysis in Section 4.2 is a separate empirical study; the paper explicitly acknowledges the imperfection of the objective metrics ('correlations around 0.85 are still not strong enough to match the human opinions'). The absence of a VVC anchor in the MOS database weakens the perceptual-superiority claim, but it does not make the claim circular, because the advertised margin is computed on objective metrics rather than derived from those MOS values. Self-citations to the authors' own GFVC methods and JVET documents are present, but the load-bearing evidence is the external VVC comparator and human subjects; no load-bearing argument reduces to a self-citation or to a definitional identity. No equation in the paper is equivalent by construction to a predicted result, and no fitted parameter is renamed as a prediction. The paper is therefore self-contained against external benchmarks for its central derivation, despite legitimate concerns about benchmark ownership and metric-choice bias that fall outside the definition of circularity.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claims rest on domain assumptions about test conditions and quality metrics rather than on new free parameters. The only free parameters are the undisclosed pruning hyperparameters of the lightweight CFTE, which prevent replication. No invented entities are introduced.

free parameters (2)
  • Per-part pruning ratio r = not reported
    In Section 5.2.1, each network part is assigned a pruning ratio r that is 'experimentally determined', but the values are not given; the claimed 90 percent parameter reduction depends on these choices.
  • Sparsity regularization weight lambda = not reported
    Eq. (7) introduces lambda to balance task loss and BatchNorm scaling-factor sparsity; without its value, the pruning pipeline cannot be reproduced.
assumptions (3)
  • domain assumption The JVET-AJ2035 test sequences, QP points, and VVC LDB configuration are a fair and representative basis for comparing GFVC against VVC.
    Section 3.2.1 adopts JVET-AJ2035 test conditions, a document co-authored by the first author, and all headline BD-rate savings are computed relative to those settings.
  • domain assumption DISTS and LPIPS are the appropriate ground-truth quality measures for the claim that GFVC 'far surpasses' VVC.
    All BD-rate savings over VVC in Tables 2 to 4 are computed on DISTS and LPIPS only; Section 4.2 shows these metrics correlate with human MOS at around 0.8 to 0.85, and the paper acknowledges visible disagreement between objective and subjective scores.
  • domain assumption The subjective ratings from 20 subjects, after rejecting 148 sequences and 1 subject, are reliable enough to rank 18 quality metrics.
    Section 4.1 describes this protocol; the small subject pool and heavy exclusion are load-bearing for the metric-correlation conclusions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding." pith.science (2026). https://pith.science/paper/BQLKQ73Q

@misc{pith2026250607369,
  author       = {Pith},
  title        = {Pith review of: Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BQLKQ73Q}},
  note         = {Machine review of arXiv:2506.07369}
}
read the original abstract

The rise of deep generative models has greatly advanced video compression, reshaping the paradigm of face video coding through their powerful capability for semantic-aware representation and lifelike synthesis. Generative Face Video Coding (GFVC) stands at the forefront of this revolution, which could characterize complex facial dynamics into compact latent codes for bitstream compactness at the encoder side and leverages powerful deep generative models to reconstruct high-fidelity face signal from the compressed latent codes at the decoder side. As such, this well-designed GFVC paradigm could enable high-fidelity face video communication at ultra-low bitrate ranges, far surpassing the capabilities of the latest Versatile Video Coding (VVC) standard. To pioneer foundational research and accelerate the evolution of GFVC, this paper presents the first comprehensive survey of GFVC technologies, systematically bridging critical gaps between theoretical innovation and industrial standardization. In particular, we first review a broad range of existing GFVC methods with different feature representations and optimization strategies, and conduct a thorough benchmarking analysis. In addition, we construct a large-scale GFVC-compressed face video database with subjective Mean Opinion Scores (MOSs) based on human perception, aiming to identify the most appropriate quality metrics tailored to GFVC. Moreover, we summarize the GFVC standardization potentials with a unified high-level syntax and develop a low-complexity GFVC system which are both expected to push forward future practical deployments and applications. Finally, we envision the potential of GFVC in industrial applications and deliberate on the current challenges and future opportunities.

Figures

Figures reproduced from arXiv: 2506.07369 by the authors.

Figure 1
Figure 1. Generalized pipeline and core techniques of Generative Face Video Coding [ [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Illustrations of compact representations for face signal. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 5
Figure 5. Rate-Distortion performance comparisons of VVC [ [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Rate-Distortion performance comparisons of VVC [ [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 6
Figure 6. Figure 6: Visual quality comparisons of GFVC algorithms and VVC [ [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Comparison of human mean opinion scores (MOSs) against [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Illustrations of SEI-based GFVC workflow at the decoder side. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Rate-Distortion performance comparisons of VVC and four SEI [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Illustrations of the proposed CFTE model with the architectural [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Rate-Distortion performance comparisons of VVC, CFTE and [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Illustrations for potential applications, challenges, and envisions in the domain of generative coding especially for GFVC techniques. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

153 extracted references · 78 canonical work pages

  1. [1]

    Generative face video coding techniques and standardization efforts: A review,

    B. Chen, J. Chen, S. Wang, and Y. Ye, “Generative face video coding techniques and standardization efforts: A review,” inData Compression Conference, 2024, pp. 103–112

  2. [2]

    Ultra-low bitrate video conferencing using deep image animation,

    G. Konuko, G. Valenzise, and S. Lathuili `ere, “Ultra-low bitrate video conferencing using deep image animation,” inIEEE Inter- national Conference on Acoustics, Speech and Signal Processing, 06 2021, pp. 4210–4214

  3. [3]

    Generative compression for face video: A hybrid scheme,

    A. Tang, Y. Huang, J. Ling, Z. Zhang, Y. Zhang, R. Xie, and L. Song, “Generative compression for face video: A hybrid scheme,” inIEEE International Conference on Multimedia and Expo, jul 2022, pp. 1–6

  4. [4]

    Beyond keypoint coding: Temporal evolution inference with compact feature representation for talking face video compression,

    B. Chen, Z. Wang, B. Li, R. Lin, S. Wang, and Y. Ye, “Beyond keypoint coding: Temporal evolution inference with compact feature representation for talking face video compression,” in Data Compression Conference, 2022, pp. 13–22

  5. [5]

    Dynamic multi-reference generative prediction for face video compression,

    Z. Wang, B. Chen, Y. Ye, and S. Wang, “Dynamic multi-reference generative prediction for face video compression,” inIEEE Inter- national Conference on Image Processing, 2022, pp. 896–900

  6. [6]

    Predictive coding for animation-based video compression,

    G. Konuko, S. Lathuili `ere, and G. Valenzise, “Predictive coding for animation-based video compression,” in2023 IEEE Interna- tional Conference on Image Processing (ICIP). IEEE, 2023, pp. 2810– 2814

  7. [7]

    Ultra- low bitrate face video compression based on conversions from 3d keypoints to 2d motion map,

    Z. Wang, B. Chen, S. Wang, S. Wang, Y. Ye, and S. Ma, “Ultra- low bitrate face video compression based on conversions from 3d keypoints to 2d motion map,”IEEE Transactions on Image Processing, vol. 33, pp. 6850–6864, 2024

  8. [8]

    Light-weighted temporal evolution inference for generative face video compres- sion,

    Z. Zhang, B. Chen, S. Yin, S. Wang, and Y. Ye, “Light-weighted temporal evolution inference for generative face video compres- sion,” in2024 IEEE 26th International Workshop on Multimedia Signal Processing (MMSP), 2024, pp. 1–6

Show all 153 references
  1. [9]

    Overview of the H.264/AVC video coding standard,

    T. Wiegand, G. J. Sullivan, G. Bjontegaard, and A. Luthra, “Overview of the H.264/AVC video coding standard,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, no. 7, pp. 560–576, 2003

  2. [10]

    Overview of the High Efficiency Video Coding (HEVC) standard,

    G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the High Efficiency Video Coding (HEVC) standard,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 12, pp. 1649–1668, 2012

  3. [11]

    Overview of the Versatile Video Coding (VVC) standard and its applications,

    B. Bross, Y.-K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sullivan, and J.-R. Ohm, “Overview of the Versatile Video Coding (VVC) standard and its applications,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 10, pp. 3736–3764, 2021

  4. [12]

    Auto-encoding variational bayes,

    D. P . Kingma and M. Welling, “Auto-encoding variational bayes,” inInternational Conference on Learning Representations, 2014, p. 14

  5. [13]

    Generative ad- versarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde- Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative ad- versarial nets,”Advances in Neural Information Processing Systems, vol. 27, 2014

  6. [14]

    Diffusion models beat gans on image synthesis,

    P . Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” inAdvances in Neural Information Processing Systems, vol. 34, 2021, pp. 8780–8794

  7. [15]

    Block-based motion estima- tion algorithms—a survey,

    M. Jakubowski and G. Pastuszak, “Block-based motion estima- tion algorithms—a survey,”Opto-Electronics Review, vol. 21, pp. 86–102, 2013

  8. [16]

    Transform coding in the vvc standard,

    X. Zhao, S.-H. Kim, Y. Zhao, H. E. Egilmez, M. Koo, S. Liu, J. Lainema, and M. Karczewicz, “Transform coding in the vvc standard,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 10, pp. 3878–3890, 2021

  9. [17]

    Synthetic highs—an exper- imental tv bandwidth reduction system,

    W. Schreiber, C. Knapp, and N. Kay, “Synthetic highs—an exper- imental tv bandwidth reduction system,”Journal of the SMPTE, vol. 68, no. 8, pp. 525–537, 1959

  10. [18]

    Visual communication at very low data rates,

    D. Pearson and J. Robinson, “Visual communication at very low data rates,”Proceedings of the IEEE, vol. 73, no. 4, pp. 795–812, 1985

  11. [19]

    Object-oriented analysis-synthesis coding of moving images,

    H. G. Musmann, M. H ¨otter, and J. Ostermann, “Object-oriented analysis-synthesis coding of moving images,”Signal Processing: Image Communication, vol. 1, no. 2, pp. 117–138, 1989

  12. [20]

    Head pose computation for very low bit-rate video coding,

    R. Lopez and T. Huang, “Head pose computation for very low bit-rate video coding,” inInternational Conference on Computer Analysis of Images and Patterns, 1995, pp. 440–447

  13. [21]

    Model- based/waveform hybrid coding for videotelephone images,

    Y. Nakaya, Y. Chuah, and H. Harashima, “Model- based/waveform hybrid coding for videotelephone images,” in International Conference on Acoustics, Speech, and Signal Processing, 1991, pp. 2741–2744 vol.4

  14. [22]

    Model-based image coding advanced video coding techniques for very low bit-rate applications,

    K. Aizawa and T. Huang, “Model-based image coding advanced video coding techniques for very low bit-rate applications,” Proceedings of the IEEE, vol. 83, no. 2, pp. 259–271, 1995

  15. [23]

    Beyond the pixel world: A novel acoustic-based face anti-spoofing system for smartphones,

    C. Kong, K. Zheng, S. Wang, A. Rocha, and H. Li, “Beyond the pixel world: A novel acoustic-based face anti-spoofing system for smartphones,”IEEE Transactions on Information Forensics and Security, vol. 17, pp. 3238–3253, 2022

  16. [24]

    Pixel- inconsistency modeling for image manipulation localization,

    C. Kong, A. Luo, S. Wang, H. Li, A. Rocha, and A. C. Kot, “Pixel- inconsistency modeling for image manipulation localization,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  17. [25]

    Recurrent face aging with hierarchical autoregressive memory,

    W. Wang, Y. Yan, Z. Cui, J. Feng, S. Yan, and N. Sebe, “Recurrent face aging with hierarchical autoregressive memory,”IEEE trans- actions on pattern analysis and machine intelligence, vol. 41, no. 3, pp. 654–668, 2018

  18. [26]

    Assessing face image quality: A large-scale database and a transformer method,

    T. Liu, S. Li, M. Xu, L. Yang, and X. Wang, “Assessing face image quality: A large-scale database and a transformer method,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 3981–4000, 2024

  19. [27]

    Memories are one-to-many mapping alleviators in talking face generation,

    A. Tang, T. He, X. Tan, J. Ling, R. Li, S. Zhao, J. Bian, and L. Song, “Memories are one-to-many mapping alleviators in talking face generation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 8758–8770, 2024

  20. [28]

    RefineFace: Refinement neural network for high performance face detection,

    S. Zhang, C. Chi, Z. Lei, and S. Z. Li, “RefineFace: Refinement neural network for high performance face detection,”IEEE Trans- actions on Pattern Analysis and Machine Intelligence, vol. 43, no. 11, pp. 4008–4020, 2021. 17

  21. [29]

    First order motion model for image animation,

    A. Siarohin, S. Lathuili `ere, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,”Advances in Neural Information Processing Systems, vol. 32, pp. 7137–7147, 2019

  22. [30]

    An- imating arbitrary objects via deep motion transfer,

    A. Siarohin, S. Lathuili `ere, S. Tulyakov, E. Ricci, and N. Sebe, “An- imating arbitrary objects via deep motion transfer,” inProceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 2372–2381

  23. [31]

    Motion representations for articulated animation,

    A. Siarohin, O. J. Woodford, J. Ren, M. Chai, and S. Tulyakov, “Motion representations for articulated animation,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 13 648–13 657

  24. [32]

    DaGAN++: Depth-aware genera- tive adversarial network for talking head video generation,

    F.-T. Hong, L. Shen, and D. Xu, “DaGAN++: Depth-aware genera- tive adversarial network for talking head video generation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 2997–3012, 2024

  25. [33]

    Compact temporal trajectory representation for talking face video compression,

    B. Chen, Z. Wang, B. Li, S. Wang, and Y. Ye, “Compact temporal trajectory representation for talking face video compression,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 11, pp. 7009–7023, 2023

  26. [34]

    A hybrid deep animation codec for low-bitrate video conferencing,

    G. Konuko, S. Lathuili `ere, and G. Valenzise, “A hybrid deep animation codec for low-bitrate video conferencing,” inIEEE International Conference on Image Processing, 2022

  27. [35]

    Beyond gfvc: A progressive face video compression framework with adaptive visual tokens,

    B. Chen, S. Yin, Z. Zhang, J. Chen, R.-L. Liao, L. Zhu, S. Wang, and Y. Ye, “Beyond gfvc: A progressive face video compression framework with adaptive visual tokens,” in2025 Data Compres- sion Conference (DCC), 2025, pp. 163–172

  28. [36]

    Pleno-Generation: A scalable generative face video compression framework with bandwidth intelligence

    B. Chen, H. Zhu, S. Yin, L. Zhu, J. Chen, R.-L. Liao, S. Wang, and Y. Ye, “Pleno-Generation: A scalable generative face video compression framework with bandwidth intelligence.”

  29. [37]

    Low bandwidth video-chat compression using deep generative mod- els,

    M. Oquab, P . Stock, D. Haziza, T. Xu, P . Zhang, O. Celebi, Y. Hasson, P . Labatut, B. Bose-Kolanu, T. Peyronelet al., “Low bandwidth video-chat compression using deep generative mod- els,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2...

  30. [38]

    Gemino: Practical and ro- bust neural compression for video conferencing,

    V . Sivaraman, P . Karimi, V . Venkatapathy, M. Khani, S. Fouladi, M. Alizadeh, F. Durand, and V . Sze, “Gemino: Practical and ro- bust neural compression for video conferencing,” in21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), 2024, pp. 569–590

  31. [39]

    Interactive face video coding: A generative compression framework,

    B. Chen, Z. Wang, B. Li, S. Wang, S. Wang, and Y. Ye, “Interactive face video coding: A generative compression framework,”IEEE Transactions on Image Processing, vol. 34, pp. 2910–2925, 2025

  32. [40]

    A generative compression framework for low bandwidth video conference,

    D. Feng, Y. Huang, Y. Zhang, J. Ling, A. Tang, and L. Song, “A generative compression framework for low bandwidth video conference,” inIEEE International Conference on Multimedia and Expo Workshop, 2021, pp. 1–6

  33. [41]

    One-shot free-view neural talking-head synthesis for video conferencing,

    T. Wang, A. Mallya, and M. Liu, “One-shot free-view neural talking-head synthesis for video conferencing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2021, pp. 10 039–10 049

  34. [42]

    Semantic neural rendering- based video coding: Towards ultra-low bitrate video conferenc- ing,

    Y. Hu, Y. Xu, J. Chang, and J. Zhang, “Semantic neural rendering- based video coding: Towards ultra-low bitrate video conferenc- ing,” inData Compression Conference, 2022, pp. 456–456

  35. [43]

    Robust ultralow bitrate video conferencing with second order motion coherency,

    Z. Chen, M. Lu, H. Chen, and Z. Ma, “Robust ultralow bitrate video conferencing with second order motion coherency,” in2022 IEEE 24th International Workshop on Multimedia Signal Processing (MMSP), 2022, pp. 1–6

  36. [44]

    Thin-plate spline motion model for image animation,

    J. Zhao and H. Zhang, “Thin-plate spline motion model for image animation,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 3647–3656

  37. [45]

    Generative human video compression with multi-granularity temporal trajectory factor- ization,

    S. Yin, B. Chen, S. Wang, and Y. Ye, “Generative human video compression with multi-granularity temporal trajectory factor- ization,”arXiv preprint arXiv:2410.10171, 2024

  38. [46]

    Bidirectional learned facial animation codec for low bitrate talking head videos,

    R. Takahashi, R. Morita, F. Kimishima, K. Iwama, and J. Zhou, “Bidirectional learned facial animation codec for low bitrate talking head videos,” in2025 Data Compression Conference (DCC), 2025, pp. 401–401

  39. [47]

    Affine transformation-based generative face video compression,

    X. Lin, X. Song, X. Zuo, X. Li, D. Gao, X. Xie, and G. Shi, “Affine transformation-based generative face video compression,” in 2025 Data Compression Conference (DCC), 2025, pp. 387–387

  40. [48]

    Neural face video compression using multiple views,

    A. Volokitin, S. Brugger, A. Benlalah, S. Martin, B. Amberg, and M. Tschannen, “Neural face video compression using multiple views,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1738–1742

  41. [49]

    A motion representation ai based video conference solution,

    X. Gu, C. Zhang, and Y. Sun, “A motion representation ai based video conference solution,” in2022 IEEE 24th International Work- shop on Multimedia Signal Processing (MMSP), 2022, pp. 1–1

  42. [50]

    Compressing video calls using synthetic talking heads,

    M. Agarwal, A. Gupta, R. Mukhopadhyay, V . P . Namboodiri, and C. Jawahar, “Compressing video calls using synthetic talking heads,” inBritish Machine Vision Conference, 2023

  43. [51]

    Enabling translatability of generative face video coding: A unified face feature transcoding framework,

    S. Yin, B. Chen, S. Wang, and Y. Ye, “Enabling translatability of generative face video coding: A unified face feature transcoding framework,” inData Compression Conference, 2024, pp. 113–122

  44. [52]

    Improved predictive coding for animation-based video compression,

    G. Konuko and G. Valenzise, “Improved predictive coding for animation-based video compression,” in2024 12th European Workshop on Visual Information Processing (EUVIP), 2024, pp. 1– 6

  45. [53]

    Multi-reference generative face video compres- sion with contrastive learning,

    G. Konukoet al., “Multi-reference generative face video compres- sion with contrastive learning,” in2024 IEEE 26th International Workshop on Multimedia Signal Processing (MMSP), 2024, pp. 1–6

  46. [54]

    Enhanced multi- resolution generative face video compression,

    R. Zhou, R.-L. Liao, B. Chen, Y. Ye, and J. Chen, “Enhanced multi- resolution generative face video compression,” in2024 IEEE 26th International Workshop on Multimedia Signal Processing (MMSP), 2024, pp. 1–4

  47. [55]

    Improving reconstruc- tion fidelity in generative face video coding using high-frequency shuttling,

    G. Konuko, G. Valenzise, and A. Trioux, “Improving reconstruc- tion fidelity in generative face video coding using high-frequency shuttling,” in2024 IEEE International Conference on Visual Commu- nications and Image Processing (VCIP), 2024, pp. 1–5

  48. [56]

    Gemino: practical and robust neural compression for video conferencing,

    V . Sivaraman, P . Karimi, V . Venkatapathy, M. Khani, S. Fouladi, M. Alizadeh, F. Durand, and V . Sze, “Gemino: practical and robust neural compression for video conferencing,” inNSDI’24, 2024

  49. [57]

    Voxceleb2: Deep speaker recognition,

    J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” inInterspeech, 2018

  50. [58]

    CelebV-HQ: A large-scale video facial attributes dataset,

    H. Zhu, W. Wu, W. Zhu, L. Jiang, S. Tang, L. Zhang, Z. Liu, and C. C. Loy, “CelebV-HQ: A large-scale video facial attributes dataset,” inEuropean Conference on Computer Vision, 2022

  51. [59]

    Test conditions and evaluation pro- cedures for generative face video coding,

    S. McCarthy and B. Chen, “Test conditions and evaluation pro- cedures for generative face video coding,”The Joint Video Experts Team of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET- AJ2035, November 2024

  52. [60]

    The unreasonable effectiveness of deep features as a perceptual met- ric,

    R. Zhang, P . Isola, A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual met- ric,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018

  53. [61]

    Image quality assessment: Unifying structure and texture similarity,

    K. Ding, K. Ma, S. Wang, and E. Simoncelli, “Image quality assessment: Unifying structure and texture similarity,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020

  54. [62]

    Image quality assessment : From error visibility to structural similarity,

    Z. Wang, “Image quality assessment : From error visibility to structural similarity,”IEEE Transactions on Image Processing, 2004

  55. [63]

    Calculation of average psnr differences between rd-curves,

    G. Bjontegaard, “Calculation of average psnr differences between rd-curves,”ITU SG16 Doc. VCEG-M33, 2001

  56. [64]

    AHG 16: Multi-resolution models for generative face video compression ,

    S. Yin, S. Wang, B. Chen, Y. Ye, G.Konuko, and G. Valenzise, “AHG 16: Multi-resolution models for generative face video compression ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0052, November 2024

  57. [65]

    Optical flow estimation using a spa- tial pyramid network,

    A. Ranjan and M. J. Black, “Optical flow estimation using a spa- tial pyramid network,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, jul 2017, pp. 2720– 2729

  58. [66]

    Study of subjective and objective quality assessment of video,

    K. Seshadrinathan, R. Soundararajan, A. C. Bovik, and L. K. Cormack, “Study of subjective and objective quality assessment of video,”IEEE Transactions on Image Processing, vol. 19, no. 6, pp. 1427–1441, 2010

  59. [67]

    Methodology for the subjective assessment of the quality of television pictures,

    Series B., “Methodology for the subjective assessment of the quality of television pictures,”Recommendation ITU-R BT, vol. 500, no. 13, 2012

  60. [68]

    Ars by vqeg,

    “Ars by vqeg,”http://www.its.bldrdoc.gov/vqeg/, 2022

  61. [69]

    Percep- tual quality assessment of face video compression: A benchmark and an effective method,

    Y. Li, B. Chen, B. Chen, M. Wang, S. Wang, and W. Lin, “Percep- tual quality assessment of face video compression: A benchmark and an effective method,”IEEE Transactions on Multimedia, pp. 1–13, 2024

  62. [70]

    Mean squared error: Love it or leave it? a new look at signal fidelity measures,

    W. Zhou and A. C. Bovik, “Mean squared error: Love it or leave it? a new look at signal fidelity measures,”IEEE Signal Processing Magazine, vol. 26, no. 1, pp. 98–117, 2009

  63. [71]

    Vmaf reproducibility: Validating a perceptual prac- tical video quality metric,

    R. Rassool, “Vmaf reproducibility: Validating a perceptual prac- tical video quality metric,” in2017 IEEE International Symposium on Broadband Multimedia Systems and Broadcasting, 2017, pp. 1–2

  64. [72]

    Fsim: A feature similarity index for image quality assessment,

    L. Zhang, L. Zhang, X. Mou, and D. Zhang, “Fsim: A feature similarity index for image quality assessment,”IEEE transactions on Image Processing, vol. 20, no. 8, pp. 2378–2386, 2011

  65. [73]

    Vsi: A visual saliency-induced in- dex for perceptual image quality assessment,

    L. Zhang, Y. Shen, and H. Li, “Vsi: A visual saliency-induced in- dex for perceptual image quality assessment,”IEEE Transactions on Image processing, vol. 23, no. 10, pp. 4270–4281, 2014. 18

  66. [74]

    Perceptual losses for real- time style transfer and super-resolution,

    J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real- time style transfer and super-resolution,” inComputer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14. Springer, 2016, pp. 694–711

  67. [75]

    Topiq: A top-down approach from semantics to distortions for image quality assessment,

    C. Chen, J. Mo, J. Hou, H. Wu, L. Liao, W. Sun, Q. Yan, and W. Lin, “Topiq: A top-down approach from semantics to distortions for image quality assessment,”IEEE Transactions on Image Processing, 2024

  68. [76]

    Ifqa: Interpretable face quality assessment,

    B. Jo, D. Cho, I. K. Park, and S. Hong, “Ifqa: Interpretable face quality assessment,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, pp. 3444–3453

  69. [77]

    Fvd: A new metric for video gen- eration,

    T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Fvd: A new metric for video gen- eration,” in7th International Conference on Learning Representations workshop, 2019

  70. [78]

    Maniqa: Multi-dimension attention network for no-reference image quality assessment,

    S. Yang, T. Wu, S. Shi, S. Lao, Y. Gong, M. Cao, J. Wang, and Y. Yang, “Maniqa: Multi-dimension attention network for no-reference image quality assessment,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1191–1200

  71. [79]

    No-reference image quality assessment in the spatial domain,

    A. Mittal, A. K. Moorthy, and A. C. Bovik, “No-reference image quality assessment in the spatial domain,”IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, 2012

  72. [80]

    Video quality assessment by reduced reference spatio-temporal entropic differencing,

    R. Soundararajan and A. C. Bovik, “Video quality assessment by reduced reference spatio-temporal entropic differencing,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 23, no. 4, pp. 684–694, 2012

  73. [81]

    Optical flow estimation using a spatial pyramid network,

    A. Ranjan and M. J. Black, “Optical flow estimation using a spatial pyramid network,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4161–4170

  74. [82]

    AHG9: Generative face video SEI message,

    B. Chen, J. Chen, S. Wang, Y. Ye, and S. Wang, “AHG9: Generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AC0088, January 2023

  75. [83]

    AHG9: Common SEI message of generative face video,

    B. Chen, J. Chen, Y. Ye, and S. Wang, “AHG9: Common SEI message of generative face video,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AD0051, April 2023

  76. [84]

    AHG9: Generative face video SEI message,

    S. McCarthy, P . Yin, G.-M. Su, A. K. Choudhury, and W. Husak, “AHG9: Generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AE0080, July 2023

  77. [85]

    AHG9: A study on generative face video SEI message,

    H.-B. Teo, J.-Y. Thong, K. Jayashree, C.-S. Lim, and K. Abe, “AHG9: A study on generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AE0088, July 2023

  78. [86]

    AHG9: Common text for proposed generative face video SEI message,

    B. Chen, J. Chen, Y. Ye, S. Wang, S. McCarthy, P . Yin, G.-M. Su, A. K. Choudhury, and W. Husak, “AHG9: Common text for proposed generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AE0280, July 2023

  79. [87]

    AHG9: Common SEI message of generative face video,

    B. Chen, J. Chen, Y. Ye, and S. Wang, “AHG9: Common SEI message of generative face video,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AE0083, July 2023

  80. [88]

    AHG9/AHG16: Software implementation of generative face video SEI message,

    B. Chen, J. Chen, R. Zou, Y. Ye, R.-L. Liao, and S. Wang, “AHG9/AHG16: Software implementation of generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0239, April 2024

  81. [89]

    A study on decoder interoperability of generative face video compression,

    B. Chen, S. Yin, J. Chen, Y. Ye, and S. Wang, “A study on decoder interoperability of generative face video compression,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET– AF0048, October 2023

  82. [90]

    AHG9: On face motion information for generative face video,

    H.-B. Teo, J.-Y. Thong, K. Jayashree, and K. A. C.-S. Lim, “AHG9: On face motion information for generative face video,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET– AF0146, October 2023

  83. [91]

    AHG9: Common text for proposed generative face video SEI message,

    B. Chen, J. Chen, Y. Ye, S. Wang, S. McCarthy, P . Yin, G.-M. Su, A. K. Choudhury, W. Husak, and G. J. Sullivan, “AHG9: Common text for proposed generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AF0234, October 2023

  84. [92]

    AHG 16: Proposed common software tools and testing conditions for generative face video compression,

    B. Chen, J. Chen, R.-L. Liao, Y. Ye, and S. Wang, “AHG 16: Proposed common software tools and testing conditions for generative face video compression,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AG0042, January 2024

  85. [93]

    AHG 16: Interoperability study on parameter translator of generative face video coding,

    S. Yin, B. Chen, J. Chen, R.-L. Liao, Y. Ye, and S. Wang, “AHG 16: Interoperability study on parameter translator of generative face video coding,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AG0048, January 2024

  86. [94]

    AHG 16: Refined parameter translator of generative face video coding,

    S. Yin, S. Wang, Z. Zhang, B. Chen, Y. Ye, R.-L. Liao, and J. Chen, “AHG 16: Refined parameter translator of generative face video coding,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AL0147, April 2025

  87. [95]

    AHG9: On the generative face video SEI message,

    M. M. Hannuksela, F. Cricri, and H. Zhang, “AHG9: On the generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AG0087, January 2024

  88. [96]

    AHG9: Usage of the neural- network post-filter characteristics SEI message to define the generator NN of the generative face video SEI message,

    M. M. Hannuksela, F. Cricriet al., “AHG9: Usage of the neural- network post-filter characteristics SEI message to define the generator NN of the generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AG0087, January 2024

  89. [97]

    AHG9/AHG16: Common text for proposed generative face video SEI message,

    J. Chen, B. Chen, Y. Ye, S. Yin, S. Wang, S. McCarthy, P . Yin, G.-M. Su, A. K. Choudhury, W. Husak, and G. J. Sullivan, “AHG9/AHG16: Common text for proposed generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET–AG0203, January 2024

  90. [98]

    AHG16: Depth- wise separable convolution for generative face video compres- sion,

    R. Zou, R.-L. Liao, B. Chen, Y. Ye, and J. Chen, “AHG16: Depth- wise separable convolution for generative face video compres- sion,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AG0139, January 2024

  91. [99]

    AHG16: Study text for common test conditions and evaluation procedures for generative face video coding (draft 1),

    S. McCarthy, P . Yin, B. Chen, Y. Ye, and S. Wang, “AHG16: Study text for common test conditions and evaluation procedures for generative face video coding (draft 1),”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AG0187, January 2024

  92. [100]

    Test conditions and evaluation pro- cedures for generative face video coding,

    S. McCarthy and B. Chen, “Test conditions and evaluation pro- cedures for generative face video coding,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AG2035, January 2024

  93. [101]

    AHG9/AHG16: Comments on generative face video SEI,

    S. Deshpande, “AHG9/AHG16: Comments on generative face video SEI,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0053, April 2024

  94. [102]

    AHG16: On the generative face video SEI message,

    H.-B. Teo, J.-Y. Thong, K. Abe, C.-S. Lim, and K. Jayashree, “AHG16: On the generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0054, April 2024

  95. [103]

    AHG9/AHG16: On key point coordinates calculation for the generative face video SEI message,

    K. Yang, Y.-K. Wang, Y. Xu, and Y. Li, “AHG9/AHG16: On key point coordinates calculation for the generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0148, April 2024

  96. [104]

    AHG16: Removal of the analysis model on the decoder side and lightweight design,

    R. Zou, R.-L. Liao, B. Chen, Y. Ye, and J. Chen, “AHG16: Removal of the analysis model on the decoder side and lightweight design,”The Joint Video Experts Team of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0109, April 2024

  97. [105]

    AHG16: Scalable representation and layered reconstruction for generative face video compression,

    B. Chen, Y. Ye, J. Chen, R.-L. Liao, S. Yin, and S. Wang, “AHG16: Scalable representation and layered reconstruction for generative face video compression,”The Joint Video Experts Team of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0110, April 2024

  98. [106]

    AHG16: Lightweight CFTE with multi-resolution support,

    R. Zou, B. Chen, R.-L. Liao, J. Chen, and Y. Ye, “AHG16: Lightweight CFTE with multi-resolution support,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0113, April 2024

  99. [107]

    AHG 16: Updated common software tools for generative face video compression,

    B. Chen, Y. Ye, G.Konuko, G. Valenzise, S. Yin, and S. Wang, “AHG 16: Updated common software tools for generative face video compression,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0114, April 2024

  100. [108]

    AHG9/AHG16: Showcase for picture fusion for generative face video sei message,

    S. Gehlot, G. Su, P . Yin, S. McCarthy, and G. J. Sullivan, “AHG9/AHG16: Showcase for picture fusion for generative face video sei message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0118, April 2024

  101. [109]

    AHG9/AHG16: The SEI message design for scalable repre- sentation and layered reconstruction for generative face video compression,

    J. Chen, B. Chen, Y. Ye, R.-L. Liao, S. Yin, and S. Wang, “AHG9/AHG16: The SEI message design for scalable repre- sentation and layered reconstruction for generative face video compression,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0127, April 2024

  102. [110]

    AHG9/AHG16: Pupil position SEI message for generative face video,

    A. Trioux, Y. Yao, F. Ma, F. Yang, and Z. W. F. Xing, “AHG9/AHG16: Pupil position SEI message for generative face video,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AH0138, April 2024

  103. [111]

    AHG9/AHG16: Update on pupil position SEI message for gen- 19 erative face video,

    F. Ma, A. Trioux, Y. Gao, Y. Yao, F. Yang, F. Xing, and Z. Wang, “AHG9/AHG16: Update on pupil position SEI message for gen- 19 erative face video,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0137, July 2024

  104. [112]

    AHG9/AHG16: Refined Methodology for Pupil Position SEI Message for Genera- tive Face Video,

    F. Ma, A. Trioux, F. Yang, F. Xing, and Z. Wang, “AHG9/AHG16: Refined Methodology for Pupil Position SEI Message for Genera- tive Face Video,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0135, November 2024

  105. [113]

    AhG16: Higher-resolution Test Sequences and Test Results for Generative Face Video Compression,

    B. Chen, S. Yin, Y. Ye, R.-L. Liao, J. Chen, and S. Wang, “AhG16: Higher-resolution Test Sequences and Test Results for Generative Face Video Compression,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0047, July 2024

  106. [114]

    AHG9/AHG16: Performance results of generative face video SEI message on high-resolution sequences,

    B. Chen, Y. Ye, R.-L. Liao, J. Chen, S. Yin, and S. Wang, “AHG9/AHG16: Performance results of generative face video SEI message on high-resolution sequences,”The JVET of ITU- T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0209, November 2024

  107. [115]

    AhG16: Improving GFVC Performance with Diverse Training Data,

    B. Chen, Y. Ye, J. Chen, R.-L. Liao, Z. Zhang, S. Yin, and S. Wang, “AhG16: Improving GFVC Performance with Diverse Training Data,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0048, July 2024

  108. [116]

    AHG9/AHG16: On generative face video SEI messages,

    J. Chen, B. Chen, Y. Ye, S. Yin, S. Wang, P . Yin, S. McCarthy, and H.-B. Teo, “AHG9/AHG16: On generative face video SEI messages,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0156, July 2024

  109. [117]

    AHG9: Some comments and editorial changes on the GFV and GFVE SEI messages,

    Y.-K. Wang, K. Yang, Y. Li, Y. Xu, and J. Chen, “AHG9: Some comments and editorial changes on the GFV and GFVE SEI messages,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0184, July 2024

  110. [118]

    AHG9: On signalling of and/or specifying GFV translator, generator, and enhancer NNs,

    Y.-K. Wang, K. Yang, Y. Li, and Y. Xu, “AHG9: On signalling of and/or specifying GFV translator, generator, and enhancer NNs,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0189, July 2024

  111. [119]

    AHG9: On GFV SEI persistence and picture presence,

    Y.-K. Wang, K. Yang, Y. Liet al., “AHG9: On GFV SEI persistence and picture presence,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0190, July 2024

  112. [120]

    AHG9: On GFV facial parameters signalling,

    Y.-K. Wang, K. Yang, Y. Li, and Y. Xu, “AHG9: On GFV facial parameters signalling,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0191, July 2024

  113. [121]

    AHG9/AHG16: On no display flag in generative face video SEI message,

    S. McCarthy, S. Gehlot, G. J. Sullivan, and P . Yin, “AHG9/AHG16: On no display flag in generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0193, July 2024

  114. [122]

    AHG9: On picture order count and timing information for gfv generated pictures,

    L. Chen, O. Chubach, Y.-W. Huang, and S. Lei, “AHG9: On picture order count and timing information for gfv generated pictures,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0186, July 2024

  115. [123]

    AHG9: On GFV picture order and timing ,

    Y.-K. Wang, K. Yang, Y. Li, and Y. Xu, “AHG9: On GFV picture order and timing ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0192, July 2024

  116. [124]

    AHG9/AHG16: On chroma key fusion for the generative face video SEI message,

    S. Gehlot, G.-M. Su, P . Yin, S. McCarthy, and G. J. Sullivan, “AHG9/AHG16: On chroma key fusion for the generative face video SEI message,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0194, July 2024

  117. [125]

    AhG9/16: Updates on GFVC common software tools and GFV SEI message ,

    B. Chen, J. Chen, Y. Ye, R.-L. Liao, S.Yin, and S. Wang, “AhG9/16: Updates on GFVC common software tools and GFV SEI message ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AI0195, July 2024

  118. [126]

    AHG9: On the GFV SEI message ,

    Y.-K. Wang, Y. Li, K. Yang, Y. Xu, J. Chen, B. Chen, and Y. Ye, “AHG9: On the GFV SEI message ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0051, November 2024

  119. [127]

    AHG9: On specifying the output order of GFV-generated pictures ,

    L. Chen, O. Chubach, Y.-W. Huang, and S. Lei, “AHG9: On specifying the output order of GFV-generated pictures ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0069, November 2024

  120. [128]

    AHG9: On signalling of instance count in generative face video SEI message ,

    H. Tan, J. Lee, J. Nam, C. Kim, J. Lim, and S. Kim, “AHG9: On signalling of instance count in generative face video SEI message ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0108, November 2024

  121. [129]

    AHG9: On the presence of translator NN filter in GFV SEI message ,

    H. Tan, J. Lee, J. Namet al., “AHG9: On the presence of translator NN filter in GFV SEI message ,”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0111, November 2024

  122. [130]

    AHG9: Usage of NNPFC SEI message to define the generator NN of the GFV SEI message ,

    M. M. Hannuksela and F. Cricri, “AHG9: Usage of NNPFC SEI message to define the generator NN of the GFV SEI message ,” The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ0132, November 2024

  123. [131]

    AHG9/AHG16: Performance results and suggestions on generative face video SEI message,

    B. Chen, Y. Ye, J. Chen, R.-L. Liao, S. Yin, Z. Zhang, S. Wang, K. Yang, Y. Li, Y. Xu, Y.-K. Wang, S. Gehlot, G.-M. Su, P . Yin, G. J. Sullivan, S. McCarthy, and H.-B. Teo, “AHG9/AHG16: Performance results and suggestions on generative face video SEI message,”The JVET of ITU-T...

  124. [132]

    AHG16: GFVC Extension of the VVC Standard ,

    L. Liu and C. Jung, “AHG16: GFVC Extension of the VVC Standard ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0068, January 2025

  125. [133]

    AHG16: Further experiments on VVC GFVC,

    L. Liuet al., “AHG16: Further experiments on VVC GFVC,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AL0101, April 2025

  126. [134]

    AHG16: QP-Adaptive GFVC ,

    W. Kang, L. Liu, and C. Jung, “AHG16: QP-Adaptive GFVC ,” The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0069, January 2025

  127. [135]

    AHG9/AHG16: SEI message and further experiments on QP-adaptive GFVC,

    W. Kang, L. Liuet al., “AHG9/AHG16: SEI message and further experiments on QP-adaptive GFVC,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AL0102, April 2025

  128. [136]

    AHG9: Com- ments on the GFV and GFVE SEI messages ,

    M. M. Hannuksela, J. Chen, B. Chen, and Y. Ye, “AHG9: Com- ments on the GFV and GFVE SEI messages ,”The JVET of ITU- T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0080, January 2025

  129. [137]

    AHG9: On miscellaneous aspects of GFV and DSCI SEI messages ,

    J. Lee, H. Tan, C. Kim, J. Nam, J. Lim, and S. Kim, “AHG9: On miscellaneous aspects of GFV and DSCI SEI messages ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET- AK0127, January 2025

  130. [138]

    AHG9: Editorial updates for GFV SEI message ,

    J. Lee, H. Tan, C. Kimet al., “AHG9: Editorial updates for GFV SEI message ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0128, January 2025

  131. [139]

    AHG9: Comments on Genera- tive Face Video SEI ,

    A. C. Sidiya and S. Deshpande, “AHG9: Comments on Genera- tive Face Video SEI ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0154, January 2025

  132. [140]

    AHG9: Supplementation of value range definition and editorial bugs fixes on the GFVE pupil position SEI messages ,

    F. Ma, A. Trioux, F. Yang, B. Li, F. Xing, and Z. Wang, “AHG9: Supplementation of value range definition and editorial bugs fixes on the GFVE pupil position SEI messages ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0164, January 2025

  133. [141]

    AHG9: Semantics and syntax fixes for generative face video SEI message ,

    J. Chen, B. Chen, and Y. Ye, “AHG9: Semantics and syntax fixes for generative face video SEI message ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0238, January 2025

  134. [142]

    AHG9: On generative face video enhance- ment SEI message,

    J. Chen, B. Chenet al., “AHG9: On generative face video enhance- ment SEI message,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0239, January 2025

  135. [143]

    AHG9: On timing information and order of pictures in GFV SEI message,

    J. Lee, H. Tan, C. Kim, J. Nam, J. Lim, and S. Kim, “AHG9: On timing information and order of pictures in GFV SEI message,” The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AK0124, January 2025

  136. [144]

    AHG9/AHG16: Colour calibration for generative face video coding ,

    J. Chen, B. Chen, Y. Ye, S. Yin, and S. Wang, “AHG9/AHG16: Colour calibration for generative face video coding ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AL0156, April 2025

  137. [145]

    AHG9: Generative face video and generative face video enhancement SEI messages for HEVC and AVC ,

    J. Chen, B. Chen, Y. Ye, S. Gehlot, G.-M. Su, P . Yin, S. McCarthy, A. Trioux, F. Yang, Y.-K. Wang, H.-B. Teo, J.-Y. Thong, K. Abe, Y. Xu, K. Yang, and Y. Li, “AHG9: Generative face video and generative face video enhancement SEI messages for HEVC and AVC ,”The JVET of ITU-T S...

  138. [146]

    AHG9: Further fixes and cleanup on GFV and GFVE SEI messages ,

    J. Chen, B. Chen, and Y. Ye, “AHG9: Further fixes and cleanup on GFV and GFVE SEI messages ,”The JVET of ITU-T SG 21 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AL0155, April 2025

  139. [147]

    Additional SEI messages for VSEI version 4 (Draft 4),

    J. Boyce, J. Chen, S. Deshpandeet al., “Additional SEI messages for VSEI version 4 (Draft 4),”The JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, doc. no. JVET-AJ2006, November 2024

  140. [148]

    Standardizing generative face video compression using supplemental enhancement informa- tion,

    B. Chen, Y. Ye, J. Chen, R.-L. Liao, S. Yin, S. Wang, K. Yang, Y. Li, Y. Xu, Y.-K. Wanget al., “Standardizing generative face video compression using supplemental enhancement informa- tion,”arXiv preprint arXiv:2410.15105, 2024

  141. [149]

    Mobilenets: Efficient convolutional neural networks for mobile vision applications,

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017

  142. [150]

    Learning efficient convolutional networks through network slimming,

    Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, and C. Zhang, “Learning efficient convolutional networks through network slimming,” in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 2736–2744

  143. [151]

    Compressing scene dynamics: A generative approach,

    S. Yin, Z. Zhang, B. Chen, S. Wang, and Y. Ye, “Compressing scene dynamics: A generative approach,” inData Compression Conference, 2025. 20 Bolin Chenreceived the B.S. degree in commu- nication engineering from Fuzhou University in July 2020 and the Ph.D. degree in computer sc...

  144. [2011]

    He joined the French Centre National de la Recherche Scientifique (CNRS) in 2012. His research interests span different fields of image and video processing, including traditional and learning-based image and video compression, light fields and point cloud coding, image/ video...

  145. [2021]

    Yan Ye(Senior Member, IEEE) received the B.S

    He serves as an Associate Editor for IEEE TRANSACTIONS ON CIRCUITS ANDSYSTEMS FORVIDEOTECHNOLOGY, IEEE TRANSAC- TIONS ONIMAGEPROCESSING, IEEE TRANSACTIONS ONMULTIMEDIA and IEEE TRANSACTIONS ONCYBERNETICS. Yan Ye(Senior Member, IEEE) received the B.S. and M.S. degrees in electr...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.