REVIEW 3 major objections 7 minor 107 references
Ultra-low-bitrate image decoding can be cast as next-frame prediction from a compact anchor using a video diffusion prior, cutting bitrate over half versus DiffC while decoding up to five times faster.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 20:39 UTC pith:2TEUNCE7
load-bearing objection Solid empirical ULB-IC recipe: visible VAE anchor + LoRA-adapted CogVideoX next-frame decoding with a latent bypass, reporting large BD-rate savings and real one-step speedup. the 3 major comments →
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Generative ultra-low-bitrate image compression can be reframed as next-frame prediction: a compact anchor that preserves scene layout is decoded first, and a video diffusion model evolves that visible anchor into a high-fidelity reconstruction, yielding better source faithfulness and perceptual quality than image-diffusion codecs that start from noise plus weak conditions.
What carries the argument
Next-frame decoding with a two-stage adapted video diffusion model (NeFIC): Stage I LoRA-adapts the model for single-interval, motion-suppressed detail synthesis from the anchor; Stage II adds conditional anchor encoding and a bypass-refinement map so generation collapses to a single informative step instead of multi-step noise.
Load-bearing premise
The method assumes that a video model pretrained on long multi-frame clips can be LoRA-adapted into a faithful single-step image decoder that only refines detail and does not introduce motion, flicker, or semantic drift once the anchor and bypass are supplied.
What would settle it
Train and evaluate the same pipeline on held-out high-resolution images outside the Flickr2W crop distribution; if LPIPS/DISTS or visual fidelity fall to or below strong image-diffusion baselines, or if one-step decoding reintroduces large semantic errors relative to multi-step, the central transfer claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes NeFIC, a paradigm for ultra-low-bitrate image compression that treats generative decoding as a virtual temporal transition: a compact VAE-coded anchor frame (geometry/semantics, low detail) is evolved into the reconstruction by a LoRA-adapted video diffusion model (CogVideoX) via next-frame prediction. A two-stage scheme first adapts the VDM for single-interval, compression-aware detail synthesis (Stage I), then introduces conditional anchor encoding and a Bypass Refinement module so multi-step denoising collapses to one-step generation (Stage II). On Kodak, CLIC2020, and DIV2K the method reports large BD-rate savings versus DiffC and other generative codecs on LPIPS/DISTS/FID/KID (and competitive PSNR/MS-SSIM), with up to about ×5 faster decoding than multi-step diffusion baselines, supported by ablations and qualitative examples.
Significance. If the results hold under independent reimplementation, the work is a meaningful contribution to perceptual ultra-low-bitrate coding: it reframes image decoding as next-frame prediction under pretrained video diffusion priors, uses an explicit visible anchor rather than opaque latent conditioning, and couples that idea with a practical one-step bypass for latency. Strengths include multi-dataset RD/RP curves, BD-rate tables against many recent generative codecs (MS-ILLM, PerCo, DiffEIC, DiffC, ResULIC, StableCodec, DLF, etc.), runtime measurements, component ablations (Stages I/II, cond-enc, bypass, aux loss, CLIP-GAN, LoRA rank), and a planned code release. The dual-domain coupling (pixel-domain anchor + generative bypass latent) is a clear conceptual and engineering advance over pure image-diffusion ULB-IC pipelines.
major comments (3)
- Sec. 3.5, Eqs. (8)–(9): One-step generation models the bypass latent as a partially denoised state at a fixed t★ = 500 (half the maximum schedule). This choice is load-bearing for both the efficiency claim and the claim that generative uncertainty is reduced without multi-step sampling, yet no sensitivity study over t★ (or alternative one-step schedules) is reported. Please ablate t★ (and, if feasible, schedule parameterization) on at least one dataset for LPIPS/DISTS/FID and a qualitative semantic-drift check; without that, robustness of the one-step bypass remains under-supported.
- Sec. 3.3–3.4 and Stage I (Eqs. 2–7): The central paradigm assumes that LoRA adaptation of CogVideoX’s 3D attention/RoPE and defocus-to-focus priors fully suppresses video-native motion, flicker, and scene transition biases for a single-interval image decoder. Table 3 and Fig. 6 show that Stage I and the anchor/bypass components matter, but residual temporal artifacts are only argued qualitatively. A compact quantitative check (e.g., optical-flow magnitude or frame-difference energy between anchor and reconstruction on CLIC/DIV2K, or a controlled two-frame consistency metric) is needed to substantiate that the transfer is complete enough to underwrite the fidelity and “next-frame” claims.
- Table 3 (Stage II row) vs abstract/Tab. 1: Relative to Stage I alone, Stage II improves LPIPS/DISTS BD-rate but worsens FID/KID (paper notes ~4–7% realism cost). The abstract’s “over 50% savings across LPIPS, DISTS, FID, and KID vs DiffC” remains valid as a baseline comparison, but the one-step speedup is not free with respect to generative realism. The main text should more explicitly frame this fidelity–realism–latency trade-off as a limitation of collapsing multi-step VDM sampling, so readers do not over-read the ×5 decode claim as strictly Pareto-improving on all perceptual axes.
minor comments (7)
- Anchor Codec architecture and entropy model are deferred to the supplementary (§4.1); a short main-text sketch (transform type, latent resolution relative to Video-VAE 1/8, entropy model family) would make the rate term R and conditional encoding path self-contained.
- Fig. 3 (“defocus-to-focus”) is illustrative of generic VDM behavior; caption/text should clarify it is not a compression experiment so readers do not treat it as direct evidence of NeFIC reconstructions.
- Notation: E_Vid vs EVid, and ξ vs ε for noise, are slightly inconsistent across Sec. 3.2–3.5; unify symbols for the Video-VAE encoder and Gaussian noise.
- Table 2: DiffC runtime is given as ranges (0.6–9.1 s etc.); state the operating point (steps/bitrate) used for the ×5 comparison so the speedup is reproducible.
- Abstract says “Code will be released at https://github.com/UnoC-727/NeFIC” while the intro says “later”; align the release statement.
- Minor typos/style: “over50%” spacing in the intro abstract block; “V AEs”/space artifacts in several places; “w/o Interval Adaptation” in text vs “w/o Stage I” in Table 3.
- Table 4 FLUX variants: briefly note training budget and whether LoRA/finetune steps match NeFIC so the IDM vs VDM comparison is not confounded by under-adaptation of the larger T2I model.
Circularity Check
No significant circularity: empirical ULB-IC method with external metrics, baselines, and standard training losses; claimed BD-rate savings and speedups are measured outcomes, not definitional.
full rationale
This is a standard empirical computer-vision systems paper. The central claims (abstract; Tab. 1 BD-rates >50% on LPIPS/DISTS/FID/KID vs DiffC; Tab. 2 runtime) are obtained by training an Anchor Codec + LoRA-adapted CogVideoX under ordinary rate-reconstruction-adversarial objectives (Eqs. 5–7, 10–11), then evaluating on held-out Kodak/CLIC2020/DIV2K against published external baselines using established perceptual and distortion metrics. No quantity is defined in terms of the reported savings; λ_R and LoRA rank are ordinary hyperparameters chosen from a discrete set, not fitted parameters later re-labeled as predictions. Citations to CogVideoX, DiffC, StableCodec, etc., are to independent prior work; there is no load-bearing uniqueness theorem or self-citation chain that forces the result. Ablations (Tab. 3–5, Fig. 6) and the FLUX comparison (Tab. 4) further treat components as falsifiable design choices rather than tautologies. The derivation chain therefore does not reduce to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (6)
- λ_R bitrate tradeoff set =
{5, 3, 1.7, 1.0, 0.5, 0.25}
- Stage I aux loss weights λ_MSE, λ_LPIPS, λ_aux =
λ_MSE=5, λ_LPIPS=1, λ_aux=0.1
- Stage II RGB/GAN weights =
1 / 2.5 / 0.5
- One-step diffusion time t★ =
500
- LoRA rank/alpha =
256/256
- Training schedule (steps, LR, crops) =
10k/50k; 1e-4→1e-5; 512–768
axioms (6)
- domain assumption Pretrained VDMs internalize a reusable defocus-to-focus temporal refinement prior that can be repurposed for single-image detail synthesis rather than motion.
- domain assumption Token concatenation with 3D attention/RoPE is a more principled conditioning path than channel-concat IDM control for anchor-faithful reconstruction.
- domain assumption A compact VAE-coded anchor can discard high-frequency detail while retaining enough geometry/semantics for faithful generative decoding at ultra-low rate.
- domain assumption Standard perceptual metrics (LPIPS, DISTS, FID, KID) and BD-rate against DiffC are adequate proxies for “fidelity and realism” of ULB-IC.
- standard math v-prediction diffusion math and LoRA finetuning of DiT attention projections preserve enough generative prior while specializing to compression.
- ad hoc to paper Conditional encoding of the anchor on Video-VAE latents plus a learned bypass map sufficiently aligns compression and generation latent spaces for one-step decoding.
invented entities (3)
-
Compact anchor frame as explicit intermediate decoding state
no independent evidence
-
Next-frame decoding paradigm for ULB-IC (NeFIC)
no independent evidence
-
Bypass Refinement module T (semantic latent bypass)
no independent evidence
read the original abstract
We present a novel paradigm for ultra-low-bitrate image compression (ULB-IC) that exploits the ``temporal'' evolution in generative image compression. Specifically, we define an explicit intermediate state during decoding: a compact anchor frame, which preserves the scene geometry and semantic layout while discarding high-frequency details. We then reinterpret generative decoding as a virtual temporal transition from this anchor to the final reconstructed image. To model this progression, we leverage a pretrained video diffusion model (VDM) as a temporal prior: the anchor frame serves as the initial frame and the original image as the target frame, transforming the decoding process into a next-frame prediction task. In contrast to image diffusion-based ULB-IC models, our decoding proceeds from a visible, semantically faithful anchor, which improves both fidelity and realism for perceptual image compression. Extensive experiments demonstrate that our method achieves superior rate-distortion performance. On the CLIC2020 test set, our method achieves over 50% bitrate savings across LPIPS, DISTS, FID, and KID compared to DiffC, while also delivering a significant decoding speedup of up to $\times$5. Code will be released at https://github.com/UnoC-727/NeFIC.
Reference graph
Works this paper leans on
-
[1]
Simoncelli
Johannes Ballé, Valero Laparra, and Eero P. Simoncelli. End-to-end optimized image com- pression. InInternational Conference on Learning Representations, 2017
2017
-
[2]
Varia- tional image compression with a scale hyperprior
Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Varia- tional image compression with a scale hyperprior. InInternational Conference on Learning Representations, 2018
2018
-
[3]
Joint autoregressive and hierarchical priors for learned image compression.Advances in neural information processing systems, 31, 2018
David Minnen, Johannes Ballé, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression.Advances in neural information processing systems, 31, 2018
2018
-
[4]
Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding
Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, and Yan Wang. Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5718–5727, 2022
2022
-
[5]
Conditional latent coding with learnable synthesized reference for deep image compression
Siqi Wu, Yinda Chen, Dong Liu, and Zhihai He. Conditional latent coding with learnable synthesized reference for deep image compression. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 12863–12871, 2025
2025
-
[6]
The devil is in the details: Window-based attention for image compression
Renjie Zou, Chunfeng Song, and Zhaoxiang Zhang. The devil is in the details: Window-based attention for image compression. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17492–17501, 2022
2022
-
[7]
Transformer-based transform coding
Yinhao Zhu, Yang Yang, and Taco Cohen. Transformer-based transform coding. InInterna- tional Conference on Learning Representations, 2022
2022
-
[8]
Learned image compression with mixed transformer- cnn architectures
Jinming Liu, Heming Sun, and Jiro Katto. Learned image compression with mixed transformer- cnn architectures. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14388–14397, 2023
2023
-
[9]
S2cformer: Reorienting learned image compression from spatial interaction to channel aggregation, 2025
Yunuo Chen, Qian Li, Bing He, Donghui Feng, Ronghua Wu, Qi Wang, Li Song, Guo Lu, and Wenjun Zhang. S2cformer: Reorienting learned image compression from spatial interaction to channel aggregation, 2025
2025
-
[10]
Mambavc: Learned visual compression with selective state spaces
Shiyu Qin, Jinpeng Wang, Yimin Zhou, Bin Chen, Tianci Luo, Baoyi An, Tao Dai, Shutao Xia, and Yaowei Wang. Mambavc: Learned visual compression with selective state spaces. arXiv preprint arXiv:2405.15413, 2024
Pith/arXiv arXiv 2024
-
[11]
Content-aware mamba for learned image compression
Yunuo Chen, Zezheng Lyu, Bing He, Hongwei Hu, Qi Wang, Yuan Tian, Li Song, Wenjun Zhang, and Guo Lu. Content-aware mamba for learned image compression. InThe Fourteenth International Conference on Learning Representations, 2026
2026
-
[12]
Frequency- aware transformer for learned image compression
Han Li, Shaohui Li, Wenrui Dai, Chenglin Li, Junni Zou, and Hongkai Xiong. Frequency- aware transformer for learned image compression. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[13]
Fanhu Zeng, Hao Tang, Yihua Shao, Siyu Chen, Ling Shao, and Yan Wang. Mambaic: State space models for high-performance learned image compression.arXiv preprint arXiv:2503.12461, 2025. 11
Pith/arXiv arXiv 2025
-
[14]
Knowledge distillation for learned image compression
Yunuo Chen, Zezheng Lyu, Bing He, Ning Cao, Gang Chen, Guo Lu, and Wenjun Zhang. Knowledge distillation for learned image compression. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4996–5006, 2025
2025
-
[15]
Scaling learned image compression models up to 1 billion.arXiv preprint arXiv:2508.09075, 2025
Yuqi Li, Haotian Zhang, Li Li, Dong Liu, and Feng Wu. Scaling learned image compression models up to 1 billion.arXiv preprint arXiv:2508.09075, 2025
Pith/arXiv arXiv 2025
-
[16]
Neural rate control for learned video compression
Yiwei Zhang, Guo Lu, Yunuo Chen, Shen Wang, Yibo Shi, Jing Wang, and Li Song. Neural rate control for learned video compression. InThe Twelfth International Conference on Learning Representations, 2023
2023
-
[17]
Yuqi Li, Haotian Zhang, Li Li, and Dong Liu. Learned image compression with hierarchical progressive context modeling.arXiv preprint arXiv:2507.19125, 2025
Pith/arXiv arXiv 2025
-
[18]
Learned image compression with dictionary-based entropy model
Jingbo Lu, Leheng Zhang, Xingyu Zhou, Mu Li, Wen Li, and Shuhang Gu. Learned image compression with dictionary-based entropy model. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 12850–12859, 2025
2025
-
[19]
Differentiable vmaf: A trainable metric for optimizing video compression codec
Jiangchuan Li, Chuqin Zhou, Yunuo Chen, and Guo Lu. Differentiable vmaf: A trainable metric for optimizing video compression codec. In2025 IEEE International Symposium on Circuits and Systems (ISCAS), pages 1–5. IEEE, 2025
2025
-
[20]
Junhao Du, Chuqin Zhou, Ning Cao, Gang Chen, Yunuo Chen, Zhengxue Cheng, Li Song, Guo Lu, and Wenjun Zhang. Large language model for lossless image compression with visual prompts.arXiv preprint arXiv:2502.16163, 2025
Pith/arXiv arXiv 2025
-
[21]
Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020
2020
-
[22]
Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020
Pith/arXiv arXiv 2010
-
[23]
Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
2020
-
[24]
Correcting diffusion-based perceptual image compression with privileged end-to-end decoder.International conference on machine learning (ICML), 2024
Yiyang Ma, Wenhan Yang, and Jiaying Liu. Correcting diffusion-based perceptual image compression with privileged end-to-end decoder.International conference on machine learning (ICML), 2024
2024
-
[25]
Towards image compression with perfect realism at ultra-low bitrates
Marlene Careil, Matthew J Muckley, Jakob Verbeek, and Stéphane Lathuilière. Towards image compression with perfect realism at ultra-low bitrates. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[26]
Towards extreme image compression with latent feature guidance and diffusion prior.IEEE Transactions on Circuits and Systems for Video Technology, 2024
Zhiyuan Li, Yanhui Zhou, Hao Wei, Chenyang Ge, and Jingwen Jiang. Towards extreme image compression with latent feature guidance and diffusion prior.IEEE Transactions on Circuits and Systems for Video Technology, 2024
2024
-
[27]
Faithdiff: Unleashing diffusion priors for faithful image super-resolution
Junyang Chen, Jinshan Pan, and Jiangxin Dong. Faithdiff: Unleashing diffusion priors for faithful image super-resolution. InCVPR, 2025
2025
-
[28]
Anle Ke, Xu Zhang, Tong Chen, Ming Lu, Chao Zhou, Jiawen Gu, and Zhan Ma. Ultra lowrate image compression with semantic residual coding and compression-aware diffusion.arXiv preprint arXiv:2505.08281, 2025
Pith/arXiv arXiv 2025
-
[29]
Dlf: Extreme image compression with dual-generative latent fusion.Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct 2025
Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Yuan Zhang, and Yan Lu. Dlf: Extreme image compression with dual-generative latent fusion.Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct 2025
2025
-
[30]
Learned image compression with discretized gaussian mixture likelihoods and attention modules
Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7939–7948, 2020. 12
2020
-
[31]
Mlic: Multi-reference entropy model for learned image compression
Wei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning, Feng Gao, and Ronggang Wang. Mlic: Multi-reference entropy model for learned image compression. InProceedings of the 31st ACM International Conference on Multimedia, pages 7618–7627, 2023
2023
-
[32]
MLIC++: Linear complexity multi-reference entropy mod- eling for learned image compression
Wei Jiang and Ronggang Wang. MLIC++: Linear complexity multi-reference entropy mod- eling for learned image compression. InICML 2023 Workshop Neural Compression: From Information Theory to Applications, 2023
2023
-
[33]
Channel-wise autoregressive entropy models for learned image compression
David Minnen and Saurabh Singh. Channel-wise autoregressive entropy models for learned image compression. In2020 IEEE International Conference on Image Processing (ICIP), pages 3339–3343. IEEE, 2020
2020
-
[34]
Glic: General format learned image compression
MingSheng Zhou and MingMing Kong. Glic: General format learned image compression. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 10815–10824, 2025
2025
-
[35]
Fidelity-controllable extreme image compression with generative adversarial networks
Shoma Iwai, Tomo Miyazaki, Yoshihiro Sugaya, and Shinichiro Omachi. Fidelity-controllable extreme image compression with generative adversarial networks. In2020 25th International Conference on Pattern Recognition (ICPR), pages 8235–8242. IEEE, 2021
2021
-
[36]
Guo-Hua Wang, Jiahao Li, Bin Li, and Yan Lu. Evc: Towards real-time neural image compression with mask decay.arXiv preprint arXiv:2302.05071, 2023
Pith/arXiv arXiv 2023
-
[37]
Generative adversarial networks for extreme learned image compres- sion
Eirikur Agustsson et al. Generative adversarial networks for extreme learned image compres- sion. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019
2019
-
[38]
High- fidelity generative image compression.Advances in neural information processing systems, 33:11913–11924, 2020
Fabian Mentzer, George D Toderici, Michael Tschannen, and Eirikur Agustsson. High- fidelity generative image compression.Advances in neural information processing systems, 33:11913–11924, 2020
2020
-
[39]
Multi-realism image compression with a conditional generator
Eirikur Agustsson, David Minnen, George Toderici, and Fabian Mentzer. Multi-realism image compression with a conditional generator. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22324–22333, 2023
2023
-
[40]
Improving statistical fidelity for neural image compression with implicit local likelihood models
Matthew J Muckley, Alaaeldin El-Nouby, Karen Ullrich, Hervé Jégou, and Jakob Verbeek. Improving statistical fidelity for neural image compression with implicit local likelihood models. InInternational Conference on Machine Learning, pages 25426–25443. PMLR, 2023
2023
-
[41]
Egic: enhanced low-bit-rate generative image compression guided by semantic segmentation
Nikolai Körber, Eduard Kromer, Andreas Siebert, Sascha Hauke, Daniel Mueller-Gritschneder, and Björn Schuller. Egic: enhanced low-bit-rate generative image compression guided by semantic segmentation. InEuropean Conference on Computer Vision, pages 202–220. Springer, 2024
2024
-
[42]
Extreme image compression using fine-tuned vqgans
Qi Mao, Tinghan Yang, Yinuo Zhang, Zijian Wang, Meng Wang, Shiqi Wang, Libiao Jin, and Siwei Ma. Extreme image compression using fine-tuned vqgans. In2024 Data Compression Conference (DCC), pages 203–212. IEEE, 2024
2024
-
[43]
Unifying generation and compression: Ultra-low bitrate image coding via multi-stage transformer
Naifu Xue, Qi Mao, Zijian Wang, Yuan Zhang, and Siwei Ma. Unifying generation and compression: Ultra-low bitrate image coding via multi-stage transformer. In2024 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6. IEEE, 2024
2024
-
[44]
Generative latent coding for ultra- low bitrate image compression
Zhaoyang Jia, Jiahao Li, Bin Li, Houqiang Li, and Yan Lu. Generative latent coding for ultra- low bitrate image compression. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26088–26098, 2024
2024
-
[45]
Hagyeong Lee, Minkyu Kim, Jun-Hyuk Kim, Seungeon Kim, Dokwan Oh, and Jaeho Lee. Neural image compression with text-guided encoding for both pixel-level and perceptual fidelity.arXiv preprint arXiv:2403.02944, 2024
Pith/arXiv arXiv 2024
-
[46]
Text+ sketch: Image compression at ultra low rates.arXiv preprint arXiv:2307.01944, 2023
Eric Lei, Yi˘git Berkay Uslu, Hamed Hassani, and Shirin Saeedi Bidokhti. Text+ sketch: Image compression at ultra low rates.arXiv preprint arXiv:2307.01944, 2023. 13
Pith/arXiv arXiv 2023
-
[47]
Perco (sd): Open perceptual compression.arXiv preprint arXiv:2409.20255, 2024
Nikolai Körber, Eduard Kromer, Andreas Siebert, Sascha Hauke, Daniel Mueller-Gritschneder, and Björn Schuller. Perco (sd): Open perceptual compression.arXiv preprint arXiv:2409.20255, 2024
Pith/arXiv arXiv 2024
-
[48]
Lossy compression with gaussian diffusion.arXiv preprint arXiv:2206.08889, 2022
Lucas Theis, Tim Salimans, Matthew D Hoffman, and Fabian Mentzer. Lossy compression with gaussian diffusion.arXiv preprint arXiv:2206.08889, 2022
Pith/arXiv arXiv 2022
-
[49]
Image compression with product quantized masked image modeling.arXiv preprint arXiv:2212.07372, 2022
Alaaeldin El-Nouby, Matthew J Muckley, Karen Ullrich, Ivan Laptev, Jakob Verbeek, and Hervé Jégou. Image compression with product quantized masked image modeling.arXiv preprint arXiv:2212.07372, 2022
Pith/arXiv arXiv 2022
-
[50]
Perceptual image compression via stable diffusion at low bitrate
Luoxu Jin and Hiroshi Watanabe. Perceptual image compression via stable diffusion at low bitrate. In2024 IEEE 7th International Conference on Multimedia Information Processing and Retrieval (MIPR), pages 626–629. IEEE, 2024
2024
-
[51]
Xiaoyue Ling, Chuqin Zhou, Chunyi Li, Yunuo Chen, Yuan Tian, Guo Lu, and Wenjun Zhang. Free-gvc: Towards training-free extreme generative video compression with temporal coherence.arXiv preprint arXiv:2602.09868, 2026
arXiv 2026
-
[52]
Idempotence and perceptual image compression
Tongda Xu, Ziran Zhu, Dailan He, Yanghao Li, Lina Guo, Yuanyuan Wang, Zhe Wang, Hongwei Qin, Yan Wang, Jingjing Liu, and Ya-Qin Zhang. Idempotence and perceptual image compression. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[53]
Lossy image compression with conditional diffusion models
Ruihan Yang and Stephan Mandt. Lossy image compression with conditional diffusion models. Advances in Neural Information Processing Systems, 36:64971–64995, 2023
2023
-
[54]
Lossy image compression with foundation diffusion models
Lucas Relic, Roberto Azevedo, Markus Gross, and Christopher Schroers. Lossy image compression with foundation diffusion models. InEuropean Conference on Computer Vision, pages 303–319. Springer, 2024
2024
-
[55]
One-step diffusion-based image compression with semantic distillation.Advances in Neural Information Processing Systems, 2025
Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Yuan Zhang, and Yan Lu. One-step diffusion-based image compression with semantic distillation.Advances in Neural Information Processing Systems, 2025
2025
-
[56]
Juan Song, Lijie Yang, and Mingtao Feng. Extremely low-bitrate image compression se- mantically disentangled by lmms from a human perception perspective.arXiv preprint arXiv:2503.00399, 2025
arXiv 2025
-
[57]
Chuqin Zhou, Xiaoyue Ling, Yunuo Chen, Jincheng Dai, Guo Lu, and Wenjun Zhang. Dual- representation image compression at ultra-low bitrates via explicit semantics and implicit textures.arXiv preprint arXiv:2602.05213, 2026
arXiv 2026
-
[58]
Diffusion-based extreme image compression with compressed feature initialization
Zhiyuan Li, Yanhui Zhou, Hao Wei, Chenyang Ge, and Ajmal Saeed Mian. Diffusion-based extreme image compression with compressed feature initialization. 2024
2024
-
[59]
Generative image coding with diffusion prior.arXiv preprint arXiv:2509.13768, 2025
Jianhui Chang. Generative image coding with diffusion prior.arXiv preprint arXiv:2509.13768, 2025
arXiv 2025
-
[60]
Hybridflow: Infusing continuity into masked codebook for extreme low-bitrate image compression
Lei Lu, Yanyue Xie, Wei Jiang, Wei Wang, Xue Lin, and Yanzhi Wang. Hybridflow: Infusing continuity into masked codebook for extreme low-bitrate image compression. InProceedings of the 32nd ACM International Conference on Multimedia, pages 3010–3018, 2024
2024
-
[61]
Picd: Versatile perceptual image compression with diffusion rendering
Tongda Xu, Jiahao Li, Bin Li, Yan Wang, Ya-Qin Zhang, and Yan Lu. Picd: Versatile perceptual image compression with diffusion rendering. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 28436–28445, 2025
2025
-
[62]
Oscar: One-step diffusion codec across multiple bit-rates.arXiv preprint arXiv:2505.16091, 2025
Jinpei Guo, Yifei Ji, Zheng Chen, Kai Liu, Min Liu, Wang Rao, Wenbo Li, Yong Guo, and Yulun Zhang. Oscar: One-step diffusion codec across multiple bit-rates.arXiv preprint arXiv:2505.16091, 2025
arXiv 2025
-
[63]
Diffpc: Diffusion-based high perceptual fidelity image compression with semantic refinement
Yichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An, Haoqian Wang, Yaowei Wang, and Bin Chen. Diffpc: Diffusion-based high perceptual fidelity image compression with semantic refinement. InThe Thirteenth International Conference on Learning Representations, 2025. 14
2025
-
[64]
Vcip 2025 ul- tra low-bitrate video compression challenge
Guo Lu, Jing Wang, Yunuo Chen, Chuqin Zhou, Yibo Shi, and Kai Lin. Vcip 2025 ul- tra low-bitrate video compression challenge. In2025 International Conference on Visual Communications and Image Processing (VCIP), pages 1–3. IEEE, 2025
2025
-
[65]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. InProceedings of the IEEE/CVF international conference on computer vision, pages 3836–3847, 2023
2023
-
[66]
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. InProceedings of the AAAI conference on artificial intelligence, volume 38, pages 4296–4304, 2024
2024
-
[67]
Controlnet-xs: Rethinking the control of text-to-image diffusion models as feedback-control systems
Denis Zavadski, Johann-Friedrich Feiden, and Carsten Rother. Controlnet-xs: Rethinking the control of text-to-image diffusion models as feedback-control systems. InEuropean Conference on Computer Vision, pages 343–362. Springer, 2024
2024
-
[68]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022
2022
-
[69]
Lossy compression with pretrained diffusion models
Jeremy V onderfecht and Feng Liu. Lossy compression with pretrained diffusion models. International Conference on Learning Representations, 2025
2025
-
[70]
Noor Fathima Ghouse, Jens Petersen, Auke Wiggers, Tianlin Xu, and Guillaume Sautiere. A residual diffusion model for high perceptual quality codec augmentation.arXiv preprint arXiv:2301.05489, 2023
Pith/arXiv arXiv 2023
-
[71]
Neural image compression with a diffusion-based decoder
Noor Fathima Khanum Mohamed Ghouse, Jens Petersen, Auke J Wiggers, Tianlin Xu, and Guillaume Sautiere. Neural image compression with a diffusion-based decoder. 2023
2023
-
[72]
Emiel Hoogeboom, Eirikur Agustsson, Fabian Mentzer, Luca Versari, George Toderici, and Lucas Theis. High-fidelity image compression with score-based generative models.arXiv preprint arXiv:2305.18231, 2023
Pith/arXiv arXiv 2023
-
[73]
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022
Pith/arXiv arXiv 2022
-
[74]
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description.arXiv preprint arXiv:2210.02399, 2022
Pith/arXiv arXiv 2022
-
[75]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023
2023
-
[76]
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.arXiv preprint arXiv:2307.04725, 2023
Pith/arXiv arXiv 2023
-
[77]
Video diffusion models.Advances in neural information processing systems, 35:8633–8646, 2022
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models.Advances in neural information processing systems, 35:8633–8646, 2022
2022
-
[78]
Efficient spatially sparse inference for conditional gans and diffusion models.Advances in neural information processing systems, 35:28858–28873, 2022
Muyang Li, Ji Lin, Chenlin Meng, Stefano Ermon, Song Han, and Jun-Yan Zhu. Efficient spatially sparse inference for conditional gans and diffusion models.Advances in neural information processing systems, 35:28858–28873, 2022
2022
-
[79]
Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792, 2022
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792, 2022. 15
Pith/arXiv arXiv 2022
-
[80]
Lingen: Towards high-resolution minute-length text-to-video generation with linear computational complexity
Hongjie Wang, Chih-Yao Ma, Yen-Cheng Liu, Ji Hou, Tao Xu, Jialiang Wang, Felix Juefei-Xu, Yaqiao Luo, Peizhao Zhang, Tingbo Hou, et al. Lingen: Towards high-resolution minute-length text-to-video generation with linear computational complexity. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 2578–2588, 2025
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.