REVIEW 2 major objections 5 minor 90 references
A lightweight predictor ranks initial noises from the prompt alone so diffusion models generate preferred images without re-rolling the dice.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 22:22 UTC pith:4WHNS7EG
load-bearing objection Clean engineering win: regress preference from noise+prompt, rank top-K, and get free prompt-difficulty feedback without touching the DM. the 2 major comments →
Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors show that a scalar human-preference score for a would-be image can be predicted well enough from only the prompt encoding and the initial noise tensor that ranking many candidate noises and forwarding the top-K systematically raises HPSv2/v3, ImageReward, and PickScore relative to unoptimized sampling and to Golden Noise, without fine-tuning the diffusion model itself.
What carries the argument
Naïve PAINE (Prompt-Aware Initial Noise Evaluator): a three-module predictor (prompt encoder with a learnable summary token, ResNet noise encoder, MLP score head) that estimates preference from (prompt embedding, initial noise). Masking the noise branch yields a Naïve-Bayes-style prior on the prompt-conditioned mean score.
Load-bearing premise
A predictor trained offline on preference scores of fully generated images will rank never-before-seen noises accurately enough that the top-K choices land in the upper tail of the true preference distribution for new prompts.
What would settle it
Hold out a fresh prompt set and a new diffusion model; generate many images from random noises, score them with the same preference metric used in training, then check whether the noises PAINE ranks highest actually produce higher measured scores than random or Golden-Noise baselines at the same compute budget.
If this is right
- Users can reduce the number of full generation cycles needed to obtain a satisfactory image for a given prompt.
- Existing Diffusers or ComfyUI pipelines can insert the predictor as a pre-generation filter with only milliseconds of added latency.
- Prompt-only mode supplies an interpretable difficulty signal so a user can rewrite a prompt before spending GPU time.
- The same selection idea applies without architecture changes to both older U-Net models and newer Diffusion Transformers.
Where Pith is reading between the lines
- If the ranking signal generalizes across preference metrics, practitioners could train once on the cheapest metric and still improve others, lowering annotation cost.
- The same noise-ranking idea should transfer to text-to-video or autoregressive image models whose quality also depends on a stochastic seed.
- When the predictor’s mean-score feedback is low, interactive prompt rewriting guided by that score may raise final quality more than noise selection alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Naïve PAINE, a lightweight plug-and-play predictor that estimates human-preference scores (primarily PickScore) of a would-be T2I image from the prompt embedding and the initial Gaussian noise latent, before reverse diffusion. Candidate noises are ranked and the top-N_img are forwarded for full generation; zeroing the noise encoder yields a prompt-only mean-score estimate framed as a naïve-Bayes prior. The method is trained offline on 5k Pick-a-Pic prompts × 20 noises per model (Hunyuan-DiT, PixArt-α, DreamShaper-XL, SDXL) with MAE + differentiable SRCC loss, and is evaluated against Golden Noise, NoiseAR and HyperNoise on HPSv2/v3, ImageReward, PickScore and GenEval, with latency and ablation studies.
Significance. If the ranking premise holds, the work supplies a practical, model-agnostic, fine-tuning-free alternative to noise mutation or RL-based optimizers that is easy to drop into Diffusers/ComfyUI pipelines and that additionally returns an interpretable prompt-difficulty signal. Multi-model, multi-benchmark tables (Tables 1–2, 7–9), hardware measurements (Tables 3–4), and ablations on target metric, loss, K and text-encoder design constitute a solid empirical package; code is released. The contribution is incremental but useful for the large community that still runs multi-sample generation on open DMs.
major comments (2)
- §3.1 Eq. (4) and §4.1 training protocol rest on the claim that offline PickScore (or similar) labels of fully generated images produce rankings accurate enough for top-K noises to systematically occupy the upper tail of the true preference distribution on new prompts. Validation SRCC of 0.74–0.87 (Table 6) is only moderate; human-preference labels are themselves noisy (authors’ own citations). The manuscript never reports an oracle gap (true top-K by post-generation score vs. PAINE-selected top-K) or error bars / significance tests on the metric deltas in Table 1. Without that measurement it is hard to know how much of the reported gains survive label noise and distribution shift.
- §4.2 / Table 1: all preference-metric gains are point estimates. Given the known variance of HPSv2/v3, ImageReward and PickScore across seeds and the multiple-comparison setting (4 models × 4 prompt sets × 4 metrics), the absence of standard errors or a simple paired test leaves open the possibility that several “best” entries are not reliably better than Golden Noise or the standard baseline. Adding these would make the central claim falsifiable rather than merely directionally consistent.
minor comments (5)
- §2.3 / Figs. 1–3: the preliminary distribution study is informative but uses only 50 prompts and 20 seeds; a short note on how sensitive the PCC matrices are to prompt sampling would strengthen the motivation.
- §3.2: the “naïve Bayes” framing is heuristic (zeroing the noise encoder is not a true likelihood). Soften the language or move the analogy to discussion so it is not read as a formal derivation.
- Table 3 vs. Table 4: parameter counts favor Golden Noise while latency favors PAINE; a single combined table of wall-clock cost for N_img = 1 and N_img = 4 would make the efficiency claim clearer.
- Supplementary Table 6 reports validation SRCC/MAE; moving a one-line summary into the main text (near §4.1) would help readers assess predictor quality without leaving the paper.
- Scattered OCR/encoding artifacts (e.g., “Na\"ive”, “PixArt-�”, “�����”) should be cleaned for the camera-ready version.
Circularity Check
No circular derivation; standard supervised ranking of noises via an offline-trained regressor on external preference labels, validated on held-out generations.
full rationale
The core pipeline (Eq. 4, Sec. 3.1) trains a lightweight network to regress external human-preference scores (primarily PickScore, also HPSv2/ImageReward) of fully-generated images from the corresponding prompt embeddings and initial noise tensors. At inference the network ranks a batch of fresh noises and forwards the top-K; the prompt-only mean estimate is obtained simply by zeroing the noise encoder. Neither step reduces the claimed quality gain to a fitted constant or to a definitional identity: the training targets are produced by independent published preference models on actual RDP outputs, the validation SRCC/MAE figures (Table 6, Sec. 4.1) are imperfect, and all reported benchmark numbers (Tables 1–2, 7–9) are measured after full generation on held-out prompt corpora. The Naïve-Bayes framing (Sec. 3.2) is an architectural interpretation, not a load-bearing uniqueness claim. No self-citation supplies a uniqueness theorem, no ansatz is smuggled via prior author work, and no equation equates a “prediction” to its own training input by construction. The method is therefore an ordinary empirical ML contribution evaluated against external baselines and metrics.
Axiom & Free-Parameter Ledger
free parameters (4)
- candidate noise count K
- learning rate / AdamW schedule
- loss mixing weight (MAE + SRCC)
- number of training prompts / images per prompt
axioms (4)
- domain assumption Human preference for a generated image given a prompt can be adequately summarized by a scalar score from an existing predictor (PickScore, HPSv2/v3, ImageReward).
- domain assumption The reverse diffusion process is a deterministic function of the initial noise and the prompt embedding once the sampler and seed are fixed.
- standard math Prompt-conditioned score distributions have finite mean and variance that can be estimated from 20 samples.
- ad hoc to paper Zeroing the noise-encoder output yields an estimate of the prompt-only mean score (Naïve Bayes framing).
invented entities (1)
-
PAINE predictor (prompt encoder + noise ResNet + score MLP)
no independent evidence
read the original abstract
Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slots at a casino, a DM will produce different results given the same user-defined inputs. This imposes a gambler's burden: To perform multiple generation cycles to obtain a satisfactory result. However, even though DMs use stochastic sampling to seed generation, the distribution of generated content quality highly depends on the prompt and the generative ability of a DM with respect to it. To account for this, we propose Na\"ive PAINE for improving the generative quality of Diffusion Models by leveraging T2I preference benchmarks. We directly predict the numerical quality of an image from the initial noise and given prompt. Na\"ive PAINE then selects a handful of quality noises and forwards them to the DM for generation. Further, Na\"ive PAINE provides feedback on the DM generative quality given the prompt and is lightweight enough to seamlessly fit into existing DM pipelines. Experimental results demonstrate that Na\"ive PAINE outperforms existing approaches on several prompt corpus benchmarks.
Reference graph
Works this paper leans on
-
[1]
AUTOMATIC1111: Stable diffusion webui.https://github.com/AUTOMATIC1111/ stable-diffusion-webui(2022)
2022
-
[2]
arXiv preprint arXiv:2412.10891 (2024)
Bai, L., Shao, S., Zhou, Z., Qi, Z., Xu, Z., Xiong, H., Xie, Z.: Zigzag diffu- sion sampling: Diffusion models can self-improve via self-reflection. arXiv preprint arXiv:2412.10891 (2024)
Pith/arXiv arXiv 2024
-
[3]
Bishop, C.M., Nasrabadi, N.M.: Pattern recognition and machine learning, vol. 4. Springer (2006)
2006
-
[4]
Black Forest Labs: Flux.https://github.com/black-forest-labs/flux(2024)
2024
-
[5]
In: International Conference on Machine Learning
Blondel, M., Teboul, O., Berthet, Q., Djolonga, J.: Fast differentiable sorting and ranking. In: International Conference on Machine Learning. pp. 950–959. PMLR (2020)
2020
-
[6]
arXiv preprint arXiv:2511.22699 (2025)
Cai, H., Cao, S., Du, R., Gao, P., Hoi, S., Hou, Z., Huang, S., Jiang, D., Jin, X., Li, L., et al.: Z-image: An efficient image generation foundation model with single-stream diffusion transformer. arXiv preprint arXiv:2511.22699 (2025)
Pith/arXiv arXiv 2025
-
[7]
canalys: Now and next for ai-capablee smartphones. Tech. rep., A Canalys Special Report (05 2024)
2024
-
[8]
ACM trans- actions on Graphics (TOG)42(4), 1–10 (2023)
Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., Cohen-Or, D.: Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM trans- actions on Graphics (TOG)42(4), 1–10 (2023)
2023
-
[9]
In: Proceedings of the 32nd ACM International Conference on Multimedia
Chen, C., Yang, L., Yang, X., Chen, L., He, G., Wang, C., Li, Y.: Find: Fine- tuning initial noise distribution with policy optimization for diffusion models. In: Proceedings of the 32nd ACM International Conference on Multimedia. pp. 6735– 6744 (2024)
2024
-
[10]
In: European Conference on Computer Vision
Chen, J., Ge, C., Xie, E., Wu, Y., Yao, L., Ren, X., Wang, Z., Luo, P., Lu, H., Li, Z.: PixArt-� : Weak-to-strong training of diffusion transformer for 4k text- to-image generation. In: European Conference on Computer Vision. pp. 74–91. Springer (2025)
2025
-
[11]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Chen, J., Xue, S., Zhao, Y., Yu, J., Paul, S., Chen, J., Cai, H., Han, S., Xie, E.: Sana-sprint: One-step diffusion with continuous-time consistency distillation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 16185–16195 (2025)
2025
-
[12]
In: The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wang, Z., Kwok, J.T., Luo, P., Lu, H., Li, Z.: Pixart-�: Fast training of diffusion transformer for photorealistic text-to-image synthesis. In: The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net (2024),https: //openreview.net/forum?id=eAKm...
2024
-
[13]
urlhttps://civitai.com (2026), accessed on January 27, 2026
Civitai: Civitai: The Generative AI Community Hub. urlhttps://civitai.com (2026), accessed on January 27, 2026
2026
-
[14]
Ai Magazine37(4), 63–66 (2016)
Cohen, P.: Harold cohen and aaron. Ai Magazine37(4), 63–66 (2016)
2016
-
[15]
Contributors, C.: Comfyui: A node-based gui for stable diffusion (2025),https: //github.com/Comfy-Org/ComfyUI
2025
-
[16]
Cyberdelia: CyberRealistic Pony - | Stable Diffusion XL Checkpoint | Civitai — civitai.com.https://civitai.com/models/443821/cyberrealistic-pony(2026), [Accessed 02-03-2026]
2026
-
[17]
Addictive Behaviors137, 107540 (2023)
Delfabbro, P., King, D., Parke, J.: The complex nature of human operant gam- bling behaviour involving slot games: Structural characteristics, verbal rules and motivation. Addictive Behaviors137, 107540 (2023)
2023
-
[18]
OpenReview.net (2024),https://openreview.net/forum?id= FPnUhsQJ5B
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Rombach, R.:Scalingrectifiedflowtransformersforhigh-resolutionimagesynthesis.In:Forty- first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net...
2024
-
[19]
Eyring, L., Karthik, S., Dosovitskiy, A., Ruiz, N., Akata, Z.: Noise hypernetworks: Amortizingtest-timecomputeindiffusionmodels.arXivpreprintarXiv:2508.09968 (2025)
Pith/arXiv arXiv 2025
-
[20]
Advances in Neural Information Processing Systems37, 125487–125519 (2024)
Eyring, L., Karthik, S., Roth, K., Dosovitskiy, A., Akata, Z.: Reno: Enhancing one-step text-to-image models through reward-based noise optimization. Advances in Neural Information Processing Systems37, 125487–125519 (2024)
2024
-
[21]
International Gambling Studies22(2), 317–336 (2022)
Ferrari, M.A., Limbrick-Oldfield, E.H., Clark, L.: Behavioral analysis of habit for- mation in modern slot machine gambling. International Gambling Studies22(2), 317–336 (2022)
2022
-
[22]
Advances in Neural Information Processing Systems36, 52132–52152 (2023)
Ghosh, D., Hajishirzi, H., Schmidt, L.: Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems36, 52132–52152 (2023)
2023
-
[23]
In: Advances in neural information processing systems
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: Advances in neural information processing systems. pp. 2672–2680 (2014)
2014
-
[24]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Guo, X., Liu, J., Cui, M., Li, J., Yang, H., Huang, D.: Initno: Boosting text- to-image diffusion models via initial noise optimization. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9380– 9389 (2024)
2024
-
[25]
He,K.,Zhang,X.,Ren,S.,Sun,J.:Deepresiduallearningforimagerecognition.In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
2016
-
[26]
Advances in Neural Information Processing Sys- tems36(2024)
He,Y.,Liu,L.,Liu,J.,Wu,W.,Zhou,H.,Zhuang,B.:Ptqd:Accuratepost-training quantization for diffusion models. Advances in Neural Information Processing Sys- tems36(2024)
2024
-
[27]
In: Moens, M.F., Huang, X., Specia, L., Yih, S.W.t
Hessel,J.,Holtzman,A.,Forbes,M.,LeBras,R.,Choi,Y.:CLIPScore:Areference- free evaluation metric for image captioning. In: Moens, M.F., Huang, X., Specia, L., Yih, S.W.t. (eds.) Proceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing. pp. 7514–7528. Association for Compu- tational Linguistics, Online and Punta Cana, Dominica...
-
[28]
Advances in neural information processing systems33, 6840–6851 (2020) Naïve PAINE 17
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Advances in neural information processing systems33, 6840–6851 (2020) Naïve PAINE 17
2020
-
[29]
Jaramillo, D.: Radiologists and their noise: variability in human judgment, fallibil- ity, and strategies to improve accuracy (2022)
2022
-
[30]
Jiang, L., Chen, R., Gao, C., Niu, D.: Raise: Requirement-adaptive evolutionary refinement for training-free text-to-image alignment (2026),https://arxiv.org/ abs/2603.00483
arXiv 2026
-
[31]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Jiang, L., Hassanpour, N., Salameh, M., Samadi, M., He, J., Sun, F., Niu, D.: Pixelman: Consistent object editing with diffusion models via pixel manipulation and generation. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 4012–4020 (2025)
2025
-
[32]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2025)
Jiang, L., Hassanpour, N., Salameh, M., Samadi, M., He, J., Sun, F., Niu, D.: Pixelman: Consistent object editing with diffusion models via pixel manipulation and generation. In: Proceedings of the AAAI Conference on Artificial Intelligence (2025)
2025
-
[33]
arXiv preprint arXiv:2408.11706 (2024)
Jiang, L., Hassanpour, N., Salameh, M., Singamsetti, M.S., Sun, F., Lu, W., Niu, D.: Frap: Faithful and realistic text-to-image generation with adaptive prompt weighting. arXiv preprint arXiv:2408.11706 (2024)
Pith/arXiv arXiv 2024
-
[34]
arXiv preprint arXiv:2510.05849 (2025)
Kalaivanan, A., Zhao, Z., Sjölund, J., Lindsten, F.: Ess-flow: Training-free guidance of flow-based models as inference in source space. arXiv preprint arXiv:2510.05849 (2025)
arXiv 2025
-
[35]
In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems
Kamali, N., Nakamura, K., Kumar, A., Chatzimparmpas, A., Hullman, J., Groh, M.: Characterizing photorealism and artifacts in diffusion model-generated images. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. pp. 1–26 (2025)
2025
-
[36]
com / models / 133005 / juggernaut-xl(2025), [Accessed 27-01-2026]
KandooAI: Juggernaut XL - Ragnarok_by_RunDiffusion | Stable Diffusion XL Checkpoint | Civitai — civitai.com.https : / / civitai . com / models / 133005 / juggernaut-xl(2025), [Accessed 27-01-2026]
2025
-
[37]
Kingma, D.P., Welling, M.: An introduction to variational autoencoders. Found. Trends Mach. Learn.12(4), 307–392 (2019).https : / / doi . org / 10 . 1561 / 2200000056,https://doi.org/10.1561/2200000056
-
[38]
Advances in neural information processing systems36, 36652–36663 (2023)
Kirstain,Y.,Polyak,A.,Singer,U.,Matiana,S.,Penna,J.,Levy,O.:Pick-a-pic:An open dataset of user preferences for text-to-image generation. Advances in neural information processing systems36, 36652–36663 (2023)
2023
-
[39]
WIRED (Aug 2017),https://www.realclearinvestigations.com/links/ 2017/08/09/meet_a_russian_casino_hacker_104376.html, feature on a St
Koerner, B.I.: Russians engineer a brilliant slot machine cheat—and casinos have no fix. WIRED (Aug 2017),https://www.realclearinvestigations.com/links/ 2017/08/09/meet_a_russian_casino_hacker_104376.html, feature on a St. Pe- tersburg group predicting outcomes by modeling slot PRNG temporal state
2017
-
[40]
In: Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision
Li, X., Liu, Y., Lian, L., Yang, H., Dong, Z., Kang, D., Zhang, S., Keutzer, K.: Q-diffusion: Quantizing diffusion models. In: Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision. pp. 17535–17545 (2023)
2023
-
[41]
Advances in neural information processing systems37, 90198–90225 (2024)
Li, Y., Jiang, H., Kodaira, A., Tomizuka, M., Keutzer, K., Xu, C.: Immiscible diffusion: Accelerating diffusion training with noise assignment. Advances in neural information processing systems37, 90198–90225 (2024)
2024
-
[42]
In: 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021
Li, Y., Gong, R., Tan, X., Yang, Y., Hu, P., Zhang, Q., Yu, F., Wang, W., Gu, S.: BRECQ: pushing the limit of post-training quantization by block reconstruction. In: 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net (2021),https://openreview. net/forum?id=POWv6hDd9XH
2021
-
[43]
arXiv preprint arXiv:2506.01337 (2025) 18 J
Li, Z., Liu, X., Zhang, X., Tan, P., Shum, H.Y.: Noisear: Autoregressing initial noise prior for diffusion models. arXiv preprint arXiv:2506.01337 (2025) 18 J. Kim et al
Pith/arXiv arXiv 2025
-
[44]
Li, Z., Zhang, J., Lin, Q., Xiong, J., Long, Y., Deng, X., Zhang, Y., Liu, X., Huang, M., Xiao, Z., Chen, D., He, J., Li, J., Li, W., Zhang, C., Quan, R., Lu, J., Huang, J., Yuan, X., Zheng, X., Li, Y., Zhang, J., Zhang, C., Chen, M., Liu, J., Fang, Z., Wang, W., Xue, J., Tao, Y., Zhu, J., Liu, K., Lin, S., Sun, Y., Li, Y., Wang, D., Chen, M., Hu, Z., Xia...
2024
-
[45]
In: 7th Inter- national Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: 7th Inter- national Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net (2019),https://openreview.net/forum? id=Bkg6RiCqY7
2019
-
[46]
In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) (2023)
Lu, S., Hu, Y., Wang, P., Han, Y., Tan, J., Li, J., Yang, S., Liu, J.: Pinat: A permutation invariance augmented transformer for nas predictor. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) (2023)
2023
-
[47]
Lykon: Dreamshaper - stable diffusion fine-tune.https://civitai.com/models/ 4384/dreamshaper(2023)
2023
-
[48]
arXiv preprint arXiv:2503.00811 (2025)
Ma, L., Cao, K., Liang, H., Lin, J., Li, Z., Liu, Y., Zhang, J., Zhang, W., Cui, B.: Evaluating and predicting distorted human body parts for generated images. arXiv preprint arXiv:2503.00811 (2025)
Pith/arXiv arXiv 2025
-
[49]
arXiv preprint arXiv:2402.17764 (2024)
Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J., Wei, F.: The era of 1-bit llms: All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764 (2024)
Pith/arXiv arXiv 2024
-
[50]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Ma, Y., Wu, X., Sun, K., Li, H.: Hpsv3: Towards wide-spectrum human preference score. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 15086–15095 (2025)
2025
-
[51]
Mance, J.: Noise: A flaw in human judgment (2022)
2022
-
[52]
In: CVPR (2024)
Mills, K.G., Han, F.X., Salameh, M., Lu, S., Zhou, C., He, J., Sun, F., Niu, D.: Building optimal neural architectures using interpretable knowledge. In: CVPR (2024)
2024
-
[53]
Mills, K.G., Niu, D., Salameh, M., Qiu, W., Han, F.X., Liu, P., Zhang, J., Lu, W., Jui, S.: Aio-p: Expanding neural performance predictors beyond image clas- sification. Proceedings of the AAAI Conference on Artificial Intelligence37(8), 9180–9189 (06 2023).https://doi.org/10.1609/aaai.v37i8.26101,https: //ojs.aaai.org/index.php/AAAI/article/view/26101
-
[54]
In: AAAI (2025)
Mills, K.G., Salameh, M., Chen, R., Hassanpour, Negar Lu, W., Niu, D.: Qua2sedimo: Quantifiable quantization sensitivity of diffusion models. In: AAAI (2025)
2025
-
[55]
In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Muennighoff, N., Yang, Z., Shi, W., Li, X.L., Fei-Fei, L., Hajishirzi, H., Zettle- moyer, L., Liang, P., Candès, E., Hashimoto, T.B.: s1: Simple test-time scaling. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. pp. 20286–20332 (2025)
2025
-
[56]
Page, Bulletin of the Computer Arts Society (18), 1–2 (1971)
Nake, F.: There should be no computer-art. Page, Bulletin of the Computer Arts Society (18), 1–2 (1971)
1971
-
[57]
Narasimhaswamy, S., Bhattacharya, U., Chen, X., Dasgupta, I., Mitra, S., Hoai, M.:Handiffuser:Text-to-imagegenerationwithrealistichandappearances.In:Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion. pp. 2468–2479 (2024)
2024
-
[58]
Nees, G.: Generative Computergraphik. Ph.D. thesis, University of Stuttgart (1969),https://archive.org/details/generative_computergraphik Naïve PAINE 19
1969
-
[59]
arXiv preprint arXiv:2507.00480 (2025)
Om, K., Sim, K., Yun, T., Kang, H., Park, J.: Posterior inference in latent space for scalable constrained black-box optimization. arXiv preprint arXiv:2507.00480 (2025)
Pith/arXiv arXiv 2025
-
[60]
ONOMAAI: Illustrious XL 2.0 - v2.0 | Illustrious Checkpoint | Civitai — civi- tai.com.https://civitai.com/models/1369089/illustrious-xl-20(2025),[Ac- cessed 27-01-2026]
arXiv 2025
-
[61]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4195–4205 (2023)
2023
-
[62]
von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., Wolf, T.: Diffusers: State-of-the-art diffusion models.https://github.com/huggingface/diffusers (2022)
2022
-
[63]
In: The Twelfth International Conference on Learning Represen- tations, ICLR 2024, Vienna, Austria, May 7-11, 2024
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: SDXL: improving latent diffusion models for high-resolution im- age synthesis. In: The Twelfth International Conference on Learning Represen- tations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net (2024), https://openreview.net/forum?id=di52zR8xgf
2024
-
[64]
O’Reilly Media, Inc
Provost, F., Fawcett, T.: Data Science for Business: What you need to know about data mining and data-analytic thinking. " O’Reilly Media, Inc." (2013)
2013
-
[65]
In: International conference on machine learning
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PmLR (2021)
2021
-
[66]
Journal of machine learning research21(140), 1–67 (2020)
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research21(140), 1–67 (2020)
2020
-
[67]
In: International conference on machine learning
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I.: Zero-shot text-to-image generation. In: International conference on machine learning. pp. 8821–8831. Pmlr (2021)
2021
-
[68]
In: CVPR (2022)
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR (2022)
2022
-
[69]
Advances in neural information processing systems35, 36479–36494 (2022)
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al.: Photorealistic text- to-image diffusion models with deep language understanding. Advances in neural information processing systems35, 36479–36494 (2022)
2022
-
[70]
arXiv preprint arXiv:2311.17042 (2023)
Sauer, A., Lorenz, D., Blattmann, A., Rombach, R.: Adversarial diffusion distilla- tion. arXiv preprint arXiv:2311.17042 (2023)
Pith/arXiv arXiv 2023
-
[71]
com/blog/archives/2017/02/predicting_a_sl.html, security commentary on the 2017 WIRED report about a Russian group reverse-engineering a Novomatic slot PRNG
Schneier, B.: Predicting a slot machine’s prng (Feb 2017),https://www.schneier. com/blog/archives/2017/02/predicting_a_sl.html, security commentary on the 2017 WIRED report about a Russian group reverse-engineering a Novomatic slot PRNG
2017
-
[72]
arXiv preprint arXiv:2510.10020 (2025)
Smith, H.D., Diamant, N.L., Trippe, B.L.: Calibrating generative models. arXiv preprint arXiv:2510.10020 (2025)
Pith/arXiv arXiv 2025
-
[73]
arXiv preprint arXiv:2010.02502 (2020)
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502 (2020)
Pith/arXiv arXiv 2010
-
[74]
arXiv preprint arXiv:2303.01469 (2023)
Song, Y., Dhariwal, P., Chen, M., Sutskever, I.: Consistency models. arXiv preprint arXiv:2303.01469 (2023)
Pith/arXiv arXiv 2023
-
[75]
Advances in Neural Information Processing Systems37(2024) 20 J
Sui, Y., Li, Y., Kag, A., Idelbayev, Y., Cao, J., Hu, J., Sagar, D., Yuan, B., Tulyakov, S., Ren, J.: Bitsfusion: 1.99 bits weight quantization of diffusion model. Advances in Neural Information Processing Systems37(2024) 20 J. Kim et al
2024
-
[76]
arXiv preprint arXiv:2410.10812 (2024)
Tang, H., Wu, Y., Yang, S., Xie, E., Chen, J., Chen, J., Zhang, Z., Cai, H., Lu, Y., Han, S.: Hart: Efficient visual generation with hybrid autoregressive transformer. arXiv preprint arXiv:2410.10812 (2024)
Pith/arXiv arXiv 2024
-
[77]
arXiv preprint arXiv:2405.18881 (2024)
Tang, Z., Peng, J., Tang, J., Hong, M., Wang, F., Chang, T.H.: Inference- time alignment of diffusion models with direct noise optimization. arXiv preprint arXiv:2405.18881 (2024)
Pith/arXiv arXiv 2024
-
[78]
arXiv preprint arXiv:2502.06999 (2025)
Venkatraman, S., Hasan, M., Kim, M., Scimeca, L., Sendera, M., Bengio, Y., Berseth, G., Malkin, N.: Outsourced diffusion sampling: Efficient posterior infer- ence in latent spaces of generative models. arXiv preprint arXiv:2502.06999 (2025)
Pith/arXiv arXiv 2025
-
[79]
In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition
Wang, J., Sun, Z., Tan, Z., Chen, X., Chen, W., Li, H., Zhang, C., Song, Y.: To- wards effective usage of human-centric priors in diffusion models for text-based human image generation. In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. pp. 8446–8455 (2024)
2024
-
[80]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Wang, R., Huang, H., Zhu, Y., Russakovsky, O., Wu, Y.: The silent assistant: Noisequery as implicit guidance for goal-driven image generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 17618–17628 (2025)
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.