Pith. sign in

REVIEW 3 major objections 2 minor 3 cited by

VideoGuard: Protecting Video Content from Unauthorized Editing

T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read VideoGuard thwarts unauthorized AI video edits with subtle perturbations

desk verdict A plausible video-protection idea buried in an abstract, but the submission is a different paper, so there is nothing here to referee. read the letter →

arxiv 2508.03480 v1 pith:IMOXWBSZ submitted 2025-08-05 cs.CV cs.AI

classification cs.CVcs.AI
keywords videoprotectionunauthorizededitingdiffusionmodelsadversarialperturbationjointframeoptimizationmotioninformationgenerativemediasafety
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

VideoGuard is a proposed method for protecting videos from being edited by generative diffusion models. It adds nearly unnoticeable perturbations to the video that interfere with the target model's editing process, making the edited output implausible and inconsistent. The method jointly optimizes perturbations across all frames rather than per frame, because video diffusion models exploit inter-frame attention and frame redundancy. The paper reports that this video-level protection outperforms baseline methods on both objective and subjective metrics.

What carries the argument

Joint frame optimization with motion fusion: all video frames are treated as a single optimization entity, and motion information extracted from the video is fused into the optimization objective. This carries the argument by ensuring the perturbation is coherent across frames and tied to the video's dynamics, so it can disrupt the inter-frame attention of the target diffusion model and produce inconsistent outputs.

What would settle it

Apply VideoGuard to a video, then edit the protected video with a diffusion model that was not part of the intended set, or after re-encoding/compressing the video; if the editing succeeds and produces a plausible result, the protection's transferability fails.

Watch

Extended reading notes

Core claim

The central claim is that video content can be shielded from unauthorized editing by a video-scale adversarial perturbation. Jointly optimizing the perturbation across the whole video, while fusing extracted motion information into the objective, forces the intended video diffusion model to produce edited results that are implausible and temporally inconsistent. This is presented as a solution to the failure of per-frame image protection methods, which cannot fully disrupt the inter-frame attention mechanism of video diffusion models.

Load-bearing premise

The perturbation is optimized to interfere with a specific 'intended' diffusion model, and the paper gives no evidence that the protection carries over to a different model or survives video compression or re-encoding.

Editorial extensions

If this is right

  • Content creators could publish videos that resist manipulation by generative editing tools, preserving authenticity.
  • Protection becomes a video-level property rather than a per-frame hack, sidestepping temporal inconsistencies that per-frame methods suffer.
  • Motion information is shown to be a usable channel for adversarial defenses against diffusion-based editing.
  • The paper's objective and subjective evaluations indicate superior protection over per-frame image-based baselines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same joint-optimization idea could be tested against video inpainting, style transfer, or other generative video tasks, not just editing diffusion models.
  • Because the perturbation is optimized for an intended model, its transferability to unseen models or robustness to compression and re-encoding remains open and is the most likely failure point.
  • Motion-fused perturbations might be detectable by automated temporal artifact filters, so stealth could be a future battleground.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The manuscript claims a method called VideoGuard that protects video content from unauthorized editing by adding 'nearly unnoticeable perturbations' that interfere with 'the intended generative diffusion models.' The abstract describes joint frame optimization and motion-information fusion, and asserts that VideoGuard outperforms all baseline methods on both objective and subjective metrics. However, the full text attached to the submission is an entirely different paper, 'When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models' (arXiv:2508.03483v3), which has no relation to VideoGuard. Consequently, the submitted material contains no description of the method, no optimization pipeline, no experiments, no baselines, and no results that could be assessed.

Significance. If the claims in the abstract were substantiated, video-level protection against unauthorized editing would be a meaningful advance beyond per-frame image-based protection, particularly because video diffusion models exploit inter-frame attention and redundancy. The proposed joint frame optimization and motion fusion are plausible ingredients for such a defense. However, the paper as submitted provides no verifiable support for any of its central claims: there is no protocol, no dataset, no target models, no metrics, no ablations, and no comparison. The only actual text available is the abstract, which alone cannot establish effectiveness or superiority over baselines. In addition, the abstract explicitly limits the protection to 'the intended generative diffusion models,' raising unresolved questions about transferability to unseen models and robustness to common preprocessing that an adversary might apply. The significance of the work therefore cannot be evaluated from the material provided.

major comments (3)
  1. [Full text (entire manuscript)] The full text of the submission is a different paper (arXiv:2508.03483v3, 'When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models') that is unrelated to the abstract's claimed VideoGuard method. As a result, the manuscript contains none of the promised content: no formulation of the perturbation optimization, no description of joint frame optimization or motion-information fusion, no objective or subjective evaluation protocols, no baseline methods, and no experimental results. The central claim that VideoGuard 'can effectively protect videos' and is 'superior to all the baseline methods' is therefore entirely unsupported in the submitted material. This is a load-bearing defect that cannot be remedied by local revision; it requires resubmission of the actual VideoGuard paper.
  2. [Abstract] Even taking the abstract in isolation, the claim of superiority over 'all the baseline methods' is unverifiable because no baselines are named, no dataset or target models are specified, and no error bars or statistical tests are reported. The abstract states that 'the results show that the protection performance of VideoGuard is superior to all the baseline methods,' but it does not provide any of the numbers or the experimental setup needed to check this claim. Without this information, the claim is an assertion rather than a result.
  3. [Abstract] The abstract defines the protection as interfering with 'the intended generative diffusion models,' which indicates a white-box or model-specific optimization. The manuscript provides no evidence that the perturbations transfer to other video diffusion models or survive common post-processing such as compression, re-encoding, resizing, or frame dropping that an unauthorized editor might apply. Because the stated goal is to 'protect videos from unauthorized malicious editing' in general, the reliance on a specific intended model is a load-bearing limitation that needs explicit demonstration of robustness or a clear scope restriction.
minor comments (2)
  1. [Abstract] The phrase 'nearly unnoticeable perturbations' is not quantified; the manuscript would benefit from stating the perturbation bound (e.g., L-infinity norm) and the perceptual metric used to verify imperceptibility.
  2. [Abstract] The sentence 'Thus, these alterations can effectively force the models to produce outputs that are implausible and inconsistent' would be clearer with a concrete definition of 'implausible' and 'inconsistent' in terms of the evaluation metrics used.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be established from the available text; the abstract describes a standard white-box adversarial-perturbation setup with no equations or fitted parameters to audit.

full rationale

The supplied material for arXiv:2508.03480 consists only of the VideoGuard abstract; the appended full text is a different, unrelated paper on demographic bias in text-to-image models. There is no derivation chain, equation set, or experimental section for VideoGuard available to inspect. The abstract's statement that perturbations 'interfere with the functioning of the intended generative diffusion models' and that performance is then demonstrated with 'objective metrics and subjective metrics' describes a conventional white-box attack/protection evaluation; it does not, on its face, define the target metric in terms of the optimization objective, fit a parameter and then relabel that fit as a prediction, or rely on a self-citation for a load-bearing premise. The concern that protection may not transfer to unseen models or survive compression is a robustness and generality limitation, not a circularity. Under the hard rule requiring a quoted equation-level reduction or other explicit circular step, no such step can be exhibited from the abstract alone. Therefore the honest finding is no significant circularity.

Assumptions & free parameters 2 free parameters · 2 assumptions · 0 invented entities

Because the submission's full text is an unrelated paper, the ledger is built from the abstract alone. The entries are the assumptions needed for the abstract's claims to hold.

free parameters (2)
  • Per-frame perturbation magnitude bound
    Implied by 'nearly unnoticeable perturbations'; no value or perceptibility threshold is given in the abstract, and this bound is load-bearing for the invisibility claim.
  • Motion-information fusion weight
    The abstract says motion information is fused into the optimization objectives; the relative weighting between attack strength and motion consistency is not reported in the abstract.
assumptions (2)
  • domain assumption Video diffusion models depend on inter-frame attention and cross-frame redundancy, so frame-wise independent perturbations fail.
    This is the stated motivation for joint-frame optimization in the abstract; if this premise is false, the central design choice is unnecessary.
  • domain assumption The objective and subjective metrics used in evaluation measure both protection effectiveness and imperceptibility.
    The abstract claims superiority on both metric types but provides no protocol; the claim of 'nearly unnoticeable' perturbations depends on these measurements.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VideoGuard: Protecting Video Content from Unauthorized Editing." pith.science (2026). https://pith.science/paper/IMOXWBSZ

@misc{pith2026250803480,
  author       = {Pith},
  title        = {Pith review of: VideoGuard: Protecting Video Content from Unauthorized Editing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IMOXWBSZ}},
  note         = {Machine review of arXiv:2508.03480}
}
read the original abstract

With the rapid development of generative technology, current generative models can generate high-fidelity digital content and edit it in a controlled manner. However, there is a risk that malicious individuals might misuse these capabilities for misleading activities. Although existing research has attempted to shield photographic images from being manipulated by generative models, there remains a significant disparity in the protection offered to video content editing. To bridge the gap, we propose a protection method named VideoGuard, which can effectively protect videos from unauthorized malicious editing. This protection is achieved through the subtle introduction of nearly unnoticeable perturbations that interfere with the functioning of the intended generative diffusion models. Due to the redundancy between video frames, and inter-frame attention mechanism in video diffusion models, simply applying image-based protection methods separately to every video frame can not shield video from unauthorized editing. To tackle the above challenge, we adopt joint frame optimization, treating all video frames as an optimization entity. Furthermore, we extract video motion information and fuse it into optimization objectives. Thus, these alterations can effectively force the models to produce outputs that are implausible and inconsistent. We provide a pipeline to optimize this perturbation. Finally, we use both objective metrics and subjective metrics to demonstrate the efficacy of our method, and the results show that the protection performance of VideoGuard is superior to all the baseline methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    TC-UAP learns a shared multi-frame adversarial perturbation that protects videos of the same identity from both fine-tuning-based and reference-based video customization, remaining effective on unseen clips and under ...

  2. Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    REINS uses supervised PCA on safety-labeled activations to find a linear direction that, when added to hidden states at roughly 50% depth in video diffusion transformers, redirects generations from unsafe to safe cont...

  3. Polarisation multiplexing ring-cavity fibre laser for dual-comb generation

    physics.optics 2025-08 unverdicted novelty 4.0 of 10

    A single-cavity polarisation-multiplexed fibre laser produces two stable optical combs with sub-millimetre ranging precision over 250 hours.

Reference graph

Works this paper leans on

47 extracted references · 31 canonical work pages · cited by 3 Pith papers

  1. [1]

    arXiv preprint arXiv:2511.21631 (2025) 14

    Bai, S., Cai, Y ., Chen, R., Chen, K., Chen, X., Cheng, Z., Deng, L., Ding, W., Gao, C., Ge, C., et al.: Qwen3-vl technical report. arXiv preprint arXiv:2511.21631 (2025) 14

  2. [2]

    In: SIGCIS (2017) 2

    Barocas, S., Crawford, K., Shapiro, A., Wallach, H.: The problem with bias: Al- locative versus representational harms in machine learning. In: SIGCIS (2017) 2

  3. [3]

    arXiv preprint arXiv:2505.20692 (2025) 1

    Barve, S., Mao, A., Shi, J.M., Juneja, P., Saha, K.: Can we debias social stereo- types in ai-generated images? examining text-to-image outputs and user percep- tions. arXiv preprint arXiv:2505.20692 (2025) 1

  4. [4]

    In: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision

    Basu, A., Babu, R.V ., Pruthi, D.: Inspecting the geographical representativeness of images from text-to-image models. In: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision. pp. 5136–5147 (2023) 3

  5. [5]

    arXiv preprint arXiv:2602.21012 (2026) 1

    Bengio, Y ., Clare, S., Prunkl, C., Andriushchenko, M., Bucknall, B., Murray, M., Bommasani, R., Casper, S., Davidson, T., Douglas, R., et al.: International ai safety report 2026. arXiv preprint arXiv:2602.21012 (2026) 1

  6. [6]

    In: Proceedings of the 2023 ACM conference on fairness, accountability, and transparency

    Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., Caliskan, A.: Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. In: Proceedings of the 2023 ACM conference on fairness, accountability, and transparency. pp. 1493–1504 (2023) 1, 2, 3, 4

  7. [7]

    https://bfl.ai/blog/flux-2 (2025), accessed: 2026-02-15 6

    Black Forest Labs: FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2 (2025), accessed: 2026-02-15 6

  8. [8]

    In: European Conference on Computer Vision

    Chinchure, A., Shukla, P., Bhatt, G., Salij, K., Hosanagar, K., Sigal, L., Turk, M.: Tibet: Identifying and evaluating biases in text-to-image generative models. In: European Conference on Computer Vision. pp. 429–446. Springer (2024) 3

Show all 47 references
  1. [9]

    In: Proceedings of the IEEE/CVF in- ternational conference on computer vision

    Cho, J., Zala, A., Bansal, M.: Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In: Proceedings of the IEEE/CVF in- ternational conference on computer vision. pp. 3043–3054 (2023) 3

  2. [10]

    arXiv preprint arXiv:2507.06261 (2025) 14

    Comanici, G., Bieber, E., Schaekermann, M., Pasupat, I., Sachdeva, N., Dhillon, I., Blistein, M., Ram, O., Zhang, D., Rosen, E., et al.: Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next genera- tion agentic capabilities. arXiv pre...

  3. [11]

    International Journal of Computer Vision 133(7), 4555–4570 (2025) 1

    Cong, Y ., Min, M.R., Li, L.E., Rosenhahn, B., Yang, M.Y .: Attribute-centric compositional text-to-image generation. International Journal of Computer Vision 133(7), 4555–4570 (2025) 1

  4. [12]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    D’Incà, M., Peruzzo, E., Mancini, M., Xu, D., Goel, V ., Xu, X., Wang, Z., Shi, H., Sebe, N.: Openbias: Open-set bias detection in text-to-image generative models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 12225–12235 (2024) 3

  5. [13]

    arXiv preprint arXiv:2310.19981 (2023) 3 16 D

    Ghosh, S., Caliskan, A.: ’person’== light-skinned, western man, and sexu- alization of women of color: Stereotypes in stable diffusion. arXiv preprint arXiv:2310.19981 (2023) 3 16 D. Choi et al

  6. [14]

    https://deepmind.google/models/imagen/ (2024), accessed: 2025-01-27 6

    Google DeepMind: Imagen 4. https://deepmind.google/models/imagen/ (2024), accessed: 2025-01-27 6

  7. [15]

    https://blog.google/innovation- and- ai/ technology/ai/nano-banana-2 (2026), accessed: 2026-06-17 14, 24

    Google DeepMind: Nano banana 2. https://blog.google/innovation- and- ai/ technology/ai/nano-banana-2 (2026), accessed: 2026-06-17 14, 24

  8. [16]

    arXiv preprint arXiv:2308.06198 (2023) 3

    Hall, M., Ross, C., Williams, A., Carion, N., Drozdzal, M., Soriano, A.R.: Dig in: Evaluating disparities in image generations with indicators for geographic diver- sity. arXiv preprint arXiv:2308.06198 (2023) 3

  9. [17]

    Advances in Neural Information Processing Systems36, 63687–63723 (2023) 3

    Hall, S.M., Gonçalves Abrantes, F., Zhu, H., Sodunke, G., Shtedritski, A., Kirk, H.R.: Visogender: A dataset for benchmarking gender bias in image-text pronoun resolution. Advances in Neural Information Processing Systems36, 63687–63723 (2023) 3

  10. [18]

    a study on used ai text-to-image in architecture

    Hanafy, N.O.: Artificial intelligence’s effects on design process creativity:" a study on used ai text-to-image in architecture". Journal of Building Engineering80, 107999 (2023) 1

  11. [19]

    Holstein, K., Wortman Vaughan, J., Daumé III, H., Dudik, M., Wallach, H.: Im- proving fairness in machine learning systems: What do industry practitioners need? In: Proceedings of the 2019 CHI conference on human factors in computing systems. pp. 1–16 (2019) 1

  12. [20]

    In: Proceedings of the Computer Vision and Pattern Recognition Conference

    Huang, H., Jin, X., Miao, J., Wu, Y .: Implicit bias injection attacks against text- to-image diffusion models. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 28779–28789 (2025) 2

  13. [21]

    IEEE Transactions on Information theory37(1), 145–151 (2002) 6

    Lin, J.: Divergence measures based on the shannon entropy. IEEE Transactions on Information theory37(1), 145–151 (2002) 6

  14. [22]

    In: European confer- ence on computer vision

    Lin, T.Y ., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European confer- ence on computer vision. pp. 740–755. Springer (2014) 6

  15. [23]

    arXiv preprint arXiv:2307.08025 (2023) 3

    Mannering, H.: Analysing gender bias in text-to-image models using object de- tection. arXiv preprint arXiv:2307.08025 (2023) 3

  16. [24]

    arXiv preprint arXiv:1802.03426 (2018) 10

    McInnes, L., Healy, J., Melville, J.: UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018) 10

  17. [25]

    In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society

    Naik, R., Nushi, B.: Social biases through the text-to-image generation lens. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. pp. 786–808 (2023) 1, 3

  18. [26]

    https://platform.openai.com/docs/ guides/image-generation (2025), accessed: 2025-07-27 6

    OpenAI: Gpt image-1: Image generation api. https://platform.openai.com/docs/ guides/image-generation (2025), accessed: 2025-07-27 6

  19. [27]

    In: Proceedings of the 25th international academic mindtrek conference

    Oppenlaender, J.: The creativity of text-to-image generation. In: Proceedings of the 25th international academic mindtrek conference. pp. 192–202 (2022) 1

  20. [28]

    arXiv preprint arXiv:2307.01952 (2023) 6

    Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: Sdxl: Improving latent diffusion models for high-resolution im- age synthesis. arXiv preprint arXiv:2307.01952 (2023) 6

  21. [29]

    In: International conference on machine learning

    Ramesh, A., Pavlov, M., Goh, G., Gray, S., V oss, C., Radford, A., Chen, M., Sutskever, I.: Zero-shot text-to-image generation. In: International conference on machine learning. pp. 8821–8831. Pmlr (2021) 1

  22. [30]

    arXiv preprint arXiv:2507.13383 (2025) 4 SODA 17

    Rastogi, C., Teh, T.H., Mishra, P., Patel, R., Wang, D., Díaz, M., Parrish, A., Da- vani, A.M., Ashwood, Z., Paganini, M., et al.: Whose view of safety? a deep dive dataset for pluralistic alignment of text-to-image models. arXiv preprint arXiv:2507.13383 (2025) 4 SODA 17

  23. [31]

    arXiv preprint arXiv:2503.02865 (2025) 1

    Raza, S., Chettiar, M.S., Yousefabadi, M., Khan, T., Lotif, M.: Fairsense-ai: Re- sponsible ai meets sustainability. arXiv preprint arXiv:2503.02865 (2025) 1

  24. [32]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10684–10695 (2022) 1

  25. [33]

    Advances in neural information processing systems35, 36479–36494 (2022) 1

    Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al.: Photorealistic text- to-image diffusion models with deep language understanding. Advances in neural information processing systems35,...

  26. [34]

    arXiv preprint arXiv:2308.00755 (2023) 1, 3, 4

    Seshadri, P., Singh, S., Elazar, Y .: The bias amplification paradox in text-to-image generation. arXiv preprint arXiv:2308.00755 (2023) 1, 3, 4

  27. [35]

    The Bell system tech- nical journal27(3), 379–423 (1948) 6

    Shannon, C.E.: A mathematical theory of communication. The Bell system tech- nical journal27(3), 379–423 (1948) 6

  28. [36]

    In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems

    Shelby, R., Rismani, S., Rostamzadeh, N.: Generative ai in creative practice: Ml- artist folk theories of t2i use, harm, and harm-reduction. In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. pp. 1–17 (2024) 1

  29. [37]

    In: Proceed- ings of the Computer Vision and Pattern Recognition Conference

    Shimoda, W., Inoue, N., Haraguchi, D., Mitani, H., Uchida, S., Yamaguchi, K.: Type-r: Automatically retouching typos for text-to-image generation. In: Proceed- ings of the Computer Vision and Pattern Recognition Conference. pp. 2745–2754 (2025) 1

  30. [38]

    In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency

    Steed, R., Caliskan, A.: Image representations learned with unsupervised pre- training contain human-like biases. In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. pp. 701–713 (2021) 3

  31. [39]

    In: International symposium on end user development

    Turchi, T., Carta, S., Ambrosini, L., Malizia, A.: Human-ai co-creation: evaluating the impact of large-scale text-to-image generative models on the creative process. In: International symposium on end user development. pp. 35–51. Springer (2023) 1

  32. [40]

    arXiv preprint arXiv:2503.08012 (2025) 1

    Vice, J., Akhtar, N., Hartley, R., Mian, A.: Exploring bias in over 100 text-to- image generative models. arXiv preprint arXiv:2503.08012 (2025) 1

  33. [41]

    CoRR (2024) 3

    Wan, Y ., Subramonian, A., Ovalle, A., Lin, Z., Suvarna, A., Chance, C., Bansal, H., Pattichis, R., Chang, K.W.: Survey of bias in text-to-image generation: Defini- tion, evaluation, and mitigation. CoRR (2024) 3

  34. [42]

    In: ICML (2021) 2

    Wang, A., Russakovsky, O.: Directional bias amplification. In: ICML (2021) 2

  35. [43]

    In: Findings of the Association for Computational Linguistics: ACL 2023

    Wang, J., Liu, X., Di, Z., Liu, Y ., Wang, X.: T2iat: Measuring valence and stereo- typical biases in text-to-image generation. In: Findings of the Association for Computational Linguistics: ACL 2023. pp. 2560–2574 (2023) 3

  36. [44]

    arXiv preprint arXiv:2508.02324 (2025) 6

    Wu, C., Li, J., Zhou, J., Lin, J., Gao, K., Yan, K., Yin, S.m., Bai, S., Xu, X., Chen, Y ., et al.: Qwen-image technical report. arXiv preprint arXiv:2508.02324 (2025) 6

  37. [45]

    In: Proceedings of the AAAI/ACM conference on AI, ethics, and society

    Wu, Y ., Nakashima, Y ., Garcia, N.: Stable diffusion exposed: Gender bias from prompt to image. In: Proceedings of the AAAI/ACM conference on AI, ethics, and society. vol. 7, pp. 1648–1659 (2024) 3

  38. [46]

    arXiv preprint arXiv:2407.00600 (2024) 3 18 D

    Xiao, Y ., Liu, A., Cheng, Q., Yin, Z., Liang, S., Li, J., Shao, J., Liu, X., Tao, D.: Genderbias-vl: Benchmarking gender bias in vision language models via counter- factual probing. arXiv preprint arXiv:2407.00600 (2024) 3 18 D. Choi et al

  39. [47]

    navy_blue

    Zhang, C., Zhang, C., Zhang, M., Kweon, I.S.: Text-to-image diffusion models in generative ai: A survey. arXiv preprint arXiv:2303.07909 (2023) 1 A Generation Parameters For GPT Image-1, we use high-quality mode with 1024×1024 resolution via OpenAI API. The remaining models ar...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.