DRL trains a discriminator on data versus base-model samples in pretrained representation space and uses its logit as reward in KL-regularized RL, cutting guidance-free FID from 9.38 to 2.62 on SiT and similar gains on other backbones.
Finetuning text-to- image diffusion models for fairness
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7verdicts
UNVERDICTED 7representative citing papers
EquiSteer reduces average demographic parity gaps by up to 87% in Stable Diffusion variants and SANA via inference-time cross-attention steering with minimal impact on image quality or alignment.
RG-TTA uses reinforcement learning at test time to gate fairness regularization by estimated bias sensitivity, reducing stereotypes on FairFace and UTKFace while improving zero-shot utility.
IR-guided diffusion injects intermediate text representations into early denoising steps to improve alignment for one-and-only objects, reporting up to 19.1pp VQAScore gains on OAO-AttackBench and other benchmarks.
VSM modulates the score Jacobian using variance guidance to reduce hallucinations in diffusion models by up to 25% on synthetic and real datasets while preserving fidelity and diversity.
Embedding Arithmetic performs vector operations in the embedding space of T2I models to mitigate bias at inference time, outperforming baselines on diversity while preserving coherence via a new Concept Coherence Score.
Implicit generative choices in diffusion models concentrate in self-attention layers; targeted ICM interventions there outperform broader debiasing methods with fewer artifacts.
citing papers explorer
-
The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL
DRL trains a discriminator on data versus base-model samples in pretrained representation space and uses its logit as reward in KL-regularized RL, cutting guidance-free FID from 9.38 to 2.62 on SiT and similar gains on other backbones.
-
EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
EquiSteer reduces average demographic parity gaps by up to 87% in Stable Diffusion variants and SANA via inference-time cross-attention steering with minimal impact on image quality or alignment.
-
Selective Test-Time Debiasing for CLIP via Reward Gating
RG-TTA uses reinforcement learning at test time to gate fairness regularization by estimated bias sensitivity, reducing stereotypes on FairFace and UTKFace while improving zero-shot utility.
-
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment
IR-guided diffusion injects intermediate text representations into early denoising steps to improve alignment for one-and-only objects, reporting up to 19.1pp VQAScore gains on OAO-AttackBench and other benchmarks.
-
Score-Control for Hallucination Reduction in Diffusion Models
VSM modulates the score Jacobian using variance guidance to reduce hallucinations in diffusion models by up to 25% on synthetic and real datasets while preserving fidelity and diversity.
-
Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models
Embedding Arithmetic performs vector operations in the embedding space of T2I models to mitigate bias at inference time, outperforming baselines on diversity while preserving coherence via a new Concept Coherence Score.
-
Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
Implicit generative choices in diffusion models concentrate in self-attention layers; targeted ICM interventions there outperform broader debiasing methods with fewer artifacts.