A teacher-student reward model learns reasoning-conditioned score distributions for text-to-image images, yielding ~89% preference accuracy and a 41% net human-preference gain when used for generator optimization.
org/abs/2405.14705
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
CONDITIONAL 2representative citing papers
Universal aesthetic alignment in image models biases outputs toward conventional beauty and penalizes anti-aesthetic prompts even when they match explicit user instructions.
citing papers explorer
-
Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
A teacher-student reward model learns reasoning-conditioned score distributions for text-to-image images, yielding ~89% preference accuracy and a 41% net human-preference gain when used for generator optimization.
-
Position: Universal Aesthetic Alignment Narrows Artistic Expression
Universal aesthetic alignment in image models biases outputs toward conventional beauty and penalizes anti-aesthetic prompts even when they match explicit user instructions.