Pith. sign in

REVIEW 2 cited by

CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12579 v3 pith:AP3A5CXR submitted 2025-02-18 cs.CV

classification cs.CV
keywords alignmentchatsgenerationmodelssamplingtext-to-imagehumantest-time
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have emerged as a dominant approach for text-to-image generation. Key components such as the human preference alignment and classifier-free guidance play a crucial role in ensuring generation quality. However, their independent application in current text-to-image models continues to face significant challenges in achieving strong text-image alignment, high generation quality, and consistency with human aesthetic standards. In this work, we for the first time, explore facilitating the collaboration of human performance alignment and test-time sampling to unlock the potential of text-to-image models. Consequently, we introduce CHATS (Combining Human-Aligned optimization and Test-time Sampling), a novel generative framework that separately models the preferred and dispreferred distributions and employs a proxy-prompt-based sampling strategy to utilize the useful information contained in both distributions. We observe that CHATS exhibits exceptional data efficiency, achieving strong performance with only a small, high-quality funetuning dataset. Extensive experiments demonstrate that CHATS surpasses traditional preference alignment methods, setting new state-of-the-art across various standard benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance

    cs.CV 2025-07 conditional novelty 6.0 of 10

    TeEFusion distills classifier-free guidance into text embeddings via linear fusion, enabling a student model to generate images in one forward pass instead of two.

  2. Reinforcement Learning for Flow-Matching Policies

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Reward-weighted flow matching and GRPO with a learned reward surrogate both improve flow-matching policies beyond a suboptimal demonstrator on simulated unicycle tasks.

Pith tools