Pith. sign in

REVIEW 5 cited by

Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11457 v1 pith:3NLPO4RK submitted 2025-04-15 cs.CV

Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception

classification cs.CV
keywords perceptiongenerativedenoisingtasksdiscriminativediffusionimagemodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

With the success of image generation, generative diffusion models are increasingly adopted for discriminative tasks, as pixel generation provides a unified perception interface. However, directly repurposing the generative denoising process for discriminative objectives reveals critical gaps rarely addressed previously. Generative models tolerate intermediate sampling errors if the final distribution remains plausible, but discriminative tasks require rigorous accuracy throughout, as evidenced in challenging multi-modal tasks like referring image segmentation. Motivated by this gap, we analyze and enhance alignment between generative diffusion processes and perception tasks, focusing on how perception quality evolves during denoising. We find: (1) earlier denoising steps contribute disproportionately to perception quality, prompting us to propose tailored learning objectives reflecting varying timestep contributions; (2) later denoising steps show unexpected perception degradation, highlighting sensitivity to training-denoising distribution shifts, addressed by our diffusion-tailored data augmentation; and (3) generative processes uniquely enable interactivity, serving as controllable user interfaces adaptable to correctional prompts in multi-round interactions. Our insights significantly improve diffusion-based perception models without architectural changes, achieving state-of-the-art performance on depth estimation, referring image segmentation, and generalist perception tasks. Code available at https://github.com/ziqipang/ADDP.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

    cs.AI 2026-05 conditional novelty 7.0

    EngiAI introduces a LangGraph-based multi-agent framework and a three-part benchmark suite for LLM-driven engineering design, reporting high task completion rates for proprietary models on Beams2D and Photonics2D problems.

  2. FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching

    cs.CV 2026-05 unverdicted novelty 7.0

    FlowDIS uses flow matching to transport image distributions to mask distributions, optionally conditioned on text, and outperforms prior DIS methods by 5.5% on F_beta^omega and 43% on MAE.

  3. Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    OTCA improves GRPO training for visual generation by estimating step importance in trajectories and adaptively weighting multiple reward objectives.

  4. EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

    cs.AI 2026-05 unverdicted novelty 6.0

    EngiAI is a multi-agent framework unifying topology optimization, retrieval, HPC orchestration, and manufacturing control, with benchmarks showing proprietary LLMs at 96-97% task completion on Beams2D and lower perfor...

  5. FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching

    cs.CV 2026-05 unverdicted novelty 6.0

    FlowDIS uses flow matching to transport image distributions to mask distributions with language guidance and PAIP training, outperforming prior DIS methods by 5.5% on F_beta^omega and 43% on MAE on DIS-TE.