Pith. sign in

REVIEW 1 cited by

VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16406 v1 pith:3OPTKOC2 submitted 2025-03-20 cs.GR cs.CVcs.MM

classification cs.GRcs.CVcs.MM
keywords interactionwordsinteractionsdiffusionimagesmodelmodelsobjects
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words. In this work, we propose VerbDiff to address the challenge of capturing nuanced interactions within text-to-image diffusion models. VerbDiff is a novel text-to-image generation model that weakens the bias between interaction words and objects, enhancing the understanding of interactions. Specifically, we disentangle various interaction words from frequency-based anchor words and leverage localized interaction regions from generated images to help the model better capture semantics in distinctive words without extra conditions. Our approach enables the model to accurately understand the intended interaction between humans and objects, producing high-quality images with accurate interactions aligned with specified verbs. Extensive experiments on the HICO-DET dataset demonstrate the effectiveness of our method compared to previous approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation

    cs.MM 2025-07 conditional novelty 6.0 of 10

    CatchPhrase improves audio-to-image generation by enriching weak class labels with LLM- and audio-caption-based prompts, filtering and retrieving the best prompt per clip, and training a mapping adapter with contrasti...

Pith tools