Pith. sign in

Generating an image from 1,000 words: Enhancing text-to-image with structured captions

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 6 cs.AI 1

years

2026 7

roles

background 1

polarities

support 1

representative citing papers

ReflectCAP: Detailed Image Captioning with Reflective Memory

cs.AI · 2026-04-14 · unverdicted · novelty 6.0

ReflectCAP distills model-specific hallucination and oversight patterns into Structured Reflection Notes that steer LVLMs toward more factual and complete image captions, reaching the Pareto frontier on factuality-coverage trade-offs.

APE: Agentic Prompt Enhancer for Image Generation and Editing

cs.CV · 2026-05-29 · unverdicted · novelty 5.0

APE post-trains small language models as single-agent or multi-agent prompt enhancers that improve visual alignment on image generation and editing benchmarks without altering the downstream visual model.

LTX-2: Efficient Joint Audio-Visual Foundation Model

cs.CV · 2026-01-06 · conditional · novelty 5.0

LTX-2 generates high-quality synchronized audiovisual content from text prompts via an asymmetric 14B-video / 5B-audio dual-stream transformer with cross-attention and modality-aware guidance.

Token-to-Token Alignment of Text Embeddings for Semantic Blending

cs.CV · 2026-06-22 · unverdicted · novelty 4.0

Token-to-Token alignment rephrases prompts into shared structure then matches token embeddings by semantic similarity, making linear interpolation a meaningful operation for blending in text-to-image models.

citing papers explorer

Showing 7 of 7 citing papers.