REVIEW 10 cited by
Improving Image Captioning with Better Use of Captions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics available in captions and leverage that to enhance both image representation and caption generation. Our models first construct caption-guided visual relationship graphs that introduce beneficial inductive bias using weakly supervised multi-instance learning. The representation is then enhanced with neighbouring and contextual nodes with their textual and visual features. During generation, the model further incorporates visual relationships using multi-task learning for jointly predicting word and object/predicate tag sequences. We perform extensive experiments on the MSCOCO dataset, showing that the proposed framework significantly outperforms the baselines, resulting in the state-of-the-art performance under a wide range of evaluation metrics.
Forward citations
Cited by 10 Pith papers
-
Trade-offs in Image Generation: How Do Different Dimensions Interact?
A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.
-
EMoE: Training-Free Expert Disagreement for Uncertainty-Aware Text-to-Image Diffusion
Expert disagreement inside pretrained MoE diffusion models, measured as latent variance at the first denoising step, gives a training-free prompt uncertainty signal that correlates with text-image alignment across languages.
-
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
A survey that organizes Omni-MLLMs into four architectural components and a taxonomy of encoding, alignment, and generation methods.
-
Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters
AdaVD removes target concepts from diffusion models by soft-projecting value vectors away from the target token direction, with a sigmoid threshold that preserves unrelated prompts.
-
Playable Game Generation
An autoregressive latent diffusion system, PlayGen, generates real-time playable Super Mario Bros and Doom sessions on an RTX 2060, with accuracy of game mechanics measured by action-recognition metrics.
-
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Selftok encodes images as diffusion-time-indexed discrete tokens, enabling a pure autoregressive VLM and visual RL with strong GenEval and DPG scores, though its claim that spatial tokens cannot support RL is not proven.
-
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
A continuous-time RL algorithm that treats diffusion scores as actions fine-tunes text-to-image models with a Girsanov-based KL regularizer, showing stability across different denoising step counts.
-
DuMo: Dual Encoder Modulation Network for Precise Concept Erasure
DuMo erases target concepts from text-to-image models by adding a frozen-backbone skip-connection eraser with learned timestep and layer modulation, reporting the best trade-off on three concept erasure benchmarks.
-
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training
Holistic CLIP trains a multi-branch image encoder with multi-to-multi contrastive learning on multiple VLM-generated captions per image and reports consistent gains over one-to-one and one-to-multi CLIP variants.
-
Cloud Platforms for Developing Generative AI Solutions: A Scoping Review of Tools and Services
A scoping review that aggregates and compares major cloud providers' generative AI tools and services, with no new empirical results.
Discussion (0). Continue with ORCID to comment.