Pith. sign in

REVIEW 10 cited by

Improving Image Captioning with Better Use of Captions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.11807 v1 pith:LQO4VHO4 submitted 2020-06-21 cs.CV cs.CL

classification cs.CVcs.CL
keywords imagecaptioningvisualbettercaptionsextensivegenerationlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics available in captions and leverage that to enhance both image representation and caption generation. Our models first construct caption-guided visual relationship graphs that introduce beneficial inductive bias using weakly supervised multi-instance learning. The representation is then enhanced with neighbouring and contextual nodes with their textual and visual features. During generation, the model further incorporates visual relationships using multi-task learning for jointly predicting word and object/predicate tag sequences. We perform extensive experiments on the MSCOCO dataset, showing that the proposed framework significantly outperforms the baselines, resulting in the state-of-the-art performance under a wide range of evaluation metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trade-offs in Image Generation: How Do Different Dimensions Interact?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.

  2. EMoE: Training-Free Expert Disagreement for Uncertainty-Aware Text-to-Image Diffusion

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Expert disagreement inside pretrained MoE diffusion models, measured as latent variance at the first denoising step, gives a training-free prompt uncertainty signal that correlates with text-image alignment across languages.

  3. From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A survey that organizes Omni-MLLMs into four architectural components and a taxonomy of encoding, alignment, and generation methods.

  4. Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AdaVD removes target concepts from diffusion models by soft-projecting value vectors away from the target token direction, with a sigmoid threshold that preserves unrelated prompts.

  5. Playable Game Generation

    cs.AI 2024-12 conditional novelty 6.0 of 10

    An autoregressive latent diffusion system, PlayGen, generates real-time playable Super Mario Bros and Doom sessions on an RTX 2060, with accuracy of game mechanics measured by action-recognition metrics.

  6. Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

    cs.CV 2025-05 reject novelty 5.0 of 10

    Selftok encodes images as diffusion-time-indexed discrete tokens, enabling a pure autoregressive VLM and visual RL with strong GenEval and DPG scores, though its claim that spatial tokens cannot support RL is not proven.

  7. Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A continuous-time RL algorithm that treats diffusion scores as actions fine-tunes text-to-image models with a Girsanov-based KL regularizer, showing stability across different denoising step counts.

  8. DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DuMo erases target concepts from text-to-image models by adding a frozen-backbone skip-connection eraser with learned timestep and layer modulation, reporting the best trade-off on three concept erasure benchmarks.

  9. Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Holistic CLIP trains a multi-branch image encoder with multi-to-multi contrastive learning on multiple VLM-generated captions per image and reports consistent gains over one-to-one and one-to-multi CLIP variants.

  10. Cloud Platforms for Developing Generative AI Solutions: A Scoping Review of Tools and Services

    cs.DC 2024-12 conditional novelty 1.0 of 10

    A scoping review that aggregates and compares major cloud providers' generative AI tools and services, with no new empirical results.

Pith tools