Pith. sign in

REVIEW 5 cited by

Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03411 v2 pith:R6HX3OCD submitted 2024-06-05 cs.CV

Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach

classification cs.CV
keywords retrievalplugircontextinteractiveapproachdialogue-formmethodologymodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we primarily address the issue of dialogue-form context query within the interactive text-to-image retrieval task. Our methodology, PlugIR, actively utilizes the general instruction-following capability of LLMs in two ways. First, by reformulating the dialogue-form context, we eliminate the necessity of fine-tuning a retrieval model on existing visual dialogue data, thereby enabling the use of any arbitrary black-box model. Second, we construct the LLM questioner to generate non-redundant questions about the attributes of the target image, based on the information of retrieval candidate images in the current context. This approach mitigates the issues of noisiness and redundancy in the generated questions. Beyond our methodology, we propose a novel evaluation metric, Best log Rank Integral (BRI), for a comprehensive assessment of the interactive retrieval system. PlugIR demonstrates superior performance compared to both zero-shot and fine-tuned baselines in various benchmarks. Additionally, the two methodologies comprising PlugIR can be flexibly applied together or separately in various situations. Our codes are available at https://github.com/Saehyung-Lee/PlugIR.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems

    cs.MA 2026-06 unverdicted novelty 6.0

    OCL is a governance layer for LLM agents that cuts unsafe executions from 88% to near-zero and raises valid success from 12% to 96% in adversarial buyer-seller negotiations across frontier LLMs.

  2. ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

    cs.CV 2026-06 unverdicted novelty 6.0

    ROGLE automates region-level supervision via Region-to-Sentence Matching and introduces the P-VLG benchmark to improve fine-grained alignment in text-based person search over CLIP-based models.

  3. DialogueVPR: Towards Conversational Visual Place Recognition

    cs.AI 2026-05 conditional novelty 6.0

    Dialogue-based place recognition lets an AI localize a place by asking clarifying questions, trained and evaluated on a GPT-4o-generated benchmark built from street-view images.

  4. GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval

    cs.CV 2025-09 conditional novelty 6.0

    GRAPE applies GRPO to an LLM query rewriter with a corpus-relative ranking reward to improve frozen CLIP retrieval by an average 4.9% Recall@10 on shifted benchmarks without retraining or re-embedding.

  5. ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

    cs.CV 2026-06 unverdicted novelty 5.0

    ROGLE introduces automated pseudo region-sentence pairs via RSM and multi-granular learning to boost fine-grained alignment in text-based person search, plus the P-VLG benchmark with over 100k annotated regions.