Pith. sign in

Paper Citation Record · LEDGER

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 7 inbound Pith citation observations for arXiv:2505.19952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19952 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:40.730204Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:10.770726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:05:41.143246Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact4
  • verified fuzzy15
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f6c19bf-fa02-4ae7-8ad0-fd89ac999fa8 · outbound

This paper cites Distribution consistency guided hashing for cross-modal retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Distribution consistency guided hashing for cross-modal retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:34.933454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:34.933454Z digest=sha256:710a5d91c5a5cc26c8959e6c3553b2164776431844ad359e4126e29e677bd033

Observation 538ab951-db2c-40b6-b51a-8f47eef29e71 · outbound

This paper cites Similarity transitivity broken-aware multi-modal hashing.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Similarity transitivity broken-aware multi-modal hashing.IEEE Trans

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.023106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.023106Z digest=sha256:cc14302ad29b90c0e8fe0a5b33f31f197bc4b36dffe0bcf8ad7da3b0057cb459

Observation 9d9f5b31-c484-4436-8dec-1c341b1491cd · outbound

This paper cites Data-aware proxy hashing for cross-modal retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Data-aware proxy hashing for cross-modal retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.062654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.062654Z digest=sha256:765c96578674dbaef4e4b0b5d50178b6eef6a27a06d5a2686af9857a693bd4dc

Observation 51b2802e-9064-479a-9e56-6f461d220588 · outbound

This paper cites Where does the performance improvement come from?: - A reproducibility concern about image-text retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Where does the performance improvement come from?: - A reproducibility concern about image-text retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.168550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.168550Z digest=sha256:c0c079a6c00352a3b35c5d5f734c952b8c2cbc720bac42c31345349d6fdd6dab

Observation a8855576-087d-471a-9ace-60361e5e8f85 · outbound

This paper cites Com- posing text and image for image retrieval - an empirical odyssey.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Com- posing text and image for image retrieval - an empirical odyssey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.291206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.291206Z digest=sha256:63e364758bd07d973081c8c2966c5aa074105cdc41a370734575706105be7e18

Observation 8c94eda2-0b0d-4e67-b969-4e28046fa20d · outbound

This paper cites Sim- ple but effective raw-data level multimodal fusion for composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Sim- ple but effective raw-data level multimodal fusion for composed image retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.951301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:35.377153Z digest=sha256:93fdd51344f7e8c7a0069a95f8b8764f97e0de77dd08295b700c5771e9a64f65

Observation 1754bc0b-4231-47f3-bbb5-5a5c2e5a0b7b · outbound

This paper cites Dual-path semantic construction network for composed query-based image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Dual-path semantic construction network for composed query-based image retrieval

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.561828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.561828Z digest=sha256:1d4760f707452746702fab04dc49cb394fd87a307d415277f0c669005ecd281a

Observation ed7928b0-7826-495d-aec2-c096137479c1 · outbound

This paper cites Multi-modal transformer with global-local alignment for composed query image retrieval.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Multi-modal transformer with global-local alignment for composed query image retrieval.IEEE Trans

Reference 8

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:43.016620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:35.694647Z digest=sha256:c6c084b9b5e813a99ea80eeb68f3210cb1fe04ed21b73a0b1202552847163c0b

Observation 301e70e5-0c60-4b16-bd29-1a0d622023ad · outbound

This paper cites Sentence-level Prompts Benefit Composed Image Retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Sentence-level Prompts Benefit Composed Image Retrieval

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.796615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.796615Z digest=sha256:96bc5cda80bdbcbdc98de0a176661b53064fb1bf037c0218bf1b4dac55e3b521

Observation fbb65504-4c99-47e3-bd34-736153cc65b2 · outbound

This paper cites Image Search with Text Feedback by Additive Attention Compositional Learning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Image Search with Text Feedback by Additive Attention Compositional Learning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:07:41.386495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:35.903005Z digest=sha256:fbbc7d8377c38ea556e2f20487cc48631c986de841e508ec749bdde31e4b94c4

Observation 30f596e3-4083-4d51-b197-01f032f122e8 · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Image retrieval on real-life images with pre-trained vision-and-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.986269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.986269Z digest=sha256:7f6dabe90c6682cd1ae7438a4ee67d76b662343ff88c0d18ed53ffc0550b6ed2

Observation 2f231df7-737b-47a7-aa08-b5b9b016f965 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented clip-based features.ACM Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Composed image retrieval using contrastive learning and task-oriented clip-based features.ACM Trans

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T14:07:41.240420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.053819Z digest=sha256:1182708890ee491c01bfe9a1f72ce06a860da3e16724d95a260dab5b407e9ce3

Observation 7dd376e6-6da1-4e3d-a81c-4eabcdadfd49 · outbound

This paper cites Fine- grained textual inversion network for zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Fine- grained textual inversion network for zero-shot composed image retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:36.160343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:36.160343Z digest=sha256:9a32b1985b218afe37a8a67ccd77cb061d41ec00729c89dfde3396d5aaa4c443

Observation 9270e1f6-a011-4e8f-9a48-b3e09ed9b047 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:07:42.601773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.254019Z digest=sha256:4a0a286343fb80cb15aec5e8ba53f875762fb884ededcee90b61cbca71a06bec

Observation 94655246-9d83-41c3-8d1f-951209eb33a1 · outbound

This paper cites MLLM-I2W: harnessing multimodal large language model for zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MLLM-I2W: harnessing multimodal large language model for zero-shot composed image retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.724543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.416601Z digest=sha256:fc9bbbc0dbb3fcd3143ffe1b8e98b9e2fad96566ef223906380484495fa88781

Observation 8cd08409-8037-41ba-8f94-b7fa225194d5 · outbound

This paper cites Kankanhalli.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Kankanhalli

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.508680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.510398Z digest=sha256:df95d937bf8b6f61fa3912ace10272bc62a7126c989760b98863e3bb286bf8f8

Observation 7c16a5f1-68a2-4fd6-9e75-b18ac824a136 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Learning transferable visual models from natural language supervi- sion

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.274574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.588283Z digest=sha256:d07404a966100292307985678f5f44805a4addb88568d40d53df0b36a1adbecf

Observation eb65907c-f8e9-45b5-bd4e-372cc5f342da · outbound

This paper cites an unresolved cited work.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:47.063818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.710310Z digest=sha256:dbee2ec2bec765bf1200f1d213c00c8c88531ae3e807556b7cc1ce2bfc2a238f

Observation d28659d0-78ad-4ad8-86c1-077abbf24ef5 · outbound

This paper cites Seeing what you miss: Vision-language pre-training with semantic completion learning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Seeing what you miss: Vision-language pre-training with semantic completion learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:36.788327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:36.788327Z digest=sha256:7cfe519f90f625569886bcfb16a3e55696bb30281a63e5072a2731ec472c4e5b

Observation a939717c-1883-4fa2-99d7-4d71ecaf91b4 · outbound

This paper cites Global and Local Semantic Completion Learning for Vision-Language Pre-training.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Global and Local Semantic Completion Learning for Vision-Language Pre-training

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:07:41.095133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.870380Z digest=sha256:b2d12e7b5e977cfcc7d405180a0685155724c76dc39a971b04d68785c1e50c3b

Observation 22b2247f-6e88-4a19-8f53-2fe7d56eb5dc · outbound

This paper cites Zero-shot composed text- image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Zero-shot composed text- image retrieval

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:46.822771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:36.972948Z digest=sha256:b2089a8202c142e55b887ff5357f75fac19a69df98efaf7277ebb232e07db1d7

Observation 7159ee2e-b43d-4a3d-a704-81725a921591 · outbound

This paper cites Compositional Image Retrieval via Instruction-Aware Contrastive Learning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.059408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.059408Z digest=sha256:a6e9fe9adad35317da312517263036bb84db0803f034d62189f9b807663d1c08

Observation 6cf8b24b-55bf-48d4-a2c2-8e94cd17290c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval LLaMA: Open and Efficient Foundation Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.160966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.160966Z digest=sha256:ca0aeb1ff800aa539aee086f19fe2ae711c545038da5f6176d5a6cec3c8d7b4d

Observation 2f83fe13-f59a-4027-9956-1c6ef3d9215c · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.240894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.240894Z digest=sha256:b52f8e8684ea50bc219e58fc45b48717f5bd686c2771133a27439259bdedd296

Observation c023c4a0-12fb-4a6f-afc6-8b68a27e560a · outbound

This paper cites Visual instruction tuning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:46.516968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:37.359497Z digest=sha256:e8204a0c92ae27391ebfeadd09bb6c89ccf2ec299b9fa0bb3d5f7c164bc576f4

Observation 4faaafec-6882-45dc-89c9-338a09a643c9 · outbound

This paper cites SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.487699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.487699Z digest=sha256:2c7aacc82c61e092849a33888f46731942c2e9841e9cba8335c88e1d4ee20359

Observation b74f79d4-0c60-463c-8822-0f8326ab012f · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Grounding language models to images for multimodal inputs and outputs

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:46.295587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:37.592129Z digest=sha256:a76856cd502e34157cc61a10cbaa8f7e741c60a02cfc69b566c2cd41c565b0d1

Observation 7b1ae7af-ba40-4db3-b761-89f4dc576f85 · outbound

This paper cites an unresolved cited work.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:46.003833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:37.696524Z digest=sha256:03154a86d0937c1f7826320c3b1c4d8fe916345cd7c61c43195af7b87d1673f2

Observation 11038ea0-4bce-4a62-bb5f-b711dbd1a21b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.755205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.755205Z digest=sha256:1b2e835cec9975c337d9ec7655d53b894ec0c1fb5981a1ffdbe90aa38d708eda

Observation 61e154bc-4477-4ca2-a4e9-eed858a478dc · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Representation Learning with Contrastive Predictive Coding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.919492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.919492Z digest=sha256:6e4eee67c48f2defcf46dd108ef8b509dd8afb0107b90f980c78e87dd0f75c42

Observation 84835a9a-54c1-4b11-9ecc-8ea290504b13 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.813610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.813610Z digest=sha256:13aecdf0c340e2bc31497150e240430d40a2c3d5bd64d0331c2dad6bf1d570c6

Observation 8cf5a0ca-7cae-4890-9b59-9ebab07e4ae0 · outbound

This paper cites Weighted gaussian loss based hamming hashing.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Weighted gaussian loss based hamming hashing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.165615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.165615Z digest=sha256:3438348fed50f519b4aac335e2631b5e6883e7bcd8a571e49321ca35c67160ac

Observation 020557ce-a9b2-4711-bb62-d7b1292c9eea · outbound

This paper cites Unsupervised hashing with semantic concept mining.Proc.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unsupervised hashing with semantic concept mining.Proc

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T14:07:40.884160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:38.080713Z digest=sha256:97e377bff661adea8af3ad79f2e1efb45fc22e7f6277ce179c637fc9b3686678

Observation 1b6d306c-3e34-41eb-aed3-f31e3e77bd37 · outbound

This paper cites Unsupervised cross-modal hashing with modality-interaction.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unsupervised cross-modal hashing with modality-interaction.IEEE Trans

Reference 34

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:42.189094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:38.379544Z digest=sha256:f30c220b7ac7d96b68ea3d7835fcba0c49fdc84ad5413b6d280b6b9063af1478

Observation d416cac1-f94a-4443-bf6e-f460ef08e3ee · outbound

This paper cites Partial-softmax loss based deep hashing.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Partial-softmax loss based deep hashing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.290845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.290845Z digest=sha256:29b09f789305f948cd6a7506aad44c1a57b276a4624fca2e3bed0d4c17314919

Observation c7699576-e266-4c04-9ef3-62c61a9746f4 · outbound

This paper cites Unsupervised cross-modal hashing via semantic text mining.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unsupervised cross-modal hashing via semantic text mining.IEEE Trans

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.566628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.566628Z digest=sha256:923e58a9a44826cb91945f465f2990a49c2c2e8a3a1c578cd089677d4db630b8

Observation af33825f-e7e6-4eb9-abe4-0863552461e6 · outbound

This paper cites Deep cross-modal proxy hashing.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Deep cross-modal proxy hashing.IEEE Trans

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.471242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.471242Z digest=sha256:dd3c4340e572476152b5e9f41f28b304fa0c22f7b428f289bbbbe79f57e2d9eb

Observation a962b845-f363-4556-9c1b-11339166c34e · outbound

This paper cites Knowledge-enhanced dual-stream zero- shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Knowledge-enhanced dual-stream zero- shot composed image retrieval

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:45.438771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:38.893398Z digest=sha256:56b2774acc76f69f031c4029ee3d39fae41af628a4d47b95589789694b1780d6

Observation 4abf15b5-e33a-4f23-a3ba-adcf8076cb5d · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Zero-shot composed image retrieval with textual inversion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:45.689940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:38.676505Z digest=sha256:2aa5d68014174b5a73ef3289e549e769cfdb60c0e90e01ead86de1bc04695333

Observation 9fbf06e4-991f-4bd4-a342-b94d9b5b34cd · outbound

This paper cites Target-guided composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Target-guided composed image retrieval

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:45.186347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:39.258733Z digest=sha256:7e8b1f877a6acb93cfccede6413bbb93cab894bd5ee354b67b32dcbb017f738a

Observation d8838646-f428-4d69-b99a-625bd9726bce · outbound

This paper cites Enhance composed image retrieval via multi-level collaborative localization and semantic activeness perception.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Enhance composed image retrieval via multi-level collaborative localization and semantic activeness perception

Reference 43

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:41.699024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:39.075665Z digest=sha256:da7b5cec611bb1bf7579e454fe46f21c52e6b7d6a5ea79a2327e9bbb7c4541f3

Observation 4d0facfa-3bef-45b5-a616-90a4e93fa75f · outbound

This paper cites Data roaming and quality assessment for composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Data roaming and quality assessment for composed image retrieval

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:07:39.516160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.516160Z digest=sha256:2d3d804585917404ae3ccc04b7dbe9cce84e6c197d9a009a87ae4029f418c276

Observation 3cc6934b-9fde-4ca2-b134-d7dec363c68b · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:39.618249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.618249Z digest=sha256:4c7046ea5312d94fe118584ad1d48f200309a48b9c2a1cf114436f2b0ae566ad

Observation 3d1f6201-d720-47ac-bb11-baf71655fb1f · outbound

This paper cites Self- training boosted multi-factor matching network for composed image retrieval.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Self- training boosted multi-factor matching network for composed image retrieval.IEEE Trans

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:44.969022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:39.359562Z digest=sha256:ca5ed53f7223a7f2a65e8372ca404ba919b55a0d4dd3939c67ed64f821cebf5d

Observation ff304f35-27d3-410d-b920-5ca08d8da2b6 · outbound

This paper cites Context- i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Context- i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:39.449109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.449109Z digest=sha256:728eeed92e7ba4fbb2708c233d05a7843e43b6c2e0dc4954ad78b7b24f238b0b

Observation ca4a488a-0f0f-4da9-bf66-635be1128220 · outbound

This paper cites Fashion IQ: A new dataset towards retrieving im- ages by natural language feedback.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Fashion IQ: A new dataset towards retrieving im- ages by natural language feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:39.932800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.932800Z digest=sha256:1c9c440a3e157cb2309baf2475908cf489b67744cb5ab0aaeef1f5a740098b84

Observation b2fb9b68-3761-4a0e-b362-2c2ddd5034b5 · outbound

This paper cites A corpus for reasoning about natural language grounded in photographs.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval A corpus for reasoning about natural language grounded in photographs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.096640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.096640Z digest=sha256:2fb5f5eaa46717a6ffa3020126fa04df0eb1563a2ef370898f2251d45c0222da

Observation 4129a7a0-e2b1-42e7-b0d3-67e399357e06 · outbound

This paper cites Mini-batch optimization of contrastive loss.Transactions on Machine Learning Research, 2024.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Mini-batch optimization of contrastive loss.Transactions on Machine Learning Research, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:44.691636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:39.674668Z digest=sha256:1c1dfa6ed9b0eebafe2be66aef380671ae67aba604346793ff7b814e42bdc09e

Observation aa12343d-fb2d-48de-8a7a-c69161ef55b3 · outbound

This paper cites Bridging mini-batch and asymptotic analysis in contrastive learning: From infoNCE to kernel-based losses.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Bridging mini-batch and asymptotic analysis in contrastive learning: From infoNCE to kernel-based losses

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:44.425859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:39.786720Z digest=sha256:149ffd4681176e31701a074f3981b0f68d83865a3a10b1a03a37cf700f13206f

Observation 60c97a70-de33-4cc0-9da2-9ccaf7a8c04e · outbound

This paper cites an unresolved cited work.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:44.117488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:39.862668Z digest=sha256:0a62618103bdb49a09128829a81ac9098e8f452ba07fa064598a0f9f0b7cfe3e

Observation ef43d3dc-2c58-44b9-951a-117d568cacea · outbound

This paper cites Language- only efficient training of zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Language- only efficient training of zero-shot composed image retrieval

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.428606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.428606Z digest=sha256:c201a7a61aefd7c9a123d4574aba07f4b78330ddc8f573332a518338477c5c41

Observation 9cc13558-3ba6-494a-8a96-f7ec2024f060 · outbound

This paper cites iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.499510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.499510Z digest=sha256:5cbc2f1eb2d7cd7f26b42a36a610a4e2e2f9aff03ba1532ed1384ff83d9d5201

Observation 60ccd460-c987-4c18-b7f0-ae258f6f9267 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.153180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.153180Z digest=sha256:085e46563d38a99881d8632ec3fbde796d385e75de56a8477b760821e761de82

Observation 7a072da1-1de3-4991-afd7-b8a93b41041c · outbound

This paper cites Bernstein, Alexander C.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Bernstein, Alexander C

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.228472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.228472Z digest=sha256:2ca2f9f1c0cd153f1e453934c24df677ac4b8d4c054b89283b9fdd87c19eb8d0

Observation 1891be75-4b69-4f6a-8173-5eb5de5d6d3a · outbound

This paper cites Decoupled weight decay regularization.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Decoupled weight decay regularization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.384950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.384950Z digest=sha256:cbb8e6c14d467c6363001028eee97bb8c0ca295df422844e5bc2ac13c815dd8f

Observation 1f139446-d555-4488-8c72-cb646f2a48bc · outbound

This paper cites Vision-by- language for training-free compositional image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Vision-by- language for training-free compositional image retrieval

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:43.833760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:40.575614Z digest=sha256:f05d0accdbc9fcdb41eaa6ee7c6fdb5ff6b6e49b2a665fdc51870ae55ae4a790

Observation 95c7083a-1aa8-49db-9d8c-7bd100e3bc53 · outbound

This paper cites Effective conditioned and composed image retrieval combining clip-based features.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Effective conditioned and composed image retrieval combining clip-based features

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.687358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.687358Z digest=sha256:afa4e8b4a7a1b1589ddc14d1373ee6c706483932f4d637452948de7d8e840b45

Observation 605c0739-24a8-4707-8b81-9bf4aab364f2 · outbound

This paper cites cap1" and image 2 with the caption.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval cap1" and image 2 with the caption

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:43.504874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:07:40.730204Z digest=sha256:9c012dfc199d42f10d115ef1c03eafca257220eee1a85fb6c4c5491ed42c05f3

Observation 05e1be0d-d4c5-4bda-a1e3-ce88948d567a · outbound

This paper cites URL https://doi.org/10.1109/ICCV51070.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval URL https://doi.org/10.1109/ICCV51070

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.798145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.798145Z digest=sha256:e73dbbab4a2f6fab047936af86b34e4a18b6e1384bbc7d062268cec93c16327c

Observation bbf5ee11-3fd1-4061-a73d-4b5b733e86cb · outbound

This paper cites URL https://doi.org/10.1145/3626772.3657727.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval URL https://doi.org/10.1145/3626772.3657727

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.458335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.458335Z digest=sha256:5e7e351c9b4e7b3a9371af3a5b98fd426a83076b27f4220ad3132574c674506a

Pith citing papers

Observation e452ef4b-739f-4f84-ad68-a889a2b58f50 · inbound

SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding cites this paper.

SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.770726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.770726Z digest=sha256:6eab4978d001b43fd2c569074b93a28f6dc24a52d1e9e3b679f414788f3664f4

Observation 9b1c5e39-0d5b-436d-98fc-ec0b8f51f774 · inbound

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories cites this paper.

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T01:02:35.070287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:02:35.070287Z digest=sha256:5b442d6c4fc5c61c767c277cf187ddea2e931e7a390e4605085932c2e4072ce9

Observation 16a6f781-dec7-417d-ae33-b5f9e59733cb · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:1345244d3ac7e2ce75ab7d7ba6c764cafc3f73d497eb43a31c0e6a15bd288544

Observation 922148f1-64cc-45c3-8c8f-05fe704e3ba4 · inbound

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval cites this paper.

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.257851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T06:40:58.445504Z digest=sha256:ddb0e077e40b97d40bd43553cef3c5979af298cf1227007a65d0d5a42665ad68

Observation 49782fa1-ea49-4103-a5ea-995e68bdb437 · inbound

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval cites this paper.

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:56.878699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:10:54.790735Z digest=sha256:592a325fb443d84a15879758985c71c1f50a2673d7ae5fc8139b0124d07cbdbf

Observation 6d51e281-51a6-4bac-bd03-c5e8505fdf78 · inbound

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval cites this paper.

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:28:19.474650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:28:19.474650Z digest=sha256:66c82b39415556e76dbc4316729b8fae1adad304e3f0ad8f2632dadce6b36390

Observation de0291f8-879f-4383-a4d6-69e97783b87b · inbound

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism cites this paper.

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:41.145002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T05:52:20.768497Z digest=sha256:63091949a3b280221d7b729fb3e89bec75e609140625713a8c9bf4550995e077