Pith. sign in

Paper Citation Record · LEDGER

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

As of 12 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2412.11087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11087 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:47.304539Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:31:53.371412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:50:50.048388Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fb24540-56e8-4360-9f8f-e6986dcef930 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.072504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.072504Z digest=sha256:64c2b9adbf7999a35c50deb5f1e433ea60575aa4d175b9ad40e6d61a63f371ee

Observation a9f5b0d6-ec8b-466c-a352-abc8c34e8974 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.077372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.077372Z digest=sha256:7de755d41278db74444569cd8d0e7fad01a5fc8539dc8653261f97ca9e4f9cf1

Observation 727a0145-d411-48a7-ad65-39c497ca8cbf · outbound

This paper cites Sentence-level Prompts Benefit Composed Image Retrieval.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Sentence-level Prompts Benefit Composed Image Retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.081805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.081805Z digest=sha256:7181d2985ae539d2fc1bd5fa35f41c029c03cbbaad3c87988df2060fb0ad8745

Observation 064b6ace-4224-4561-be2d-2e86603ad21f · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.980959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.086242Z digest=sha256:7b9b08bce1e566993e26ce7dc2ea0998f45c426af520169707422be6eb717312

Observation dd9be402-77fe-4010-b660-7eeeec471912 · outbound

This paper cites L.; Berg, A.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval L.; Berg, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:47.969858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.089784Z digest=sha256:9224e6669b3c4d3e75e919e3ed64117f9b55c1c7d4a5cd81dc0fd51864018bf1

Observation 6ce4b15a-ff37-4188-a439-13b8ca563391 · outbound

This paper cites D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.093309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.093309Z digest=sha256:e8f5e15dc641c137a43ea5466e325a494844d67b1fc4e8f274d9f0b88d5d3466

Observation ea39360f-46e6-4b14-b7ae-72b02ed00b32 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.951530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.097963Z digest=sha256:61b3344606e2817a4c60f85c0be349ee618237bbda62f01b3decf28709f134ef

Observation 9249f51c-0d35-48e9-a670-ef8862423a9a · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.101264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.101264Z digest=sha256:404b4e4374fb32810fde1c54249e1420880c5a8e00a5960dc305971235d58692

Observation 92e8164d-392a-4a45-8e4b-4b200b236fe0 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.104874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.104874Z digest=sha256:3b4e65d53fe2abdd15ca5d47879ac0829b7aa1dd59d85e0278653952697efe31

Observation 631d5c9e-d0aa-439e-985b-88480d072afb · outbound

This paper cites ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.108405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.108405Z digest=sha256:fb114d0b515043fe41c7261806ef4886b9a7774118eb7d1f6ba5c17b65826fc1

Observation 4973a0fb-8a17-47c5-bdf0-48ff6745052e · outbound

This paper cites S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:47.931167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.112664Z digest=sha256:a15ebfa46392da2db74f58d3d76f9eea97f733254b4e856590cf78149d9336c6

Observation 27f9bddc-f97c-4171-8b23-0714a9ce819d · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.116387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.116387Z digest=sha256:763b48663c0cf00f3e5e8d242ab95a75829c838d47196371c8a7b7cd82c6edfe

Observation c1974dd6-5e0e-46f7-afe5-c803db7f9efe · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.917791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.120566Z digest=sha256:dd15561d4ec4a12457f13bbdaf40db5b493537b98b20b16c2d45b4fcb2fce40f

Observation 5b237d09-9245-4d6b-80fb-235cab2dcb04 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.906077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.124246Z digest=sha256:985d46a683451d7e1090f256740e77d0504e3e5eb4c1726536ffb788d3d5b65b

Observation fce0905e-7f11-4e6f-9f28-6774c2fd175e · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.895282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.128141Z digest=sha256:a50f4eb206c26e62d6f69679318a552d64280ebe76f88fc6171c3b6f2a02ae20

Observation ef0a1cfb-6a81-498f-aaa9-de20bb8b4924 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.884739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.131595Z digest=sha256:8451fea9762c267069ae75693b7e0ad1ea1611aed1eb282756231aa90dec4ded

Observation 8c8290c0-af88-45f2-abd0-a459b4ecd52a · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.873396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.135577Z digest=sha256:5adf80a34e909f8b1423ea443a8a46827d3a31aab773dcb44c7be54658cf6135

Observation 4c176027-f57b-43fe-9489-4e11646f569c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.139250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.139250Z digest=sha256:5c5ee8316315efc03a4644aab83e681ccf7da0b6a9ee0b761a96e4711bd4b377

Observation cc39824b-a43e-44aa-8c39-2eeb3f6eaa6f · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.142767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.142767Z digest=sha256:351ec79afac125022446636416eabf0ae127ac211ac35160813e9ac2d4db1a1a

Observation 509bffd9-bbe0-425d-a872-31402c8ee968 · outbound

This paper cites A.; and Manning, C.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval A.; and Manning, C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.145792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.145792Z digest=sha256:54f24aab8be44b9ffa1a630fc2d2dce6a1a150596cf383a900edfa184db27bd0

Observation dab740d2-3bb6-4637-9ecb-e65e974755ef · outbound

This paper cites A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.148781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.148781Z digest=sha256:24a355c6187e0884fb92f56dfd9d2386f98efb24a9c27a338be546e2be3f610f

Observation 0b0799d0-d39e-4c7e-8208-1ab26c86643a · outbound

This paper cites Vision-by-Language for Training-Free Compositional Image Retrieval.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Vision-by-Language for Training-Free Compositional Image Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.152427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.152427Z digest=sha256:fbe757046094bf3c08cd2a1258f6f9495c963876cb94b4ec19fefc0c7709fda5

Observation b82718db-b4c4-4818-91ec-5d3d25385df1 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.847330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.155913Z digest=sha256:b5661483385d54cd8e06e23df16b85d2d38437fa92bf930b4ff85b3e76adbb6b

Observation 7d86ae43-c7fc-479b-bc15-e9d1593981ae · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Adam: A Method for Stochastic Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.159571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.159571Z digest=sha256:7f61a6b2ec50c89b273ef4d1f1f60930a6654926bcdd1c5fae8cccd86e096660

Observation 91f68c5f-8c88-4ee8-960a-f9b9dc32b077 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.835166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.164035Z digest=sha256:a852bc1173495f0d91c500afd25145a2bda196a870a892f162d6ebcb5c2cbd74

Observation 0132dc7e-1cec-41a4-9c4e-fdfcc16c05ad · outbound

This paper cites Data Roaming and Quality Assessment for Composed Image Retrieval.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Data Roaming and Quality Assessment for Composed Image Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.167708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.167708Z digest=sha256:e99e63558aa6443bb475526ce1b9d0cc5d85a161abc19b11f0f43127601d4081

Observation 69f869d3-b441-4c75-9e42-c5642a8396c0 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.171646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.171646Z digest=sha256:0613021e040eda483a8cffc190b860735316c6b7f5ff45b9859488c166effa6e

Observation a65b9e6d-9794-4c1b-8b79-7825e177d0e4 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.175513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.175513Z digest=sha256:83ff5015dd89f539ca3a090831dcfec176fc6d90de6736683f6978df3e12e505

Observation 041e9df5-d825-44df-8817-753de58a5356 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.179120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.179120Z digest=sha256:8ede6fc41257e46c2f651056f573634922b3a93e6e3090a536df072eee15eba2

Observation 894eadbe-dadf-4997-9037-3ca42f8f56fe · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Evaluating Object Hallucination in Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.183013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.183013Z digest=sha256:053d453ca045831f3e11dd5603189f18c64a3d2c3933dbcdcf9c63ecc5f5ff50

Observation 09255caa-b43c-4572-b1b4-756ea2c990c7 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Improved Baselines with Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.186912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.186912Z digest=sha256:0841b9124ba7430d61a2ed6d4ac0728f948dfa653bf3b9b7b41d098baf02e147

Observation db129af9-232b-421b-a51c-ffc2462569ea · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval MMBench: Is Your Multi-modal Model an All-around Player?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.190182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.190182Z digest=sha256:4ae7966abf9966a2c4f55ed07eb73fb0af6414632bec5780e537784627e5e86a

Observation 9c6279a2-7dbd-4356-a9fd-1a300c492206 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.807648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.193804Z digest=sha256:19eddeddfe8e268b8c11443e24729f5127315f2ac85718ce1e0053c96e1c2b5d

Observation 0d246f28-a1f9-487b-a393-eb2c7b077dff · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.795602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.196669Z digest=sha256:5a7a508f26f1cc3e832697dca42515dd649933d24f93bbb13df957512474f641

Observation 822d1bec-c915-4244-88cb-fa5d1786eab1 · outbound

This paper cites Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.199484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.199484Z digest=sha256:1adc88fae7d3c34b68b383ed1bfb7e10168c89fac2cee4f8617f7524379aa729

Observation 2beb4b4a-b894-4e0e-b389-5f4ab2c24db5 · outbound

This paper cites Fine-Tuning LLaMA for Multi-Stage Text Retrieval.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Fine-Tuning LLaMA for Multi-Stage Text Retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.202636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.202636Z digest=sha256:ada92a278b7a7d7f17816be41261642ffe75eb212f5dd9d56206784895bad9ff

Observation d5a39774-fd91-441e-9753-a11741d1c624 · outbound

This paper cites K.; and Chakraborty, A.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval K.; and Chakraborty, A

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.206294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.206294Z digest=sha256:f08bb28b6e6cd39723c914b2617274cd0ada6447189169876218744c9864d5f3

Observation 8cbbff1d-dbc6-4e95-84c2-f9d74f256452 · outbound

This paper cites SGPT: GPT Sentence Embeddings for Semantic Search.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval SGPT: GPT Sentence Embeddings for Semantic Search

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.209922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.209922Z digest=sha256:70ff75d43b80230ee2ecc4446ba2ffda393a7dc649d6f6e2d7352352fbfd82ed

Observation 7a2ecf54-2ea2-4fca-9c19-b297e9ac1ba9 · outbound

This paper cites Generative Representational Instruction Tuning.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Generative Representational Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.213998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.213998Z digest=sha256:d1f4a019c21952909ebd7de3a9a95fa56ca7e9bff4e0976ad76861c6b196135d

Observation 205da4d2-3dce-411c-8f3e-daf13d0f8dc2 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.217824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.217824Z digest=sha256:b2b5c74f52187592140c2e7e6ab0e58e6c95b8875e0fbc996167a9f174912887

Observation fdd48168-0ce0-4c6e-9570-987871562290 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.221844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.221844Z digest=sha256:613d2060f4e7bcb04d1a2c913964123cbce07ab1fbe84b2f69131ce0155cdef8

Observation 95fa12b0-5b56-4198-9901-d81d658bcb4d · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.763751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.225620Z digest=sha256:e8df42e7e8f16f16ea84fd89a863f2b5405a91674be38a3f1a9cb83a179ad543

Observation b0aaf17e-3454-4555-849a-b062c4180004 · outbound

This paper cites G.; Malinowski, M.; Pascanu, R.; Battaglia, P.; and Lillicrap, T.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval G.; Malinowski, M.; Pascanu, R.; Battaglia, P.; and Lillicrap, T

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:47.753157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.229206Z digest=sha256:e9e005fb625c73cd228cdb3e46a113dda67c23e8aa039848a6e0c9e5738ab1cc

Observation 0c56815c-e770-4609-9eec-1287024e33ba · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.232840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.232840Z digest=sha256:c6ba2334aff0f8d8904d81325433dbbe758a8be645ee564d9ab610907b384789

Observation 3ac7e409-e237-4800-83db-7be3e5894c67 · outbound

This paper cites A Corpus for Reasoning About Natural Language Grounded in Photographs.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval A Corpus for Reasoning About Natural Language Grounded in Photographs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.236367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.236367Z digest=sha256:d813ac0ddcdf8689009cd5534d6e421ab7dbc7dd0e86b959dad9031b1346f788

Observation 4fa180c8-e150-48d4-ad96-8008474f3364 · outbound

This paper cites Training-free Zero-shot Composed Image Retrieval with Local Concept Reranking.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Training-free Zero-shot Composed Image Retrieval with Local Concept Reranking

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.239606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.239606Z digest=sha256:98f38d81aee3f98b52dec76c2cf66e78c7f76b104a9557cea6ca96ac32efa784

Observation 3645044f-2070-4130-88db-9d6714d33017 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.740836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.243146Z digest=sha256:262b50939638958ec0bea64a0fb19df90064a5f470bc80eb96dd49f84df61e3f

Observation cb588802-6fb7-48b3-a0af-1a7328346ce8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.246222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.246222Z digest=sha256:05d4bf6858b0f820f56687fbc2a6da518739422a6d6eadce26923594c3539460

Observation ab887efd-101c-41a6-9015-109824a3b76f · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.727791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.250381Z digest=sha256:06e90df7c9b051b67b756449904c7018de63e8cae4ee439330e7610e38cdfef3

Observation b799e0cf-2820-45d6-991e-a2dfde406b13 · outbound

This paper cites Evaluation and Analysis of Hallucination in Large Vision-Language Models.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Evaluation and Analysis of Hallucination in Large Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.254339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.254339Z digest=sha256:239f33b463c8c900938235842cd4f0819e42acbd7d8839d1d055690ddaf7f3eb

Observation 5b2309af-10e4-462b-9d65-9c5a6205a06b · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.258153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.258153Z digest=sha256:2baa9669e70330d15dd2c41552bca66ff20d99c23e04fd6017cfbf48a14fe193

Observation 60d3d41a-88a4-4690-937b-ebf4e679913a · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.709440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.261872Z digest=sha256:ed5cafb906feede69d0e0f8e14f83b9c96d7e5611dbb647e86fd48ecd598350f

Observation 51ad0a63-dcc2-4c5f-83cc-1caffdad4311 · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.698621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.265736Z digest=sha256:3805e3e309feb27902cebdf08a7d9c38f66499324f800610ec5a93ecd3ebf606

Observation 804d1bc3-1f7f-4aec-9fbb-823628419a0f · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.688103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.269897Z digest=sha256:d51e8ce8824a88f6cb58343ccf5f04b2e7fc0a5a5a33aac7934549289f1c1c98

Observation 6e473cf3-9d1f-45bf-a40b-2ad714ab883f · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.677071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.273631Z digest=sha256:9bc618ba6e6f5ef15ef28b8ad814d388941a1aaad6d1d71d30ac09f853c8d899

Observation 9c3ff6a2-6eba-485f-a0fa-a705fcdeb841 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.276893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.276893Z digest=sha256:cd5d787714ce2f01c9ea336232a84fe6c95776e1d1e38b7e7486084b88640119

Observation 5004d9f0-efa5-4d6c-956a-19b78d0446eb · outbound

This paper cites Unified Vision and Language Prompt Learning.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unified Vision and Language Prompt Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.281351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.281351Z digest=sha256:5d53c69bb448b505700ac47a555c7b88ac0caf248c917475340230daae106776

Observation 42aa9c8d-351b-4ef2-be75-b2fd0a4dc6b2 · outbound

This paper cites Progressive Learning for Image Retrieval with Hybrid-Modality Queries.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Progressive Learning for Image Retrieval with Hybrid-Modality Queries

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.285162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.285162Z digest=sha256:e5459f38fdb65678aced30553f9f7c304912e92f0a3975b735dd3d1c3ff8041a

Observation ddc6c121-5f93-40bb-b94f-f1e32b4e0bc6 · outbound

This paper cites C.; and Liu, Z.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval C.; and Liu, Z

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.288651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.288651Z digest=sha256:7c77f8dffcff86ba269877d1688fcb8be17f57ce97ec90e62094612eb8456f0e

Observation 372f6699-56bc-4c03-8978-d3f62914bdb5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.291989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.291989Z digest=sha256:9caeb4b1496933da5f76c0e4d5abac927d3213c8d25572e8667c958c06c6d3dd

Observation ace799f6-7218-4536-8016-cce571eb8dbc · outbound

This paper cites an unresolved cited work.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:47.659404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T15:20:47.296031Z digest=sha256:42a74d58163098447cbb04a14b3936cc01de7ca526237e2f0273d04ba6800f51

Observation d25233e3-11bc-4165-b1de-a7de402374fb · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval , " * write output.state after.block = add.period write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.299661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.299661Z digest=sha256:72efdc8d7ac82504d33c4461060b2f453b7fd3f2e4b9915db7582a8200d509ca

Observation 22567bc8-5e40-4428-8c68-e7cac127861f · outbound

This paper cites write newline.

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval write newline

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:47.304539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:47.304539Z digest=sha256:013d11196eb2a9bc5d72caf8a3318597b127d68a02757ddca78cef3877706934

Pith citing papers

Observation 068c18a3-4b06-4d85-a674-58f6890816df · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.051194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:9a11b7e7fd4413cb7443a78bfe1481dff225edbbc45864a2de62ebcfe2d609f9