Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:47.304539Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2412.11087.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:47.304539Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T19:31:53.371412Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T22:50:50.048388Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5fb24540-56e8-4360-9f8f-e6986dcef930 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f5b0d6-ec8b-466c-a352-abc8c34e8974 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727a0145-d411-48a7-ad65-39c497ca8cbf · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Sentence-level Prompts Benefit Composed Image Retrieval
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064b6ace-4224-4561-be2d-2e86603ad21f · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd9be402-77fe-4010-b660-7eeeec471912 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval L.; Berg, A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ce4b15a-ff37-4188-a439-13b8ca563391 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea39360f-46e6-4b14-b7ae-72b02ed00b32 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9249f51c-0d35-48e9-a670-ef8862423a9a · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e8164d-392a-4a45-8e4b-4b200b236fe0 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 631d5c9e-d0aa-439e-985b-88480d072afb · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4973a0fb-8a17-47c5-bdf0-48ff6745052e · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 27f9bddc-f97c-4171-8b23-0714a9ce819d · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1974dd6-5e0e-46f7-afe5-c803db7f9efe · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b237d09-9245-4d6b-80fb-235cab2dcb04 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fce0905e-7f11-4e6f-9f28-6774c2fd175e · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef0a1cfb-6a81-498f-aaa9-de20bb8b4924 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c8290c0-af88-45f2-abd0-a459b4ecd52a · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c176027-f57b-43fe-9489-4e11646f569c · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc39824b-a43e-44aa-8c39-2eeb3f6eaa6f · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509bffd9-bbe0-425d-a872-31402c8ee968 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval A.; and Manning, C
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab740d2-3bb6-4637-9ecb-e65e974755ef · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b0799d0-d39e-4c7e-8208-1ab26c86643a · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Vision-by-Language for Training-Free Compositional Image Retrieval
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82718db-b4c4-4818-91ec-5d3d25385df1 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d86ae43-c7fc-479b-bc15-e9d1593981ae · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Adam: A Method for Stochastic Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f68c5f-8c88-4ee8-960a-f9b9dc32b077 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0132dc7e-1cec-41a4-9c4e-fdfcc16c05ad · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Data Roaming and Quality Assessment for Composed Image Retrieval
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f869d3-b441-4c75-9e42-c5642a8396c0 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65b9e6d-9794-4c1b-8b79-7825e177d0e4 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041e9df5-d825-44df-8817-753de58a5356 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 894eadbe-dadf-4997-9037-3ca42f8f56fe · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Evaluating Object Hallucination in Large Vision-Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09255caa-b43c-4572-b1b4-756ea2c990c7 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Improved Baselines with Visual Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db129af9-232b-421b-a51c-ffc2462569ea · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval MMBench: Is Your Multi-modal Model an All-around Player?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6279a2-7dbd-4356-a9fd-1a300c492206 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0d246f28-a1f9-487b-a393-eb2c7b077dff · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 822d1bec-c915-4244-88cb-fa5d1786eab1 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2beb4b4a-b894-4e0e-b389-5f4ab2c24db5 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Fine-Tuning LLaMA for Multi-Stage Text Retrieval
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a39774-fd91-441e-9753-a11741d1c624 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval K.; and Chakraborty, A
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbbff1d-dbc6-4e95-84c2-f9d74f256452 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval SGPT: GPT Sentence Embeddings for Semantic Search
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2ecf54-2ea2-4fca-9c19-b297e9ac1ba9 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Generative Representational Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 205da4d2-3dce-411c-8f3e-daf13d0f8dc2 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd48168-0ce0-4c6e-9570-987871562290 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95fa12b0-5b56-4198-9901-d81d658bcb4d · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b0aaf17e-3454-4555-849a-b062c4180004 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval G.; Malinowski, M.; Pascanu, R.; Battaglia, P.; and Lillicrap, T
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0c56815c-e770-4609-9eec-1287024e33ba · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac7e409-e237-4800-83db-7be3e5894c67 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval A Corpus for Reasoning About Natural Language Grounded in Photographs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa180c8-e150-48d4-ad96-8008474f3364 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Training-free Zero-shot Composed Image Retrieval with Local Concept Reranking
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3645044f-2070-4130-88db-9d6714d33017 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cb588802-6fb7-48b3-a0af-1a7328346ce8 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval LLaMA: Open and Efficient Foundation Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab887efd-101c-41a6-9015-109824a3b76f · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b799e0cf-2820-45d6-991e-a2dfde406b13 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Evaluation and Analysis of Hallucination in Large Vision-Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2309af-10e4-462b-9d65-9c5a6205a06b · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d3d41a-88a4-4690-937b-ebf4e679913a · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51ad0a63-dcc2-4c5f-83cc-1caffdad4311 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 804d1bc3-1f7f-4aec-9fbb-823628419a0f · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e473cf3-9d1f-45bf-a40b-2ad714ab883f · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c3ff6a2-6eba-485f-a0fa-a705fcdeb841 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5004d9f0-efa5-4d6c-956a-19b78d0446eb · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unified Vision and Language Prompt Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42aa9c8d-351b-4ef2-be75-b2fd0a4dc6b2 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Progressive Learning for Image Retrieval with Hybrid-Modality Queries
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc6c121-5f93-40bb-b94f-f1e32b4e0bc6 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval C.; and Liu, Z
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372f6699-56bc-4c03-8978-d3f62914bdb5 · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace799f6-7218-4536-8016-cce571eb8dbc · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d25233e3-11bc-4165-b1de-a7de402374fb · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval , " * write output.state after.block = add.period write newline
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22567bc8-5e40-4428-8c68-e7cac127861f · outbound
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval write newline
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 068c18a3-4b06-4d85-a674-58f6890816df · inbound
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.