Pith. sign in

Paper Citation Record · LEDGER

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

As of 6 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2606.07032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07032 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:44:36.059582Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:12:46.295455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T20:38:56.147115Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact15
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5eb7b69d-7cea-4f65-b4b5-bf98e3f4e976 · outbound

This paper cites Heterogeneous feature fusion and cross-modal alignment for composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Heterogeneous feature fusion and cross-modal alignment for composed image retrieval,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:c7e125c96b896ab55bb614e057e0a6b24ea75611f4581abca9fb10f08ded8421

Observation 1f39d5ed-8f5d-46c4-a85f-b6547350f427 · outbound

This paper cites Effective conditioned and composed image retrieval combining clip-based features,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Effective conditioned and composed image retrieval combining clip-based features,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:080821a409825f0a5ac2495631635de111dbed8e4700f56308212ea9f2880df8

Observation 446c99c6-ef44-4505-83ae-9550143215d7 · outbound

This paper cites Target-guided composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Target-guided composed image retrieval,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:c27d621df22ba11ceb3eb5bd137dcfec0756eb290956d901e118772ec93e945f

Observation f86d1e4c-ec94-49f9-8b56-91f138383314 · outbound

This paper cites Self-training boosted multi-factor matching network for composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Self-training boosted multi-factor matching network for composed image retrieval,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:57e0faf960e300c1c701715e8bacb4ef35749d7d18903afaf34547e654817fa5

Observation 098c753e-95f1-451a-aa96-33012075331a · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Pic2word: Mapping pictures to words for zero-shot composed image retrieval,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:8f82d9eb825af9386cba554c1c69aad15ba0c2f703db2bcd443ec44528fc8d86

Observation c698c4d4-8b46-4981-baea-43f4377dd961 · outbound

This paper cites Zero-Shot Composed Image Retrieval with Textual Inversion.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Zero-Shot Composed Image Retrieval with Textual Inversion

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.005749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:503a86448a024f36ec4fe63212dff7f1b731eb395396074fec55fb83bfa23cae

Observation 811d0319-d9fc-425e-a80b-aeedb1bce59a · outbound

This paper cites Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.017796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:f00cf69ae0d456bd85e644ece23589b2f849695c45ada6caa246b4c3a4e5ce16

Observation 61df25ed-969e-491e-b733-e76afef4b15d · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Image retrieval on real-life images with pre-trained vision-and-language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:520211a27dcef47ae83edbd23e07cd64538584debcbfc914e677140d55cd1a26

Observation 6c004942-a749-468e-b7d6-9f9e9f5b7529 · outbound

This paper cites Genecis: A benchmark for general con- ditional image similarity,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Genecis: A benchmark for general con- ditional image similarity,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:d45222df900095fb90739c187d520a26039a67f865a5b829aa61873de96b7989

Observation 3fbfeff2-130c-4258-82aa-563179dc996a · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Fashion iq: A new dataset towards retrieving images by natural language feedback,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:0cbe5cdafb702ba3f79a88dda729cc0f4bb61088893377e38180a167a85d5f08

Observation 6f368c76-baa5-4926-a2e3-c0cd4e035b9e · outbound

This paper cites Data roaming and quality assessment for composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Data roaming and quality assessment for composed image retrieval,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:a1741ee4c9d7ae68fc4b2666aeb92948c1081892034cd893cc6b4b0a5c8622db

Observation 13eba261-1956-4688-82c0-79d8b4230741 · outbound

This paper cites Dual compositional learning in interactive image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Dual compositional learning in interactive image retrieval,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:048040baffaba27b3596ef021cbdf34ba3026b61e08e708d7d68f2d15d041812

Observation 8e658440-3cd7-4d36-a488-ac70a1765f23 · outbound

This paper cites Visual compositional learning for human-object interaction detection,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Visual compositional learning for human-object interaction detection,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:d43906f65acb6ad883d3af4d0f333c4cee4393b12259f0be1f86021387b89d64

Observation 35531127-c7e7-43c0-bc4c-5fcfe1a569be · outbound

This paper cites Leveraging large vision-language model as user intent-aware encoder for composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Leveraging large vision-language model as user intent-aware encoder for composed image retrieval,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:94daa5c6a00c2a80ad72f0c72cff0c186e3b41da7682182fee209310615c97ff

Observation 80abc98f-ad59-43c9-953e-697fe69a88f7 · outbound

This paper cites Ccin: Compositional conflict identification and neutralization for composed image retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Ccin: Compositional conflict identification and neutralization for composed image retrieval

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:0a0755a2deb4143bdec97b001c9ef93de52b7bd3b3b3c67b8583866498382d8e

Observation 223b5389-6303-4ff6-87a3-d48cbf8a8c98 · outbound

This paper cites Covr: Learning composed video retrieval from web video captions,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Covr: Learning composed video retrieval from web video captions,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:04154c6ca53769bdd55ef0145724a1d742da0102a530ba1bcf3f48bc7bfe2954

Observation 9a300e96-ae04-4fdc-bbea-44ed1606ac90 · outbound

This paper cites Bi-directional training for composed image retrieval via text prompt learning,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Bi-directional training for composed image retrieval via text prompt learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:d673cc3d9c8a20ce2b36d1c423dcd3a18317372da7935be3ff67e25d2049bd1a

Observation 4a57f7e5-dfe0-4d86-8bfc-d7cf4f10665d · outbound

This paper cites Sentence-level prompts benefit composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Sentence-level prompts benefit composed image retrieval,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:fac78feb8f32a09440ac8994f6e5322c62cddc3b632fc71ead4cb4eb2caf1333

Observation af264afd-b0bf-4239-a6dd-2b3daa670635 · outbound

This paper cites Dynamic bit-wise semantic transformer hashing for multi-modal retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Dynamic bit-wise semantic transformer hashing for multi-modal retrieval,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:0aaa60a24bcf4c98e9326796ab384bea9b5c686f441a1081dd45695e162c5488

Observation c70e21b3-44d3-4e86-9122-ce985088fa6e · outbound

This paper cites Fashion retrieval via graph reasoning networks on a similarity pyramid,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Fashion retrieval via graph reasoning networks on a similarity pyramid,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:5603e628706a6f6c365c9079d87197a5993cd4e3502b397dd2b8c4c925096a6c

Observation d0606636-f7bc-4fbf-924b-cb049d0fcb3e · outbound

This paper cites Multilin- gual text-to-image person retrieval via bidirectional relation reasoning and aligning,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Multilin- gual text-to-image person retrieval via bidirectional relation reasoning and aligning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:e0fdc7520c2f2ada050aa3c1cb2d0fb122a2c4f6aa8330c877661610d84fbeb1

Observation 45f5d826-d653-4e7f-a43c-75a1eb053701 · outbound

This paper cites A corpus of natural language for visual reasoning,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets A corpus of natural language for visual reasoning,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:7c03bf64d347fbdb88897e5b812f8c7b6b47fda86e3c28ebff8c85c709c0ef4b

Observation 8b87c6a8-a8eb-46b3-a188-ca79681bc9e6 · outbound

This paper cites Microsoft coco: Common objects in context,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Microsoft coco: Common objects in context,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:5b82e6741abe1f181391d224d14abd46c153774d99636abbe594cc04e0e78d06

Observation 58dcc51d-ebf8-481b-9c47-3f1d29371f49 · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Imagenet: A large-scale hierarchical image database,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:2472a29444121578fd8366908f039adeb291a019ced78f9df0037839c507d686

Observation 02f864c4-5c23-4855-af8d-c9afd093d007 · outbound

This paper cites “this is my unicorn, fluffy.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets “this is my unicorn, fluffy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:18fd7856a5f5add4b352287860e041c70e3642512ead033d8b78150e40d63249

Observation 90f49ea8-c0b9-492f-879b-1ade5371772b · outbound

This paper cites Language-only Efficient Training of Zero-shot Composed Image Retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Language-only Efficient Training of Zero-shot Composed Image Retrieval

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:08.989476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:386daad0094073a611ff303d73c7c3690fbde5bacee37da10d84ca490e2d2a2f

Observation cac05558-6818-457b-b095-d2c2dfe52d60 · outbound

This paper cites Vision-by-Language for Training-Free Compositional Image Retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Vision-by-Language for Training-Free Compositional Image Retrieval

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.025375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:edad1d5b8893742ed067d62bb74de85014940457d0ad1e70df9a0fb30c57eb47

Observation 40b226c5-07cf-4ef0-93f1-ad464089aa88 · outbound

This paper cites Ldre: Llm-based divergent reasoning and ensemble for zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Ldre: Llm-based divergent reasoning and ensemble for zero-shot composed image retrieval,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:867fa92208e28aa6c731415c77f8587b66c7b96bdf12ccf9089437b89703879b

Observation 5e28f45c-55e3-42c7-8413-2a545eb79560 · outbound

This paper cites Semantic editing increment benefits zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Semantic editing increment benefits zero-shot composed image retrieval,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:78183d8d801e24545df1468a10df85a650b2af051ff1771ef710f2a2d39c83bf

Observation 8e8c9a2e-76f0-4497-9393-7f2a70470a7e · outbound

This paper cites Mllm-i2w: Harnessing multimodal large language model for zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Mllm-i2w: Harnessing multimodal large language model for zero-shot composed image retrieval,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:57c2635ce6e6a628b5de2275542aae52470b2c18f2c9d4b3af673431651b6fbb

Observation 5d391179-443e-4514-9012-8a18f3bcd8ae · outbound

This paper cites Generative zero-shot composed image retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Generative zero-shot composed image retrieval

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:6633ee2ca098fb84766afdc3743c0d2d66f6dc472ae6139984af4fa3a85ebe77

Observation b1474ca5-80ca-45b4-b404-be260001f46e · outbound

This paper cites Active supervised cross- modal retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Active supervised cross- modal retrieval,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:d11d399910017a30bc3bfdec2032a61b3479c07b79c72d0bf9dbe5b798fe2a3e

Observation 67a2aa83-5855-40ed-820d-2001c897fb77 · outbound

This paper cites Prvr: Partially relevant video retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Prvr: Partially relevant video retrieval,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:4aec95e540279ddf9f362a465247b626a073a31de2e04e9a260ffc65b4f673f5

Observation bfca25d9-5ee4-400d-81d8-d4350352e56b · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Learning transferable visual models from natural language supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:5575e6f88081cd3b53a9f9eb7998b914cba79ba042b14396d841d2a0e9c28c8f

Observation 4c693f89-9606-40b9-93d4-5df3c99a1f6a · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Laion- 5b: An open large-scale dataset for training next generation image-text models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:d42812945a060f1995652ea08ccc04707f1a14a8f607dd2ccb4d9ee864c902f7

Observation 59eccd9a-9721-4058-90e5-90e42f438a10 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Datacomp: In search of the next generation of multimodal datasets,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:7e1098adced5ffed50f9eb7f01be3fa1729e9b48f972abde39771f3956a35add

Observation c60950f5-d858-4187-b3ff-ca3b937b1ef4 · outbound

This paper cites Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:52694ea19dc0afffbc3eda41d9507ce6a9e8602e8f23eb2f7068009ab0322f7e

Observation 4febbffb-4a82-474f-925e-42fa4de6a4a5 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:08.991831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:b72867857e1df9d97e78f0bee3d500e78e06d86e55434aea6fa3f4b49b90a61e

Observation 854a9759-a307-45cc-97f0-b1ff1b61f5eb · outbound

This paper cites Data Filtering Networks.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Data Filtering Networks

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:09.020161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:911004d4275b65005a8cc5418d86c2e2fa002bc7db1d603c63f6aafeaff712e2

Observation 4699784f-86a7-439d-a3f1-5f274009e3bb · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Composing text and image for image retrieval-an empirical odyssey,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:2f46a240a6d0e3882c942b400acf84173be71b14444b1e3d19eff8d13ac73bbd

Observation 1b5b1f11-ad27-49c3-b91b-328a77b0155c · outbound

This paper cites Image search with text feedback by visiolinguistic attention learning,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Image search with text feedback by visiolinguistic attention learning,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:3f15858d423112f4cb9dc5b3fd128a531e47d75f698c7df673199da2051230eb

Observation 4b8bcb31-f9e5-437c-a337-78f0d9f49ebf · outbound

This paper cites Image search with text feedback by deep hierarchical attention mutual information maximization,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Image search with text feedback by deep hierarchical attention mutual information maximization,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:02dca5a3ba390fee7e284f1800ef82bd103a4040fefc4c8fc45abdceabe267d7

Observation 339a7394-8da7-4230-b304-ac52ccb1f705 · outbound

This paper cites Cosmo: Content-style modulation for image retrieval with text feedback,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Cosmo: Content-style modulation for image retrieval with text feedback,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:e4a6627eb25e127d415f66170f18085084f4e247437e1e06faf401ba152ba3cd

Observation ca95bf01-3fb3-4801-bd50-208b661f02a8 · outbound

This paper cites Zero-shot Composed Text-Image Retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Zero-shot Composed Text-Image Retrieval

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:08.987164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:ff6e53b383e6db2e65162bd3907bd4d7e97e14ba1e9a9c22124ffd23076a9815

Observation 353a1d1d-8fc0-4e9d-a00d-9ac6ee0b83a5 · outbound

This paper cites isearle: Improving textual inversion for zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets isearle: Improving textual inversion for zero-shot composed image retrieval,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:f9ac25157829f4dd1a0b7e90f55bdf0ad8d14abc244c0bb2f462f2f75fdc60db

Observation ee94aa73-edf9-4c8a-ac25-b50617eeda2b · outbound

This paper cites Towards codebook-free deep probabilistic quantization for image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Towards codebook-free deep probabilistic quantization for image retrieval,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:769c1c2e90fa6fc14ca75e0450dcc8b84d25520dfb291b7ba799eb6f2b550c7f

Observation 7b6676c1-f11d-444d-94fa-b86d58417b1c · outbound

This paper cites Mixture of subspaces image representation and compact coding for large-scale image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Mixture of subspaces image representation and compact coding for large-scale image retrieval,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:0d10a36e3fb37d23eee7646c9dc8f4b88452145c14d6de2b80ac79913c8465af

Observation ed908753-9182-4225-8342-ceb476273c57 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:08.996267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:c992a2ae43ac4b2401131d454fd279f512c23e136aacd20ba702eef29a86af46

Observation d534b49c-fdb6-4a5c-9f93-29288e968081 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Uniter: Universal image-text representation learning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:de85aa1af7da8bc5d5891b11c7b62be5ab14f2ea69b1971de4fbb66218d5dbed

Observation 8daaf6b4-1826-4e49-bfb4-74f98aee3e18 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:09.000692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:96e4ef107770befe068ae361cd4edebc04e13fa4394f78b0a59cfbb2120bbd22

Observation 59e1d280-1f91-4d23-a280-b6030a469e67 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:32bc388100e55a179b8301661e1caad25f72a0e813a847f10722a723f51ff1b8

Observation d0c9f4f9-3bd0-4ea7-8314-066962fa14d6 · outbound

This paper cites an unresolved cited work.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:385e2974337273f9f130a13a05395f55ae7eb3b10a1a81489fc6f9f978a99e02

Observation 4ad469ad-cf4b-40aa-a0e3-ab7a4bf8d44e · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:3833b6a6921a6bbb24e66e6249088c18b9c6ecf56feaface97d4c7df95c8f519

Observation e1d4918e-fecc-46a0-9a86-32c4c830a764 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.010267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:4e18aaff71283df058c2bcfa7d30601669b5f62910dfb39a46b87fde5bc7e5fa

Observation 0641a4dc-e4b0-4809-9c1d-44519274eb43 · outbound

This paper cites Optimization of rank losses for image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Optimization of rank losses for image retrieval,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:4f252486824b0feafd748262e2e7cd287157c527e8305dc1234c30dca3723bd1

Observation e57909c8-9423-4715-aa00-b864e8c239d7 · outbound

This paper cites Attack as defense: Proactive adversarial multi-modal learning to evade retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Attack as defense: Proactive adversarial multi-modal learning to evade retrieval,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:1198b322a0bee562005789b2977b0dc8499cf325b66c018ac35e592f9b756483

Observation ede8a08f-34e0-428a-ab8d-4d63e2aee859 · outbound

This paper cites Attention is all you need,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Attention is all you need,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:735fe1cb163e26e051096f6f002573c8c92d4c1352ffc10a18ae83cef8a0fd50

Observation 894ce067-9aae-4741-b2de-00a27a7d88da · outbound

This paper cites Conditioned and composed image retrieval combining and partially fine-tuning clip-based features,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Conditioned and composed image retrieval combining and partially fine-tuning clip-based features,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:1fb44b45d05967c1dd20ff193cf936b0d0372e214345f2c68c9a156dc3ac46c1

Observation 7dd07608-5bad-41cd-ade2-fc06ea6d55f3 · outbound

This paper cites Fame-vil: Multi-tasking vision-language model for heterogeneous fashion tasks,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Fame-vil: Multi-tasking vision-language model for heterogeneous fashion tasks,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:463a3352ee468297563df81a081f3a12bc45433bf7ac2493f3991018177263a0

Observation e5899d17-ab65-4ff5-a9aa-06b38d89f85b · outbound

This paper cites GPT-4 Technical Report.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets GPT-4 Technical Report

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:09.007916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:27310ab7980d16055addb87cc3673d86179545734b719334a46c82daf8811b69

Observation 8fc72112-78cd-4984-a35f-f23dfde16900 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:08.998421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:9dfedbdb83204c56fc4c9f52f7e351e95a611479cf2ab03315f07f9b63e95d53

Observation ea7d97e4-1a6c-4a93-8a8b-b65bd210e62b · outbound

This paper cites Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:40eb92e7760f7ec44b25a3ee3700034b434820e536122a430f788dae09ec9b02

Observation 84da6844-188f-481e-8e8b-d348186503b3 · outbound

This paper cites Livestar: Live streaming assistant for real-world online video understanding,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Livestar: Live streaming assistant for real-world online video understanding,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:7a8a8e8a7b3289f842ba45be95a3849d6d3b511bd6ce07e133cdaa145a0e7d86

Observation e322133c-fdce-411d-a25c-5b0bdb5004bc · outbound

This paper cites Querystream: Advancing streaming video understanding with query-aware pruning and proactive response,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Querystream: Advancing streaming video understanding with query-aware pruning and proactive response,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:c9e47f036ea12b24dda5e17ddc99498b4eb929083545ef7eded34bb755e865fe

Observation 87837274-f1b6-4a2a-a0d5-fce1a8b971cc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Gemini: A Family of Highly Capable Multimodal Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:09.002889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:2eb8a2f008aade82e0288732068117305d3e2063e9a67d50b43c7f742259882f

Observation 849c4da3-674c-47c1-aedf-dc96647ad78d · outbound

This paper cites Reason-before-retrieve: One-stage reflective chain- of-thoughts for training-free zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Reason-before-retrieve: One-stage reflective chain- of-thoughts for training-free zero-shot composed image retrieval,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:5a615126f636e557733a32a7b44ad9f5683f5bd84f3dc913bedb9fb32501e1b2

Observation 32752517-fb52-4f96-9490-f87197612446 · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Merlot reserve: Neural script knowledge through vision and language and sound,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:3f99344ee56eb19f4b32a5704da2b750dde36c3b7762593d5a57159e9d24449e

Observation dd0dc6ae-f94d-4901-986a-e936aa0474f9 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:09.012450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:9df1f0c256682ae06a788c5aae60d431a79d9cdb770aa24ee1cf23fc4602a654

Observation 45819137-670d-4572-be5a-8ff5c0445cea · outbound

This paper cites Vision-by-language for training-free compositional image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Vision-by-language for training-free compositional image retrieval,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:612f51bd978ca7f8d6dfdcbdac358bd714410b29df16fd97579af3f41e3b1680

Observation 286c6393-1da7-4f5c-9a2d-ff7a455c6348 · outbound

This paper cites "this is my unicorn, fluffy.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets "this is my unicorn, fluffy

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:0baa61142b7a23cb5aef7d68675b88168b2f71bf4d00f54ec1a5a8338336a762

Observation e188e77e-1f62-4589-b813-835bb1d6960e · outbound

This paper cites iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:08.994177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:3204aabc32eda6a6aba32cf5b06e280aee575c4b84e930d2cd2402e0d8166654

Observation 0d8b53f1-430a-4377-b492-5bd2ece3deec · outbound

This paper cites Language-only training of zero-shot composed image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Language-only training of zero-shot composed image retrieval,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:a1b8085f2111d63856f4cc07d6db3a6af2ea53717e58105786f2d1e40b632be2

Observation 22f97239-b564-4257-886d-d9d58a6ce561 · outbound

This paper cites ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.015326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:95958715bab7b81146b2749ac8b43d9e03696e7acc896cbe1b55c2645a3f873f

Observation d48b8729-0fc2-40d1-855b-1ad3ad73701a · outbound

This paper cites Amc: Adaptive multi-expert collaborative network for text-guided image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Amc: Adaptive multi-expert collaborative network for text-guided image retrieval,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:1a7bf45ea8dc5bc8ba264ea8f4168f2985032fe4d25f80c1f450e8cb408f0543

Observation b87902bf-c34b-49a5-9146-4bcc2b93b382 · outbound

This paper cites Modality-Agnostic Attention Fusion for visual search with text feedback.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Modality-Agnostic Attention Fusion for visual search with text feedback

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:09.022557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:5a13af637c0a3a422acb371a0829c77cfd88f27cf465057a1a1795054017addb

Observation bbd71be3-e5c7-4cca-a8c9-c14b10453cc5 · outbound

This paper cites Dynamic weighted combiner for mixed-modal image retrieval,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Dynamic weighted combiner for mixed-modal image retrieval,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:54fb09ef2cd2a4a01ea35b104fa8a462529fd70b4a50d8c15b8d696317e8fcae

Observation b8cc411e-8009-467a-9582-1e02c021364b · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented clip-based features,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Composed image retrieval using contrastive learning and task-oriented clip-based features,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:44fbf7e9c1667c35ba01b84ecd52d78076cf96e2c304c4aa453c0fececd8812d

Observation 280e856e-9043-4b41-bb6a-e0c918432f1f · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T22:44:36.059582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:050974dfef7c705f1b9f9e87dcfd5e6d13d12e33a273d53c9bfbcff1b02517ec

Pith citing papers

Observation 3062e5b6-9dfd-4c90-876c-a06af031dc4d · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.149055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:972a834571bfdf7c6bf98c7e418d4d3eacb0c820cca4792256e176b6881a6f63