Pith. sign in

Paper Citation Record · LEDGER

UniCoRN: Unified Commented Retrieval Network with LMMs

As of 17 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2502.08254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08254 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:54:16.148845Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy38
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38d6ddc2-c253-4ba6-a33a-97550b330624 · outbound

This paper cites Pixtral 12B.

UniCoRN: Unified Commented Retrieval Network with LMMs Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.808010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.808010Z digest=sha256:e660f34bda0b21e379f1863c311044fae2d2bf57a30f9a7b0aa63760bc821f1f

Observation c558e5bf-2f8e-45f7-b00a-23bdc8214473 · outbound

This paper cites Learning attribute representations with local- ization for flexible fashion search.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning attribute representations with local- ization for flexible fashion search

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.812393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.812393Z digest=sha256:133bc05d41fb62dfb12047c9cb8ccf6ac42ae80c3605d34628c72a24c37e7f50

Observation 2a2b8b81-bc1c-4b50-9095-641d36a52911 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.815852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.815852Z digest=sha256:2a6d5f2541d6338b3ca749649838c7d338da9bfb2f1f0ddce2b7cd3df2ed07d0

Observation b52e3554-78f2-4aa7-a5e2-19a3cc08e183 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

UniCoRN: Unified Commented Retrieval Network with LMMs The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.819291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.819291Z digest=sha256:3632182cd87d22bf1858c72c7074483d848f07c2e5211e03ebe6e3034bd51c9b

Observation 992f8b97-ed79-42cf-be80-18f80ad192f2 · outbound

This paper cites Vqa: Visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Vqa: Visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.822811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.822811Z digest=sha256:e4ae2c7e176337e397792e7bff6c2a10dbe27da7bbb1f1eea9bc336791474051

Observation d9590b7f-e4e7-4517-aa01-fb2b35046b20 · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

UniCoRN: Unified Commented Retrieval Network with LMMs Effective conditioned and composed im- age retrieval combining clip-based features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.826342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.826342Z digest=sha256:98e73e9ea852f744a14ea346149009986344f601307fb68aec3c139d16d0e08c

Observation 8034d24e-c299-4e48-9ba7-8187cac9e97b · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

UniCoRN: Unified Commented Retrieval Network with LMMs Zero-shot composed image retrieval with textual inversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.829572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.829572Z digest=sha256:b7d38ba4aeb478c6af28dad6f7223b33baf397f1fce3d7e334332fab27cd755f

Observation c2c2326d-a5c4-4552-bb9f-fcde8a26abd2 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts.

UniCoRN: Unified Commented Retrieval Network with LMMs Vlmo: Unified vision-language pre-training with mixture-of-modality-experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.833054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.833054Z digest=sha256:a622598f7bec7dbab2a2d7c019667efad315e28c4a8223951888cf89fb03572c

Observation 06e54353-c76a-4c4a-a521-e9f159eade34 · outbound

This paper cites Tomayto, tomahto.

UniCoRN: Unified Commented Retrieval Network with LMMs Tomayto, tomahto

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.836907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.836907Z digest=sha256:9a24c6586aa9dc53bb45ff1fbc40114f38a80b4eb9aa9e5f0ab61dce04d96b6f

Observation d4a09c30-4cdf-40c6-a317-5863e240dfcf · outbound

This paper cites A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding.

UniCoRN: Unified Commented Retrieval Network with LMMs A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.840236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.840236Z digest=sha256:676658a1e917946b435d64ce9098d6b287682736398fb7933e18bbc450a5a73a

Observation bbd4ef3c-0289-4215-9ebd-abf83f7d2828 · outbound

This paper cites Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo.

UniCoRN: Unified Commented Retrieval Network with LMMs Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.843793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.843793Z digest=sha256:6cd1c09b5e14257cfecf0206871313ed7a3992ef89a62da2e2f6cfeb03eccdd9

Observation a669727d-692c-4509-ba45-bcbbc5839f3d · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented gener- ation for multimodal llms.

UniCoRN: Unified Commented Retrieval Network with LMMs Wiki-llava: Hierarchical retrieval-augmented gener- ation for multimodal llms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.847113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.847113Z digest=sha256:e0b304125ac9023ab7f4eaf4542640e45169c832e7347c3ff2ef38aa35ec4211

Observation 9a0755fe-5633-45ce-a0dc-c5be88c1260e · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

UniCoRN: Unified Commented Retrieval Network with LMMs MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.850481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.850481Z digest=sha256:4181eda627259c324c64783fde25f23919789c173d2d56178a9a64d7b367008f

Observation 1369fc8e-ca41-41cd-ab5d-3da27cc877e4 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

UniCoRN: Unified Commented Retrieval Network with LMMs PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.853944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.853944Z digest=sha256:df239e0865669d03b580bf0bdbecba6c33fd40a95d8df023122817cfe30c788f

Observation bed027d4-31ba-4427-b29e-619051466185 · outbound

This paper cites Image search with text feedback by visiolinguistic attention learn- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Image search with text feedback by visiolinguistic attention learn- ing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.857665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.857665Z digest=sha256:263bfdc303fc71780a703ec9781172d5953bb8bd664afa559994aa736bc8edbf

Observation 3add3c2b-6dd9-44a2-b881-b45c5638c29a · outbound

This paper cites Image search with text feedback by visiolinguistic attention learn- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Image search with text feedback by visiolinguistic attention learn- ing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.860861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.860861Z digest=sha256:e8ee0487f7a215b7dce4cd9b6dba70ce5f294f0d96fc89810548f4ba3373eb73

Observation 3109636d-67a5-415a-baff-0e63cb21a129 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

UniCoRN: Unified Commented Retrieval Network with LMMs Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.864357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.864357Z digest=sha256:f12731363a48d14dd7a61f834f90fbfb876c2c547f44b12ad3966d45d32ef80d

Observation 4ef45408-0c61-4f29-abdb-4b6c0de8ec1d · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.868000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.868000Z digest=sha256:5da7184ec735515f96bc95aedfe59c8656e125bd23b2d1ea0642f5eea6a6b277

Observation 425d9efd-8881-43fa-b621-d4a2072ea104 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.871260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.871260Z digest=sha256:63bfaa1eda4bbf6aafd8335a91aa54e88c58363c8884e704edb6d195254ed26f

Observation bafc4799-8cb8-4f13-a03e-a855dd5c3956 · outbound

This paper cites Meteor universal: Lan- guage specific translation evaluation for any target language.

UniCoRN: Unified Commented Retrieval Network with LMMs Meteor universal: Lan- guage specific translation evaluation for any target language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.874549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.874549Z digest=sha256:361e80644f7e259f2c9cdc074130671656f00a6d70d9f834f4e508f048289ebc

Observation fcec966b-795a-4e81-85ad-db0c89a83a69 · outbound

This paper cites Hyper- bolic image-text representations.

UniCoRN: Unified Commented Retrieval Network with LMMs Hyper- bolic image-text representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.877970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.877970Z digest=sha256:3e59e81390987b06ce4d3df32de31f2e051e40550ae322efbb5b1fabdf01108c

Observation 5ff08aba-4c39-439f-b204-4a38fe95f237 · outbound

This paper cites Toutanova.

UniCoRN: Unified Commented Retrieval Network with LMMs Toutanova

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.881162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.881162Z digest=sha256:f39b45b80dd39bbd0b6d28afff7e6302c8b18573ca8aea33aae31df48b18a2e0

Observation 1073d162-89b2-4794-983c-09792b776666 · outbound

This paper cites The Llama 3 Herd of Models.

UniCoRN: Unified Commented Retrieval Network with LMMs The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.884750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.884750Z digest=sha256:6ff2706b929865dd8dffd7b7cf4751235a9906d8b88280f4f8935f5bf7585d13

Observation 16874e68-f265-4254-b4cf-48e4ff28f391 · outbound

This paper cites Entities as Experts: Sparse Memory Access with Entity Supervision.

UniCoRN: Unified Commented Retrieval Network with LMMs Entities as Experts: Sparse Memory Access with Entity Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.888711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.888711Z digest=sha256:e25dad092063102988d87e11e46bf2fb7211ba1349102a1e05b464122592038f

Observation 033cdd40-4ff8-48c0-8b12-7f28458d3ef7 · outbound

This paper cites Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretrain- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretrain- ing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.892674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.892674Z digest=sha256:96f4a8a60fb71f451579f3d2e980309a50635ce007a9b303e60dd1e0fd88e102

Observation dc07771b-3808-4fa9-8b0a-7919c7744148 · outbound

This paper cites SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger.

UniCoRN: Unified Commented Retrieval Network with LMMs SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:16.446151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.895975Z digest=sha256:c0a5a73caf8d5a94b30e4046222c0f53329dbcdb342b90009ed7e40db46f9655

Observation 4f9696d8-933d-4d91-8b84-0d05be0fc9a4 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

UniCoRN: Unified Commented Retrieval Network with LMMs Making LLaMA SEE and Draw with SEED Tokenizer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.899573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.899573Z digest=sha256:bd0f67aa2854001d5a88af263d49a5b9a78a9ed3b6fe040f881cc6954ab0f66d

Observation 67f67b07-c0a4-4475-8e0d-f043e95af363 · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

UniCoRN: Unified Commented Retrieval Network with LMMs Cyclip: Cyclic contrastive language-image pretraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.902988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.902988Z digest=sha256:7d2bbf41dcdfef80c8586f46f663b48fef19bf65510cf99564fc1a79e228bac4

Observation f34a8fe8-19ff-46b6-9529-34b77634a9ef · outbound

This paper cites Fashionvlp: Vision language transformer for fashion re- trieval with feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashionvlp: Vision language transformer for fashion re- trieval with feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.946462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.906213Z digest=sha256:8e2a69bbf51506933c8a008a3cdcd6a4d8f51eeacb003963957cbc0432da38f0

Observation 53408d33-2e4e-41c6-8db8-c566824c86b0 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.935833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.909463Z digest=sha256:6e86587f60db0dc703ce30852b809e966d46c1f598c47afe00f2cd152f28c97d

Observation a5810dab-5893-47e6-b4fe-c280901bc02d · outbound

This paper cites Language-only training of zero-shot com- posed image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Language-only training of zero-shot com- posed image retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.925444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.912732Z digest=sha256:4c1629102a7f06f16d2a1005cb6d13bf0a12d61d167a12bc66b7af9e8cf38d85

Observation efdad6e1-a744-4786-9a69-dbe3c71d3d7b · outbound

This paper cites Dialog-based interactive image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Dialog-based interactive image retrieval

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.915546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.915885Z digest=sha256:e59831bec4ee666c4333f9540f663f0f8edc0241941838e51c8ec95aaaed749b

Observation 1ffed5d6-3957-444a-8568-7c2949410145 · outbound

This paper cites Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.919336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.919336Z digest=sha256:1dcd1265f0c11e56cba57a84e4d1265cca2b4ca2c16508d8c0a8333241f79dca

Observation c70ead6c-7350-4f67-9286-00a82bb5d4b7 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

UniCoRN: Unified Commented Retrieval Network with LMMs Vizwiz grand challenge: Answering visual questions from blind people

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.922958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.922958Z digest=sha256:74d098445321d250095b5474fe8f2d130b8cd02af76f62e744c6c89dcfc470cf

Observation 4eca5215-0015-4a15-99d2-9b72b539afee · outbound

This paper cites Retrieval augmented language model pre- training.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval augmented language model pre- training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.926563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.926563Z digest=sha256:ae0a8004c6d69bd7e0a3b084d37a9f70b74398c124f91282c739ba54ca2a7770

Observation a3130d19-48be-4574-bc75-17cbf7b07ff9 · outbound

This paper cites Au- tomatic spatially-aware fashion concept discovery.

UniCoRN: Unified Commented Retrieval Network with LMMs Au- tomatic spatially-aware fashion concept discovery

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.893871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.929768Z digest=sha256:204dba927cb6fb417525d2eb022457549557417947c9b333baa1400a3fe261d6

Observation 8472be4a-c96f-495f-b465-d90702b91cfa · outbound

This paper cites Learning attribute-driven disentangled represen- tations for interactive fashion retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning attribute-driven disentangled represen- tations for interactive fashion retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.884577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.933138Z digest=sha256:c445cc0e3943d6f7c5d48731598ef9fa64cc881da6a68c80af3f7e3c818a8cf6

Observation 6e15b028-1da2-4866-996b-1d52a2d3d25a · outbound

This paper cites Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities.

UniCoRN: Unified Commented Retrieval Network with LMMs Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.874761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.936531Z digest=sha256:dfa0940c07976502cb50f691b5d117e6353bb37f798f7764698aceb7f32a3bba

Observation 97fb0400-2a4e-4e51-8ccb-fad1c82afafc · outbound

This paper cites Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory.

UniCoRN: Unified Commented Retrieval Network with LMMs Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.865028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.939817Z digest=sha256:b8e5405ad376243d75506f038c3f236a0501c309008fcece58dbe6f4b34983f3

Observation 88d8d982-2371-488e-9f1d-7496f0ef343d · outbound

This paper cites Openclip, 2021.

UniCoRN: Unified Commented Retrieval Network with LMMs Openclip, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.855413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.943126Z digest=sha256:2fd81f12eb029f069719671bfb8936ebf693eb8bb02158d6840d19cf0ec2f766

Observation aeb94a3a-7f2b-4238-a29d-50dc08a65f5f · outbound

This paper cites Mantis: Interleaved multi-image instruction tuning, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Mantis: Interleaved multi-image instruction tuning, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.845876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.946618Z digest=sha256:da6cdfd2f2b7fcbb3bc549df25a5509274cbbf1594064d0648ff6ee9f98cdff1

Observation 11b4e977-6a6e-4ec8-8e66-6549b62d0d4e · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.949814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.949814Z digest=sha256:7d88cd30cfe94800fff5a01ed38cf15136d2cc10bf81d650529102f85ea57a74

Observation 40be1038-6517-4deb-a45e-396b5470b8e5 · outbound

This paper cites Dense Passage Retrieval for Open-Domain Question Answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Dense Passage Retrieval for Open-Domain Question Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.953157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.953157Z digest=sha256:58e556a438f2c5857eccd71d18c7f882e58d1660374c529fc4b1014d3de37d7d

Observation a2540e4b-15b0-4ffc-b70e-47acf8f5fe44 · outbound

This paper cites Vision-by-Language for Training-Free Compositional Image Retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Vision-by-Language for Training-Free Compositional Image Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.956672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.956672Z digest=sha256:54707f18a776d2d53ba4dd68a28c10f0205ef9a003b82aa7d3c769a21cc3600b

Observation 2fefdf49-ab73-4198-b080-6f82ff96ad3f · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

UniCoRN: Unified Commented Retrieval Network with LMMs Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.835520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.960045Z digest=sha256:e2448b9c06b046974d275999fe6aeecb0fd78b42380d3e533e2722a205c6c2c0

Observation 449e2b1a-41d0-4cf8-bf8d-2f5a379d0626 · outbound

This paper cites Grounding language models to images for multimodal in- puts and outputs.

UniCoRN: Unified Commented Retrieval Network with LMMs Grounding language models to images for multimodal in- puts and outputs

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.824825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.963359Z digest=sha256:fe36ff424abb095dee3f32017c3ebf33dd014a29765c8ec4348769577c2e45eb

Observation a4e72049-9e5a-47b4-9670-8f0756296a55 · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.966675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.966675Z digest=sha256:37050be8a3ed3ebffce6299e72bb3fcc0211bca706bf2c71f850dfa1a4b943ed

Observation 24c7f62a-e6d4-4eb0-85a4-49e7e456ee08 · outbound

This paper cites Latent Retrieval for Weakly Supervised Open Domain Question Answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Latent Retrieval for Weakly Supervised Open Domain Question Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.970255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.970255Z digest=sha256:6b8a4dcf1cbce88c70bf522e4d36885fd01849f482dd79a3bf556da8e4d731bd

Observation a26127f8-216b-4c52-99f4-09b99c72868e · outbound

This paper cites Chatting makes perfect: Chat-based image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Chatting makes perfect: Chat-based image retrieval

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.814878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.974015Z digest=sha256:af10efafaa78d321f1b7c275db3b63fdac6886fa4c186378aaabf3ed219e1f70

Observation 707197bd-11e0-4063-856c-5150dd2e5bbf · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.805378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.977277Z digest=sha256:cd74e7937db1822a2f37dc25ac74ffd0be931665ab89bb06666f5511fa0f17f1

Observation c1f5bef1-4759-4615-879e-d60368870870 · outbound

This paper cites Retrieval-augmented genera- tion for knowledge-intensive nlp tasks, 2021.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval-augmented genera- tion for knowledge-intensive nlp tasks, 2021

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.795786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.980488Z digest=sha256:b6b6f8bc71e8b1137abb9169d696c7a81d41e1c02e661cebc10f3c9af266ac4d

Observation 96d48774-0e2e-4239-a0e7-7e03dee3e2a1 · outbound

This paper cites Textbind: Multi-turn interleaved multimodal instruction- following in the wild.

UniCoRN: Unified Commented Retrieval Network with LMMs Textbind: Multi-turn interleaved multimodal instruction- following in the wild

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.785517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.983696Z digest=sha256:febc7269b52d335f607cd9e3539392b4bc22b68cc774e347205c44833dc8fdb1

Observation fd3196f1-a8bc-467e-954c-0fcacac5b9c5 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

UniCoRN: Unified Commented Retrieval Network with LMMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.987047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.987047Z digest=sha256:e7635ee27d585dcfbe8d7dec8ca98ed407d4a8226d4606ec8bf03899addc1852

Observation 2e4636fd-fcbc-470c-b437-31dd485575b2 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

UniCoRN: Unified Commented Retrieval Network with LMMs Rouge: A package for automatic evaluation of summaries

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.990318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.990318Z digest=sha256:cc125f0567adee0d77ccbb0e13569dafc298190c7060ebe9bbd0d18b4fdcd18d

Observation 591ce93e-7e28-434c-9f98-8987dc53e7b8 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

UniCoRN: Unified Commented Retrieval Network with LMMs MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.993546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.993546Z digest=sha256:aa6fd75e70cab7558d5e0d1c8af4ea4e513e81a0a168e813d89c710b6400928b

Observation 7ee84982-63be-4cba-a493-145152d42283 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

UniCoRN: Unified Commented Retrieval Network with LMMs Improved baselines with visual instruction tuning, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.762848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:15.997037Z digest=sha256:d21336e8e34d26bcf3d929198059bf5c0a7ffaf0faf0304e9965e7d61582b95b

Observation 88b233ce-38d6-4a2a-9fd8-41b3558cbc82 · outbound

This paper cites Visual instruction tuning, 2023.

UniCoRN: Unified Commented Retrieval Network with LMMs Visual instruction tuning, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.751844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.000680Z digest=sha256:76769dcfd32214b40eb227a643e0ff96d01fa74221fd5562c8a8546eee1f48a4

Observation 5d36c078-a63d-47a0-9cb2-52bd215e595f · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

UniCoRN: Unified Commented Retrieval Network with LMMs Image retrieval on real-life images with pre-trained vision-and-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.739618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.003945Z digest=sha256:acefc542764554418f45a80b0955d9ebce48b9110e3a6a64de8e51434856e402

Observation cc3c7b49-a857-4c7a-bcc2-088717bdaeae · outbound

This paper cites Image retrieval on real-life images with pre- trained vision-and-language models.

UniCoRN: Unified Commented Retrieval Network with LMMs Image retrieval on real-life images with pre- trained vision-and-language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.007282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.007282Z digest=sha256:7ff110ceaea2bac5df5ac6e944304e3fcc6d5124e6c33504e18e6735133c8377

Observation 4bb2271a-8d20-46fd-ae2b-601b1da7987c · outbound

This paper cites Bi-directional training for composed im- age retrieval via text prompt learning.

UniCoRN: Unified Commented Retrieval Network with LMMs Bi-directional training for composed im- age retrieval via text prompt learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.723592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.010587Z digest=sha256:f565bc93895d8eaa597e594a1844377138099e83babaa2809cad456de4be4732

Observation 6d68932e-aee0-4a36-9613-52f3a90496fe · outbound

This paper cites Three facets of visual and verbal learners: Cognitive ability, cognitive style, and learning preference.

UniCoRN: Unified Commented Retrieval Network with LMMs Three facets of visual and verbal learners: Cognitive ability, cognitive style, and learning preference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.713434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.013721Z digest=sha256:d5974869aa304ce98a8f670972bc23f1c76f5d2baafc9198b9d5bd2298cea341

Observation e1354ed6-e14f-4de9-8f5e-bbb089595898 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.017042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.017042Z digest=sha256:9eac46168f61952d5cd269d707268f06e13278924800437f124fc462252dec47

Observation f95c9202-b167-4158-9d88-f3e1a0bdc10d · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

UniCoRN: Unified Commented Retrieval Network with LMMs Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.703498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.020707Z digest=sha256:d25a4ca11527e352d8075816eec2d55e2c499a90a2fb74237039de8b40dd35b8

Observation 4b4b17d4-a55e-4555-89c1-548ddd5e7829 · outbound

This paper cites Gpt-4 technical report, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Gpt-4 technical report, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.024210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.024210Z digest=sha256:fa42a094ed637503818d021a6e248dbb4dc835460faa1d8e5a6d289e48cab489

Observation 342eb1c4-f58d-44fa-931b-f0ff86bb93fd · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

UniCoRN: Unified Commented Retrieval Network with LMMs Bleu: a method for automatic evaluation of machine translation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.687287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.027601Z digest=sha256:5efdcbbf19184dc78dd425cc45bc521234c199b4da0bd4ec100fc3c0e0f0309a

Observation a94d9f92-3caa-43da-9e5f-d4e191272feb · outbound

This paper cites Rora-vlm: Robust retrieval-augmented vision language models, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Rora-vlm: Robust retrieval-augmented vision language models, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.677710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.031009Z digest=sha256:afc5803a3a5aeaec1849f673e2b0a92e43445ac63888067b6a65d1adbc6dc19c

Observation 19c69085-513f-47e0-9035-33bc8a69694c · outbound

This paper cites Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation.

UniCoRN: Unified Commented Retrieval Network with LMMs Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.035119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.035119Z digest=sha256:c5f27bcb37e57aff83b2aad5e0193d65b0eee32b5c2c9aea9e19580b50b42d74

Observation 7b0afb44-c106-4e9a-8e51-81e36db8fc49 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning Transferable Visual Models From Natural Language Supervision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.038739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.038739Z digest=sha256:e268c3e73a0d55cf17ea058bbae24b43a29da13654a73fa9ef89824254538af5

Observation 645a6847-16b2-456f-b901-afd2f19c4813 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

UniCoRN: Unified Commented Retrieval Network with LMMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.042435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.042435Z digest=sha256:6fc052a36b212d87808db4666a1b10619884c333cf89366994e64b9605aa4533

Observation 94e1b93c-2782-482f-8579-efdd03bdd4a6 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

UniCoRN: Unified Commented Retrieval Network with LMMs A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.667678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.045951Z digest=sha256:f8f3738e12570e166420e76a1613190deacfd40acb78cb7020319395fdb187fb

Observation 084a3a0d-94b5-4e4c-b964-2dc08e73e7bd · outbound

This paper cites Kvqa: Knowledge-aware visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Kvqa: Knowledge-aware visual question answering

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.657548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.049269Z digest=sha256:eb68e64113c769eb850be8804f9b304a6848226ab6845b77b13dae65e29a668c

Observation 5932d71f-7388-4f8e-885e-54ca21648e15 · outbound

This paper cites Towards vqa models that can read.

UniCoRN: Unified Commented Retrieval Network with LMMs Towards vqa models that can read

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.646486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.052474Z digest=sha256:c93dc65fb19b946fe3ed14d9b8a2e0638e96169bb00c7aaa1702da10e314f175

Observation b0f5bc17-34e7-4571-bb41-9805ece0c23a · outbound

This paper cites Knowledge-enhanced dual-stream zero-shot composed im- age retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Knowledge-enhanced dual-stream zero-shot composed im- age retrieval

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.635984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.056213Z digest=sha256:9b809093a7dcee73ec007ebfae495e71eaf6acfc948353156ddde7ded60e97c9

Observation e5afdb99-5d1e-4271-a2e3-02333f732556 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCoRN: Unified Commented Retrieval Network with LMMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.059617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.059617Z digest=sha256:718cc36677f57fbd56009ae6ef7487421edfba61f631866f0240b148f794271d

Observation 17392f1e-64e6-4937-80c0-e3260e4c6c9f · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.625967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.063185Z digest=sha256:2a3734f73e72a253255bf4f5f4f287db058f89af334f6c4449d21e94a4e628ef

Observation 918b30d8-a3b9-4390-9cec-7a5a6d442a61 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

UniCoRN: Unified Commented Retrieval Network with LMMs MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.066474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.066474Z digest=sha256:83562360ff526c0f5079ccf1eb4b6f76700267b4e88cddd404344b37405bcaac

Observation faeba532-7ee5-473f-a4ec-791afc12438a · outbound

This paper cites Genecis: A benchmark for general conditional image similarity.

UniCoRN: Unified Commented Retrieval Network with LMMs Genecis: A benchmark for general conditional image similarity

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.615682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.069876Z digest=sha256:94073344be010e7f9e96189ee867b87007ca52b005785251354215a2bab9d45d

Observation c4307a2c-2981-4e25-ac03-72277d72af1e · outbound

This paper cites Composing text and image for image retrieval - an empirical odyssey.

UniCoRN: Unified Commented Retrieval Network with LMMs Composing text and image for image retrieval - an empirical odyssey

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.605302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.073150Z digest=sha256:f043a43e3d15ab37c3e6b20469c0b8195afbde32b409630e0ce4b6c1a754db91

Observation 8a47854d-ddaf-4635-ba28-df7b6e540b35 · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

UniCoRN: Unified Commented Retrieval Network with LMMs Composing text and image for image retrieval-an empirical odyssey

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.076600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.076600Z digest=sha256:11063e71056d7e35bf1cdac27d02c7a501c80f17aebcb50d6406479f3a2dd1fd

Observation d573bb81-416d-4a88-87cc-ffb22d910204 · outbound

This paper cites Cross-modal feature alignment and fusion for com- posed image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Cross-modal feature alignment and fusion for com- posed image retrieval

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.589090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.079980Z digest=sha256:af8a6d66006f2fb2b211b538a8dcd2ada120648f933ec989b3f914e282bd8fc4

Observation 891a069f-594a-423f-a1ca-f90029ac0587 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniCoRN: Unified Commented Retrieval Network with LMMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.083564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.083564Z digest=sha256:886afbd3a2f4251ea8018dc96b788bfa24b9270d2218e1589cbe69ee65e250f7

Observation 182dbbdc-3271-4672-97ca-4db2bebad9f2 · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.578986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.087345Z digest=sha256:16a43bcb23d1427967e191b6473838ef26bcd6d97b0f081a21f269faee27c1c1

Observation 8462c595-82f9-4c73-9e8a-f598044ab3da · outbound

This paper cites UniIR: Training and Benchmarking Universal Multimodal Information Retrievers.

UniCoRN: Unified Commented Retrieval Network with LMMs UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.090827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.090827Z digest=sha256:8c1ae8c5c7c53efad5625b6777149a84935fc60de9359e904c06ec53e60e6f3e

Observation 7faf818e-0027-4616-b6a6-dba6c7623a02 · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.569127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.094338Z digest=sha256:987f33487724d90349887822e29e8e8d4681b917a790d80c60948c1144bfbf77

Observation f480dfe3-c184-478b-a138-44100814778a · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.558907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.097683Z digest=sha256:d7ad831be898d9bb37e3016f8b866f52d1161fec7c1bd7cc204dcb23c64cd0e2

Observation f981b359-073d-4059-9f45-42a1fba8df2e · outbound

This paper cites Visual question answer- ing: A survey of methods and datasets.

UniCoRN: Unified Commented Retrieval Network with LMMs Visual question answer- ing: A survey of methods and datasets

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.101159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.101159Z digest=sha256:d9d63e2b6506ee4911d29bea2097c5308f7cd5046fd7ffb4625d2f5f72655607

Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.104506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.104506Z digest=sha256:041f3596b2085c2de3940692c27ddef5765dc9683f103f1a366987db041a35e0

Observation 22ec78be-1e28-48a4-af82-7030b566ab26 · outbound

This paper cites EchoSight: Advancing visual- language models with Wiki knowledge.

UniCoRN: Unified Commented Retrieval Network with LMMs EchoSight: Advancing visual- language models with Wiki knowledge

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.542820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.108201Z digest=sha256:a769c1e4b19a9f9ba784f82ed7f7a2621e0be57fada0ed883b93a5978ea264d8

Observation 44e175a0-ec5e-4e91-9d78-8d43d48b7cc9 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

UniCoRN: Unified Commented Retrieval Network with LMMs FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.111778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.111778Z digest=sha256:71e98048899d230b4afdb19079fffe19690f916c8ccec6f2ea9760b9560c8537

Observation f3622b38-edbd-4faf-8681-7c8a7e89af82 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

UniCoRN: Unified Commented Retrieval Network with LMMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.115407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.115407Z digest=sha256:207cd21cbf1bcf865e6fcd851cf3c17b070a04e9b549844db30bee0ad1c7c267

Observation 962fba5b-9b5b-4d40-b138-8a495af5dd8b · outbound

This paper cites A Survey on Multimodal Large Language Models.

UniCoRN: Unified Commented Retrieval Network with LMMs A Survey on Multimodal Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.119075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.119075Z digest=sha256:465def20e242b614f483e62a230b399aab9efce9ac044e6fe8f78299acfbd40a

Observation 43394a8b-a74f-4e12-ad5a-39f6911df660 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

UniCoRN: Unified Commented Retrieval Network with LMMs CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.122413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.122413Z digest=sha256:7c946e5fd3529405851dfd26c06f65f0f3162179df4e3d0414f680085bab37e6

Observation 00b79990-a65d-40ee-9b7d-cc78b3d733dc · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniCoRN: Unified Commented Retrieval Network with LMMs Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.126629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.126629Z digest=sha256:98960256c9764af5c53bbc20f9ee8e73abe817b72ea47b8a64c9c5ed0cdf4c82

Observation 3382d6d5-1c24-40b4-a9db-3d32e9e0b111 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

UniCoRN: Unified Commented Retrieval Network with LMMs VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.130352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.130352Z digest=sha256:dc0cb535ae965bb1479c65daca8ab5c5806ea4fb0735579a9b2c5257b92c4342

Observation 5ebec1ff-b09f-4a7e-ba09-4104836ad76b · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

UniCoRN: Unified Commented Retrieval Network with LMMs A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.134206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.134206Z digest=sha256:4590e27ecc9f045ad83a54f973afc60a1be1204ef7737a0123d7061ebdead7d2

Observation a892aaff-44e9-478b-91be-e937e72ad97f · outbound

This paper cites Sigmoid loss for language image pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs Sigmoid loss for language image pre-training

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.532426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.137913Z digest=sha256:f9a3ba5aba9fb7a8fc564ca3285f94ec20344281105444ff4d3e97b8f81f6c48

Observation 151d9cde-c624-4470-b210-bd4e414baa0f · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

UniCoRN: Unified Commented Retrieval Network with LMMs MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.141327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.141327Z digest=sha256:3cee5899eaa0247f6bfa63ebf588886e8194f45b527f9c716742c39d9b36b03a

Observation a4528704-16cb-4ea1-b256-d898499e5c46 · outbound

This paper cites Non-Contrastive Learning Meets Language-Image Pre-Training.

UniCoRN: Unified Commented Retrieval Network with LMMs Non-Contrastive Learning Meets Language-Image Pre-Training

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:16.184255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.145165Z digest=sha256:70e35398d39c9026cf445d73a6323d5dabcc3f69ad0f7f2f7f3f5c6e24ef70a8

Observation e6957db2-f559-4b6d-bc2e-ba3a7e2a2f08 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

UniCoRN: Unified Commented Retrieval Network with LMMs MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.522347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T05:54:16.148845Z digest=sha256:a2e1f890b2b76be99f6cc3f2eb1df613b87de54448b30c1e573665084ed4d3c5

Pith citing papers

No inbound Pith citation observations are available.