Pith. sign in

Paper Citation Record · LEDGER

UniCoRN: Unified Commented Retrieval Network with LMMs

As of 8 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2502.08254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08254 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:54:16.148845Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy38
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38d6ddc2-c253-4ba6-a33a-97550b330624 · outbound

This paper cites Pixtral 12B.

UniCoRN: Unified Commented Retrieval Network with LMMs Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.808010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.808010Z digest=sha256:9dda8de359b96f75ef7248bc369f7bd09e4a8cb3fa38679f80676f0ab7a54933

Observation c558e5bf-2f8e-45f7-b00a-23bdc8214473 · outbound

This paper cites Learning attribute representations with local- ization for flexible fashion search.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning attribute representations with local- ization for flexible fashion search

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.812393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.812393Z digest=sha256:1c06a5c9e4c0a6f5b8dddcde6f0cedeea6b30f7d89df71e4d50aa77d3ec3e59d

Observation 2a2b8b81-bc1c-4b50-9095-641d36a52911 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.815852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.815852Z digest=sha256:abb4da4d639386418aa1ae50c2f01d5ef5cebea0b02431c65cd7d96300a9473a

Observation b52e3554-78f2-4aa7-a5e2-19a3cc08e183 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

UniCoRN: Unified Commented Retrieval Network with LMMs The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.819291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.819291Z digest=sha256:cb5254069441ed48824d1ffbcac644b245d4b2cde822e37299d28f92d6e6d0b4

Observation 992f8b97-ed79-42cf-be80-18f80ad192f2 · outbound

This paper cites Vqa: Visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Vqa: Visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.822811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.822811Z digest=sha256:407fd710097632ae48621900a68a5d9b688dadf0ac5a496fd23256a8ee14c6b0

Observation d9590b7f-e4e7-4517-aa01-fb2b35046b20 · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

UniCoRN: Unified Commented Retrieval Network with LMMs Effective conditioned and composed im- age retrieval combining clip-based features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.826342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.826342Z digest=sha256:6a144e622ea6e648ab008d9c812acd0d20ead07974923fd8f936199b526feb01

Observation 8034d24e-c299-4e48-9ba7-8187cac9e97b · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

UniCoRN: Unified Commented Retrieval Network with LMMs Zero-shot composed image retrieval with textual inversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.829572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.829572Z digest=sha256:d0b866225467648318a39b0d94967b4d2dcc88ceb424c58a857c55bea8649654

Observation c2c2326d-a5c4-4552-bb9f-fcde8a26abd2 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts.

UniCoRN: Unified Commented Retrieval Network with LMMs Vlmo: Unified vision-language pre-training with mixture-of-modality-experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.833054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.833054Z digest=sha256:efd77b35b2b71fc3f40f6967a3fc9057d452141671744dde2a2505345b2b367e

Observation 06e54353-c76a-4c4a-a521-e9f159eade34 · outbound

This paper cites Tomayto, tomahto.

UniCoRN: Unified Commented Retrieval Network with LMMs Tomayto, tomahto

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.836907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.836907Z digest=sha256:16fb670cca982ce54ecde5bd9afbc18b400f551e594ec9f120ffc6979c675ce6

Observation d4a09c30-4cdf-40c6-a317-5863e240dfcf · outbound

This paper cites A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding.

UniCoRN: Unified Commented Retrieval Network with LMMs A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.840236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.840236Z digest=sha256:5b3e00c2325652275c38cf981c8801881993d221334d12f347b65da22b53dd78

Observation bbd4ef3c-0289-4215-9ebd-abf83f7d2828 · outbound

This paper cites Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo.

UniCoRN: Unified Commented Retrieval Network with LMMs Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.843793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.843793Z digest=sha256:3747dec626f1781e2c6fb56f4508c576b448c2085ee0a31b0d40cba2e372a450

Observation a669727d-692c-4509-ba45-bcbbc5839f3d · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented gener- ation for multimodal llms.

UniCoRN: Unified Commented Retrieval Network with LMMs Wiki-llava: Hierarchical retrieval-augmented gener- ation for multimodal llms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.847113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.847113Z digest=sha256:40c2910b117a995d1b7b599d5de0df14badca62969d394dcaad686cd82406405

Observation 9a0755fe-5633-45ce-a0dc-c5be88c1260e · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

UniCoRN: Unified Commented Retrieval Network with LMMs MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.850481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.850481Z digest=sha256:8de51a648fc9ad7f3e41fe4edf9968dada30379f0891df920df07a029108731a

Observation 1369fc8e-ca41-41cd-ab5d-3da27cc877e4 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

UniCoRN: Unified Commented Retrieval Network with LMMs PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.853944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.853944Z digest=sha256:be04d3d4be9e9cb74c07ca9f140ff978e12aa6541554156ea7484da1158023fb

Observation bed027d4-31ba-4427-b29e-619051466185 · outbound

This paper cites Image search with text feedback by visiolinguistic attention learn- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Image search with text feedback by visiolinguistic attention learn- ing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.857665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.857665Z digest=sha256:f2292e7e0ddabf3023846b4e3018b6dbcd06a8f35e335aefc89b44397aa03819

Observation 3add3c2b-6dd9-44a2-b881-b45c5638c29a · outbound

This paper cites Image search with text feedback by visiolinguistic attention learn- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Image search with text feedback by visiolinguistic attention learn- ing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.860861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.860861Z digest=sha256:d8c7f23e808ae8c681268f4b0f31b8912159c6db1801190284788b641f4c5bd9

Observation 3109636d-67a5-415a-baff-0e63cb21a129 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

UniCoRN: Unified Commented Retrieval Network with LMMs Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.864357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.864357Z digest=sha256:a2753170cec9e78d8ab4bb9be59ae7d8be2117a5656bb3e3f4c493bc9caf20f6

Observation 4ef45408-0c61-4f29-abdb-4b6c0de8ec1d · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.868000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.868000Z digest=sha256:e5fe2c6b3eff3e974d8b851dda4a562c92bfab9dba32c0dc4ad92bafdcd3550f

Observation 425d9efd-8881-43fa-b621-d4a2072ea104 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.871260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.871260Z digest=sha256:cfc5ffce68f6bff9b9865be01f1eafb118de93ec8831c880950e8fd005a031e4

Observation bafc4799-8cb8-4f13-a03e-a855dd5c3956 · outbound

This paper cites Meteor universal: Lan- guage specific translation evaluation for any target language.

UniCoRN: Unified Commented Retrieval Network with LMMs Meteor universal: Lan- guage specific translation evaluation for any target language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.874549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.874549Z digest=sha256:d76a2c5bff389af980d6c821c443b3d7e3a5690f2b554d7fe6da3747ef7d2b0e

Observation fcec966b-795a-4e81-85ad-db0c89a83a69 · outbound

This paper cites Hyper- bolic image-text representations.

UniCoRN: Unified Commented Retrieval Network with LMMs Hyper- bolic image-text representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.877970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.877970Z digest=sha256:18217e852ac12b25d4127b0c5e190f49d81642f7e29d1f7f9909042eab0ed30b

Observation 5ff08aba-4c39-439f-b204-4a38fe95f237 · outbound

This paper cites Toutanova.

UniCoRN: Unified Commented Retrieval Network with LMMs Toutanova

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.881162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.881162Z digest=sha256:79ee8a0b25cae631bd46b5387fee614a1380035efe2c6185d50562a5e533e24b

Observation 1073d162-89b2-4794-983c-09792b776666 · outbound

This paper cites The Llama 3 Herd of Models.

UniCoRN: Unified Commented Retrieval Network with LMMs The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.884750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.884750Z digest=sha256:09a6e392d1a2b090835f9e85c6a5b070fb6121a74edb422e136e6e2299b6d325

Observation 16874e68-f265-4254-b4cf-48e4ff28f391 · outbound

This paper cites Entities as Experts: Sparse Memory Access with Entity Supervision.

UniCoRN: Unified Commented Retrieval Network with LMMs Entities as Experts: Sparse Memory Access with Entity Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.888711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.888711Z digest=sha256:744af67cfb33a819a6e7f36510b187128e5c3241e3dffc1603a4e1e341d6fd84

Observation 033cdd40-4ff8-48c0-8b12-7f28458d3ef7 · outbound

This paper cites Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretrain- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretrain- ing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.892674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.892674Z digest=sha256:c352253e115c284bee35183d68ae6bdc48a4c92e5aabd27853daf6e6cbe87bc5

Observation dc07771b-3808-4fa9-8b0a-7919c7744148 · outbound

This paper cites SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger.

UniCoRN: Unified Commented Retrieval Network with LMMs SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:16.446151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.895975Z digest=sha256:235652e64f9b644ca55cb1e088626657d6e7d904a9b6222cb7d369b1de460b1d

Observation 4f9696d8-933d-4d91-8b84-0d05be0fc9a4 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

UniCoRN: Unified Commented Retrieval Network with LMMs Making LLaMA SEE and Draw with SEED Tokenizer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.899573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.899573Z digest=sha256:d20eff50a6a52d3499d3b62039f65024c8e849f551f306770bb3bf855ceefcd5

Observation 67f67b07-c0a4-4475-8e0d-f043e95af363 · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

UniCoRN: Unified Commented Retrieval Network with LMMs Cyclip: Cyclic contrastive language-image pretraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.902988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.902988Z digest=sha256:c9adb3f8439cbd96172a4305f9fd178aed09bc8ea6fdfb6095ed91f8b6c67afd

Observation f34a8fe8-19ff-46b6-9529-34b77634a9ef · outbound

This paper cites Fashionvlp: Vision language transformer for fashion re- trieval with feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashionvlp: Vision language transformer for fashion re- trieval with feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.946462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.906213Z digest=sha256:be427faa55bfd7d31f9ccdd09291764a273b837d5af3f0fb808e299282cb449b

Observation 53408d33-2e4e-41c6-8db8-c566824c86b0 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.935833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.909463Z digest=sha256:0efcf57723cfa46cc6463a791f8d5b6a1487c430e57f784e0010956abc361a3d

Observation a5810dab-5893-47e6-b4fe-c280901bc02d · outbound

This paper cites Language-only training of zero-shot com- posed image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Language-only training of zero-shot com- posed image retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.925444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.912732Z digest=sha256:5ff4e677afc6f51ce5ba85896ed6024a950a69747470ba2ab2f89fbe18c08385

Observation efdad6e1-a744-4786-9a69-dbe3c71d3d7b · outbound

This paper cites Dialog-based interactive image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Dialog-based interactive image retrieval

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.915546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.915885Z digest=sha256:8f8e35417b51e9fdb502ea8de898aec42cb8a6c00bb10ed2ed72d2567198dacf

Observation 1ffed5d6-3957-444a-8568-7c2949410145 · outbound

This paper cites Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.919336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.919336Z digest=sha256:870fdfb1a2a47ec961f426a1dd152a90226314d470f92b2f90517e4191768893

Observation c70ead6c-7350-4f67-9286-00a82bb5d4b7 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

UniCoRN: Unified Commented Retrieval Network with LMMs Vizwiz grand challenge: Answering visual questions from blind people

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.922958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.922958Z digest=sha256:fda9ac2aa840e8c13cb7a4e0209ebbfa044e3fb3841e3e58d525fb64233d564f

Observation 4eca5215-0015-4a15-99d2-9b72b539afee · outbound

This paper cites Retrieval augmented language model pre- training.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval augmented language model pre- training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.926563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.926563Z digest=sha256:18abff8fa89f4c5b0bb7a7531b4f2eb35729ffdd7919c950269189d4d531b117

Observation a3130d19-48be-4574-bc75-17cbf7b07ff9 · outbound

This paper cites Au- tomatic spatially-aware fashion concept discovery.

UniCoRN: Unified Commented Retrieval Network with LMMs Au- tomatic spatially-aware fashion concept discovery

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.893871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.929768Z digest=sha256:65c7c6beaa20c2468f39ee97d62a0a6c23c0bc536ea060f3345d8a8c9c2856d2

Observation 8472be4a-c96f-495f-b465-d90702b91cfa · outbound

This paper cites Learning attribute-driven disentangled represen- tations for interactive fashion retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning attribute-driven disentangled represen- tations for interactive fashion retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.884577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.933138Z digest=sha256:2caa3d4ae9ff90e8b997ffa32c980d72bffe454d0882d2eb1ba09250d1167470

Observation 6e15b028-1da2-4866-996b-1d52a2d3d25a · outbound

This paper cites Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities.

UniCoRN: Unified Commented Retrieval Network with LMMs Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.874761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.936531Z digest=sha256:3bf6ba2e8944736f9af57b38b7a9d2b3c936abe4c7f7487f3fa8740989617272

Observation 97fb0400-2a4e-4e51-8ccb-fad1c82afafc · outbound

This paper cites Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory.

UniCoRN: Unified Commented Retrieval Network with LMMs Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.865028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.939817Z digest=sha256:078bf082464efa9b0dc2daed94c9c686eefab1d6de34578a027d6a567ba4831b

Observation 88d8d982-2371-488e-9f1d-7496f0ef343d · outbound

This paper cites Openclip, 2021.

UniCoRN: Unified Commented Retrieval Network with LMMs Openclip, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.855413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.943126Z digest=sha256:b8efb0c0b6fc83659f00634928ec6688f6ef432164563d404ab91d00ec9e89b0

Observation aeb94a3a-7f2b-4238-a29d-50dc08a65f5f · outbound

This paper cites Mantis: Interleaved multi-image instruction tuning, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Mantis: Interleaved multi-image instruction tuning, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.845876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.946618Z digest=sha256:6ff2eda53242c554b0d94180e87fbe355dc1a7356f70b280fe32c68b6e084d17

Observation 11b4e977-6a6e-4ec8-8e66-6549b62d0d4e · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.949814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.949814Z digest=sha256:a57d696efdb637cbb53f1a5d822b080fdb06f68e1bc53bcbe6f8095a80634797

Observation 40be1038-6517-4deb-a45e-396b5470b8e5 · outbound

This paper cites Dense Passage Retrieval for Open-Domain Question Answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Dense Passage Retrieval for Open-Domain Question Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.953157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.953157Z digest=sha256:10dc39cb2a00cdc9c6da2012856eda2760a8aa705d3106434ec48ceb30f21f4f

Observation a2540e4b-15b0-4ffc-b70e-47acf8f5fe44 · outbound

This paper cites Vision-by-Language for Training-Free Compositional Image Retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Vision-by-Language for Training-Free Compositional Image Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.956672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.956672Z digest=sha256:25e10bcda8a5dc8e0b380a2d0a47f01e50ff4a6371d64ddcce76d43e87898b27

Observation 2fefdf49-ab73-4198-b080-6f82ff96ad3f · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

UniCoRN: Unified Commented Retrieval Network with LMMs Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.835520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.960045Z digest=sha256:cfcecd6417889bec960dcb07a26b43e2d5f00e61653a97c6cb72e8ebd7f0f203

Observation 449e2b1a-41d0-4cf8-bf8d-2f5a379d0626 · outbound

This paper cites Grounding language models to images for multimodal in- puts and outputs.

UniCoRN: Unified Commented Retrieval Network with LMMs Grounding language models to images for multimodal in- puts and outputs

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.824825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.963359Z digest=sha256:ee06b5f2cf8808792d724086307cfb94b9e2f65c32e09031aa235dde19aa9814

Observation a4e72049-9e5a-47b4-9670-8f0756296a55 · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.966675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.966675Z digest=sha256:5c11c0caa20da6c4097d4cbb279e345585278618f68e4f53cec5f02057abde23

Observation 24c7f62a-e6d4-4eb0-85a4-49e7e456ee08 · outbound

This paper cites Latent Retrieval for Weakly Supervised Open Domain Question Answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Latent Retrieval for Weakly Supervised Open Domain Question Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.970255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.970255Z digest=sha256:01801b172db109f9b19a4b4f1a1074b23a0c59563d1dcbfa810efab625f7ce5a

Observation a26127f8-216b-4c52-99f4-09b99c72868e · outbound

This paper cites Chatting makes perfect: Chat-based image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Chatting makes perfect: Chat-based image retrieval

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.814878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.974015Z digest=sha256:56ace74fa54d26bcdfabcdfb885536ef4b029f65eebeb75527846a65d51ab835

Observation 707197bd-11e0-4063-856c-5150dd2e5bbf · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.805378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.977277Z digest=sha256:ad0ad5901ca068a468255fef8d9bce9e87a88019f63a4ef7ad9f6f0fa002ac70

Observation c1f5bef1-4759-4615-879e-d60368870870 · outbound

This paper cites Retrieval-augmented genera- tion for knowledge-intensive nlp tasks, 2021.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval-augmented genera- tion for knowledge-intensive nlp tasks, 2021

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.795786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.980488Z digest=sha256:56ee34ba6dc181c0fa8e929d493c55bbf7456c79df498124bf67f4d0f40425fe

Observation 96d48774-0e2e-4239-a0e7-7e03dee3e2a1 · outbound

This paper cites Textbind: Multi-turn interleaved multimodal instruction- following in the wild.

UniCoRN: Unified Commented Retrieval Network with LMMs Textbind: Multi-turn interleaved multimodal instruction- following in the wild

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.785517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.983696Z digest=sha256:61500e1a4789c4cda8d79b83920ee11eafa435585b13b3e12b76cee8d2872169

Observation fd3196f1-a8bc-467e-954c-0fcacac5b9c5 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

UniCoRN: Unified Commented Retrieval Network with LMMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.987047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.987047Z digest=sha256:5e4e6cc1df2d1169a8b74b5a7cf236f87bec593d387d5700863b4c0d44c31e62

Observation 2e4636fd-fcbc-470c-b437-31dd485575b2 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

UniCoRN: Unified Commented Retrieval Network with LMMs Rouge: A package for automatic evaluation of summaries

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.990318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.990318Z digest=sha256:160f3f3003a26e52d605c3c8e8c89147246f15d3cf0d4a8cb390c208656134f5

Observation 591ce93e-7e28-434c-9f98-8987dc53e7b8 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

UniCoRN: Unified Commented Retrieval Network with LMMs MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.993546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.993546Z digest=sha256:52ff72caa4a4e3c5746cacc58f6f3436d34ec0b6298422d4e24dd2555649d703

Observation 7ee84982-63be-4cba-a493-145152d42283 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

UniCoRN: Unified Commented Retrieval Network with LMMs Improved baselines with visual instruction tuning, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.762848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:15.997037Z digest=sha256:be913cbcbba15f14d98992f7f9f03ecb35521622d128ade0ba6f8302cb5a0aff

Observation 88b233ce-38d6-4a2a-9fd8-41b3558cbc82 · outbound

This paper cites Visual instruction tuning, 2023.

UniCoRN: Unified Commented Retrieval Network with LMMs Visual instruction tuning, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.751844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.000680Z digest=sha256:ed7707a6aad1ae5e69d84290668e0af49f9c7379c68fcef6438716dd179d3f68

Observation 5d36c078-a63d-47a0-9cb2-52bd215e595f · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

UniCoRN: Unified Commented Retrieval Network with LMMs Image retrieval on real-life images with pre-trained vision-and-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.739618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.003945Z digest=sha256:3f8e8cbab94a668799ed288b5f22adc6eb6e8219025d9e25ddb5a378e4089b79

Observation cc3c7b49-a857-4c7a-bcc2-088717bdaeae · outbound

This paper cites Image retrieval on real-life images with pre- trained vision-and-language models.

UniCoRN: Unified Commented Retrieval Network with LMMs Image retrieval on real-life images with pre- trained vision-and-language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.007282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.007282Z digest=sha256:6f6db9c9bd0df25051cc648c91662621a7eb70257d132133e65cf04c5282773c

Observation 4bb2271a-8d20-46fd-ae2b-601b1da7987c · outbound

This paper cites Bi-directional training for composed im- age retrieval via text prompt learning.

UniCoRN: Unified Commented Retrieval Network with LMMs Bi-directional training for composed im- age retrieval via text prompt learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.723592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.010587Z digest=sha256:c7d9dadb29cdac6dd01602690d5a35ca6b94908245c40f343f54adfae3abc3d3

Observation 6d68932e-aee0-4a36-9613-52f3a90496fe · outbound

This paper cites Three facets of visual and verbal learners: Cognitive ability, cognitive style, and learning preference.

UniCoRN: Unified Commented Retrieval Network with LMMs Three facets of visual and verbal learners: Cognitive ability, cognitive style, and learning preference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.713434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.013721Z digest=sha256:fe88e1a5475cd09223bed0e1927f7cb0c18a9329bb0c61c9a41471abfcb54dba

Observation e1354ed6-e14f-4de9-8f5e-bbb089595898 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.017042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.017042Z digest=sha256:d0b140aa2259957d17583d7dd5c12def6899b30b9cb2a1c6a5e1af43207b37db

Observation f95c9202-b167-4158-9d88-f3e1a0bdc10d · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

UniCoRN: Unified Commented Retrieval Network with LMMs Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.703498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.020707Z digest=sha256:9192c006117c1917088569d3d6413ea26e1f645f1529264187abd3e403516fbf

Observation 4b4b17d4-a55e-4555-89c1-548ddd5e7829 · outbound

This paper cites Gpt-4 technical report, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Gpt-4 technical report, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.024210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.024210Z digest=sha256:2d0babf2b891276becaa631f2e21cfab6d6ca402b4080dafa2eba125349a1621

Observation 342eb1c4-f58d-44fa-931b-f0ff86bb93fd · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

UniCoRN: Unified Commented Retrieval Network with LMMs Bleu: a method for automatic evaluation of machine translation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.687287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.027601Z digest=sha256:c503450391ad82fd75ed014acbf5a0e6e92679ff24fb08357e91d9e7d318af88

Observation a94d9f92-3caa-43da-9e5f-d4e191272feb · outbound

This paper cites Rora-vlm: Robust retrieval-augmented vision language models, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Rora-vlm: Robust retrieval-augmented vision language models, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.677710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.031009Z digest=sha256:184186c27b363c8425f129876049bdbc768243b3fac94d10a8f70c6bf071843a

Observation 19c69085-513f-47e0-9035-33bc8a69694c · outbound

This paper cites Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation.

UniCoRN: Unified Commented Retrieval Network with LMMs Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.035119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.035119Z digest=sha256:92ba0921d318e80395805efacdcedda20f3978b3feb67fd196b9ac8787055e6b

Observation 7b0afb44-c106-4e9a-8e51-81e36db8fc49 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning Transferable Visual Models From Natural Language Supervision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.038739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.038739Z digest=sha256:1d041d438ce596fc619715721eafae8584246d8c1f1ff84e67857f562e2f8a59

Observation 645a6847-16b2-456f-b901-afd2f19c4813 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

UniCoRN: Unified Commented Retrieval Network with LMMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.042435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.042435Z digest=sha256:e9149279413c93eface34bfbd5b69f17d81d1bd8f27fc3bb4558681ce89ad7a0

Observation 94e1b93c-2782-482f-8579-efdd03bdd4a6 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

UniCoRN: Unified Commented Retrieval Network with LMMs A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.667678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.045951Z digest=sha256:412824375637252e1cf057a9288d3328064571709bd32050fdb3f443ea9e14d0

Observation 084a3a0d-94b5-4e4c-b964-2dc08e73e7bd · outbound

This paper cites Kvqa: Knowledge-aware visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Kvqa: Knowledge-aware visual question answering

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.657548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.049269Z digest=sha256:1627cc7c1757c90912b7ef87ef2214019380e602b7bcce9d9e5883aee55706e4

Observation 5932d71f-7388-4f8e-885e-54ca21648e15 · outbound

This paper cites Towards vqa models that can read.

UniCoRN: Unified Commented Retrieval Network with LMMs Towards vqa models that can read

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.646486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.052474Z digest=sha256:c92578e6078c49dd103c98efdf4f87904c5b7fa9a1af6707f8b2a9f3e62f7cc0

Observation b0f5bc17-34e7-4571-bb41-9805ece0c23a · outbound

This paper cites Knowledge-enhanced dual-stream zero-shot composed im- age retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Knowledge-enhanced dual-stream zero-shot composed im- age retrieval

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.635984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.056213Z digest=sha256:25100cf20adc72c78adb5132013ec176d4cc335f95026744a18354fd0c285819

Observation e5afdb99-5d1e-4271-a2e3-02333f732556 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCoRN: Unified Commented Retrieval Network with LMMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.059617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.059617Z digest=sha256:953893fe3eff368749de67ea5fd6a2fedea240f54ee3bfe011a4be9fa2344921

Observation 17392f1e-64e6-4937-80c0-e3260e4c6c9f · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.625967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.063185Z digest=sha256:bee138db616635ed5fd633fdb2f2fd7a7b1da388407f4da0dd1b825abd3874fb

Observation 918b30d8-a3b9-4390-9cec-7a5a6d442a61 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

UniCoRN: Unified Commented Retrieval Network with LMMs MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.066474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.066474Z digest=sha256:2720c679a4ccbf05f2631f376ad4421bab2dc22cf9666dda2d279b36509c54dd

Observation faeba532-7ee5-473f-a4ec-791afc12438a · outbound

This paper cites Genecis: A benchmark for general conditional image similarity.

UniCoRN: Unified Commented Retrieval Network with LMMs Genecis: A benchmark for general conditional image similarity

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.615682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.069876Z digest=sha256:61d78d7cf7f503257507a235b6d4a628fe3e7938d4dd7c83ea4f2b85a8183e22

Observation c4307a2c-2981-4e25-ac03-72277d72af1e · outbound

This paper cites Composing text and image for image retrieval - an empirical odyssey.

UniCoRN: Unified Commented Retrieval Network with LMMs Composing text and image for image retrieval - an empirical odyssey

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.605302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.073150Z digest=sha256:69c816e38b22173b5c8da0b0eebd765ad60b98c10e2e089fb5ff14feed63800f

Observation 8a47854d-ddaf-4635-ba28-df7b6e540b35 · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

UniCoRN: Unified Commented Retrieval Network with LMMs Composing text and image for image retrieval-an empirical odyssey

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.076600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.076600Z digest=sha256:391a3ceacf449ce36b37ffaa57510bcb0a6fd6f9b01383839bd593153b9effb8

Observation d573bb81-416d-4a88-87cc-ffb22d910204 · outbound

This paper cites Cross-modal feature alignment and fusion for com- posed image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Cross-modal feature alignment and fusion for com- posed image retrieval

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.589090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.079980Z digest=sha256:bf6c8bb3718716c691b92465aa084fb619211c14a6a8db95799d9a2cd9af1f49

Observation 891a069f-594a-423f-a1ca-f90029ac0587 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniCoRN: Unified Commented Retrieval Network with LMMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.083564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.083564Z digest=sha256:5e5f256f75d249e3423275039a2b869d5012ea455de022acb736e6bb6581608e

Observation 182dbbdc-3271-4672-97ca-4db2bebad9f2 · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.578986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.087345Z digest=sha256:63b30ee38e3f2f80cf06eb75a68d0bc32b9e8c5254de170a023711a20ba316ba

Observation 8462c595-82f9-4c73-9e8a-f598044ab3da · outbound

This paper cites UniIR: Training and Benchmarking Universal Multimodal Information Retrievers.

UniCoRN: Unified Commented Retrieval Network with LMMs UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.090827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.090827Z digest=sha256:f652ab52ede725cca5f4c44488610446e8f6c0e80b44fc2da9bd65ea1b8f18af

Observation 7faf818e-0027-4616-b6a6-dba6c7623a02 · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.569127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.094338Z digest=sha256:17633d1020afa93a807de092b08344f05930ed0be55d6275878aa088a3c97d06

Observation f480dfe3-c184-478b-a138-44100814778a · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.558907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.097683Z digest=sha256:8d1e4dad2224e45cac7c64b89533f2a629b0b8122a226cab167fbe9a14d4983e

Observation f981b359-073d-4059-9f45-42a1fba8df2e · outbound

This paper cites Visual question answer- ing: A survey of methods and datasets.

UniCoRN: Unified Commented Retrieval Network with LMMs Visual question answer- ing: A survey of methods and datasets

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.101159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.101159Z digest=sha256:705d4044f324c950899144aa02288c42ae44cd452430c5b4a58ee2bac81584a5

Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.104506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.104506Z digest=sha256:ea31d479a81915d27d8ecb8753e47c0c3b8296031acbd15cb31921e31ec1fded

Observation 22ec78be-1e28-48a4-af82-7030b566ab26 · outbound

This paper cites EchoSight: Advancing visual- language models with Wiki knowledge.

UniCoRN: Unified Commented Retrieval Network with LMMs EchoSight: Advancing visual- language models with Wiki knowledge

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.542820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.108201Z digest=sha256:4cc998958c404ce3434845735490f5e41551cca683ff0c50174acfe37fc3b7c8

Observation 44e175a0-ec5e-4e91-9d78-8d43d48b7cc9 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

UniCoRN: Unified Commented Retrieval Network with LMMs FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.111778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.111778Z digest=sha256:23cbe9f5b7a4cd5626f85a3df2064df5735a7faabe8c0c153e9730d57017c1b1

Observation f3622b38-edbd-4faf-8681-7c8a7e89af82 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

UniCoRN: Unified Commented Retrieval Network with LMMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.115407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.115407Z digest=sha256:49f55b89930a961d39941c56e0045c85f1e631a908decd8ca3d06d323eebdfbf

Observation 962fba5b-9b5b-4d40-b138-8a495af5dd8b · outbound

This paper cites A Survey on Multimodal Large Language Models.

UniCoRN: Unified Commented Retrieval Network with LMMs A Survey on Multimodal Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.119075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.119075Z digest=sha256:8f3488649dc56bb7b410ac5cea8fa6f1792328718ec1f67bf1109df33d320781

Observation 43394a8b-a74f-4e12-ad5a-39f6911df660 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

UniCoRN: Unified Commented Retrieval Network with LMMs CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.122413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.122413Z digest=sha256:42a1e37fc883ce8af051b699a50f09679067e1eaa7ffcf72ff959f42335156ed

Observation 00b79990-a65d-40ee-9b7d-cc78b3d733dc · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniCoRN: Unified Commented Retrieval Network with LMMs Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.126629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.126629Z digest=sha256:9c59db9b7fdbde9cb9761013347435c09882b2fcf6b8090c0e67c2b3d97f64e4

Observation 3382d6d5-1c24-40b4-a9db-3d32e9e0b111 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

UniCoRN: Unified Commented Retrieval Network with LMMs VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.130352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.130352Z digest=sha256:e07ac3173e90236f1e2125230fa022adfae51e65d02d726826b73be2ba143fe5

Observation 5ebec1ff-b09f-4a7e-ba09-4104836ad76b · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

UniCoRN: Unified Commented Retrieval Network with LMMs A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.134206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.134206Z digest=sha256:e434973b0725e3ead0197c6264476e25289106c7e21d50c50133233a73127c49

Observation a892aaff-44e9-478b-91be-e937e72ad97f · outbound

This paper cites Sigmoid loss for language image pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs Sigmoid loss for language image pre-training

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.532426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.137913Z digest=sha256:6fdbf3cc46aa6d811c873086d0d5cc490f2caa1ba7598561b6d3e8ee291f925b

Observation 151d9cde-c624-4470-b210-bd4e414baa0f · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

UniCoRN: Unified Commented Retrieval Network with LMMs MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.141327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.141327Z digest=sha256:d0c6b6dfc981626503a0b26c2a825969388cf0c5845b0e10471aab698b637096

Observation a4528704-16cb-4ea1-b256-d898499e5c46 · outbound

This paper cites Non-Contrastive Learning Meets Language-Image Pre-Training.

UniCoRN: Unified Commented Retrieval Network with LMMs Non-Contrastive Learning Meets Language-Image Pre-Training

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:16.184255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.145165Z digest=sha256:e4f8ad5bd9a90d05c236c7ce55f64d8e672e6e32dc848650b5a38f7eb8e3a2eb

Observation e6957db2-f559-4b6d-bc2e-ba3a7e2a2f08 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

UniCoRN: Unified Commented Retrieval Network with LMMs MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.522347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T05:54:16.148845Z digest=sha256:38560a2c19511efc1ab7652a8d56ad963b66baccabc73b37f8841b31fd4c1908

Pith citing papers

No inbound Pith citation observations are available.