Pith. sign in

Paper Citation Record · LEDGER

Illuminating Visual Identity in Universal Multimodal Embeddings

As of 9 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.01794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01794 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:43:17.507439Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved55
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2363b7d3-81bd-4afc-a3da-8fe99693d3db · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Illuminating Visual Identity in Universal Multimodal Embeddings Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.346553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.346553Z digest=sha256:da6b5c1610d00dc8f22611402d9e83cc0b6a375be08eccb9742b8c72d2b389b1

Observation bd4f6e0e-b313-4b2c-89cf-b70b0de3e3dc · outbound

This paper cites Unicom: Universal and Compact Representation Learning for Image Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Unicom: Universal and Compact Representation Learning for Image Retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.403233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.403233Z digest=sha256:093a99a0cf3f79ba2ee93de8622e6ea86a418aff57acbc5de75bdb3423f7d832

Observation 88636538-b85b-48d0-8d11-55c14ce56b5b · outbound

This paper cites Qwen2.5-VL Technical Report.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.493030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.493030Z digest=sha256:98b216c12b65c9401e3af19cafd0f211e1ab27eb1a4929c1d6a48ed506610bd0

Observation 6dd964af-ecc0-4d72-b6a0-0b501e81ebdb · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Illuminating Visual Identity in Universal Multimodal Embeddings LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.611224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.611224Z digest=sha256:728c97e742c56a4023d5cede63824eeae96c7f49d89a815549e28f9eb77ab635

Observation a64de592-e3dc-479b-8d0b-49625f0cf298 · outbound

This paper cites Flame: Frozen large language models enable data-efficient language-image pre- training.

Illuminating Visual Identity in Universal Multimodal Embeddings Flame: Frozen large language models enable data-efficient language-image pre- training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.727986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.727986Z digest=sha256:6ce72fb8ccfa44654b21781b7b94f7e224098e7ee57e55218e2f02ec89d89e97

Observation 9e7d481e-e174-4d49-a57f-18746709650d · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Illuminating Visual Identity in Universal Multimodal Embeddings mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.812234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.812234Z digest=sha256:60dd3565ac5b8aae0f1b55bb64ef1260f719a677d0c4ce30e53384be779769de

Observation ce543fff-5dde-46e1-bc81-2031f0eb2874 · outbound

This paper cites Murag: Multimodal retrieval-augmented generator for open question answering over images and text.

Illuminating Visual Identity in Universal Multimodal Embeddings Murag: Multimodal retrieval-augmented generator for open question answering over images and text

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.932581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.932581Z digest=sha256:ff626aee0247b50317d83e3ec9bc9c501bd34f5ca35873576f94a0bb2a3fc8bf

Observation c0ab3143-6da5-4689-99ea-34a6d71ff4a8 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Illuminating Visual Identity in Universal Multimodal Embeddings Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.058096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.058096Z digest=sha256:988096a6a713f31624d43c259f7694adef5c498165b60f7e845583b3eef6c20e

Observation e2d25ea6-cd73-4d60-9e7a-3ef01090b387 · outbound

This paper cites Opengpt-4o-image: A com- prehensive dataset for advanced image generation and edit- ing.arXiv preprint arXiv:2509.24900, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Opengpt-4o-image: A com- prehensive dataset for advanced image generation and edit- ing.arXiv preprint arXiv:2509.24900, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.225875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.225875Z digest=sha256:cc712f0b07c1b1453a6a114352432977b7d7e2132d5d6ee628ba6065879d24ea

Observation ced016ae-bb8a-4efa-8861-ef6d4e4a0e19 · outbound

This paper cites Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014,.

Illuminating Visual Identity in Universal Multimodal Embeddings Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.394360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.394360Z digest=sha256:99b829d72fc2e82d4bd451b9238564f85849f6826311c6f48238108067b1a5c6

Observation 1503431a-f199-4920-89fd-2aa1bdbb0418 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Arcface: Additive angular margin loss for deep face recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.494829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.494829Z digest=sha256:6a192184d74f7d63acf13e548b137a7489d3c5ee24291cf56f857af2175085c3

Observation 84c5eb15-19db-4409-9a10-8ec3f53b88d7 · outbound

This paper cites Efficient and discriminative image feature extrac- tion for universal image retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Efficient and discriminative image feature extrac- tion for universal image retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.665447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.665447Z digest=sha256:8924f3541dca87f73c47868e690059f891e4f58575fe7157066e9f067a359e9a

Observation 6675d07d-4f80-441a-9de8-552a76f22541 · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Illuminating Visual Identity in Universal Multimodal Embeddings DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.825424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.825424Z digest=sha256:060131b73e0358904d5bcd9181f19b6b9a4ee5032a424222e5fe8250e11634cb

Observation 3e1ac40a-362b-4648-b680-886b4dcee7b4 · outbound

This paper cites Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup.

Illuminating Visual Identity in Universal Multimodal Embeddings Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.991222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.991222Z digest=sha256:344c0411d17bad7f448479ab12e951a2b39d297240723fe0641fba6b71dc854b

Observation 036f4965-2c3d-44eb-95c6-f35eb5cbe0a1 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Illuminating Visual Identity in Universal Multimodal Embeddings SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.155373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.155373Z digest=sha256:ba9d644cb1dce26e497ea3f806485725cdfc290e8ce9b03e092a78c6226c9c1f

Observation a2d55248-5423-4de2-bd52-c4b22dd1ce2b · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.239595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.239595Z digest=sha256:28fddac6b36ac6bff707e3c808193d1353fccb0f228e3422bd61ce992187eb06

Observation 809043bb-8bd1-4a11-8630-d35cce3060a4 · outbound

This paper cites Unime-v2: Mllm-as-a-judge for uni- versal multimodal embedding learning.arXiv preprint arXiv:2510.13515, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Unime-v2: Mllm-as-a-judge for uni- versal multimodal embedding learning.arXiv preprint arXiv:2510.13515, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.281060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.281060Z digest=sha256:aabb7cec3c08f0db174dc38e56cb7ab7d8883901b365ddcd3e872a496df88cb4

Observation c6a7724b-16fd-42e0-8e92-c11485f0cb06 · outbound

This paper cites jina-embeddings-v4: Universal embeddings for multimodal multilingual retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-embeddings-v4: Universal embeddings for multimodal multilingual retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.356047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.356047Z digest=sha256:125cf18a28df7801df5f71371386e870564fe0077ec13a7540b662fc3587cded

Observation 44a9300e-5520-4cee-ae65-51ce664fb9c2 · outbound

This paper cites Ms-celeb-1m: A dataset and benchmark for large-scale face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Ms-celeb-1m: A dataset and benchmark for large-scale face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.495874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.495874Z digest=sha256:089a9ed801643d61baaa744ad90a0c123efdf1c0b411c659b73b26e0bc45d4d8

Observation f71a18a8-5f8a-4d10-ae47-c58f4fb9ad44 · outbound

This paper cites In Defense of the Triplet Loss for Person Re-Identification.

Illuminating Visual Identity in Universal Multimodal Embeddings In Defense of the Triplet Loss for Person Re-Identification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.565316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.565316Z digest=sha256:1905e89cc803c39b16fd898dc43e3f636f8f26502f60c2fafd680843731faed8

Observation 07e17851-bee6-46b0-aa54-affa8b247483 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Illuminating Visual Identity in Universal Multimodal Embeddings Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.678112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.678112Z digest=sha256:8e5a17bd1511a8d86b630096f43079d42cc3a6c7de020b8e1624ac39bd8af491

Observation 31614ed6-7a1a-41c3-bc9d-bc694a947a15 · outbound

This paper cites Llm2clip: Powerful language model unlocks richer visual representation.arXiv preprint arXiv:2411.04997, 2024.

Illuminating Visual Identity in Universal Multimodal Embeddings Llm2clip: Powerful language model unlocks richer visual representation.arXiv preprint arXiv:2411.04997, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.751502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.751502Z digest=sha256:571ac676588471e28cbc33d458936d1293c0bb620c27c4187da157ea032c07bf

Observation 7e5169f5-85b8-4aad-a134-ca2008b0f153 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Illuminating Visual Identity in Universal Multimodal Embeddings Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.861413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.861413Z digest=sha256:f143fbb1ded97a0ad1abdf56e377fce40ed166ea3377a29b9210c65bf2be55b7

Observation aec1c32d-63af-40d8-a722-e6794830ce73 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Illuminating Visual Identity in Universal Multimodal Embeddings E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.906861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.906861Z digest=sha256:1ef826abd29c81b6ecc7e80c7f9842c4ff6666b0d2f3281b30cb445291f0bef9

Observation 8c618cf0-6d98-493e-af7c-dc803f18a35b · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

Illuminating Visual Identity in Universal Multimodal Embeddings Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.560698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:16.006633Z digest=sha256:9e00b70fbd446ac22a2cdad62efdaca48b34efb80bd627dbdc98b8459e0a3ec2

Observation 68e4c1a0-4d0d-436a-b85c-8bffd421c7ac · outbound

This paper cites Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.106644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.106644Z digest=sha256:6db115af74fbe718b7f4c2aac7dfbb6cbd03f61d18b6b3c10bca4a59da55761c

Observation 69f665ef-4a45-4d80-a72b-0487e08fa2e7 · outbound

This paper cites Ilias: Instance-level image retrieval at scale.

Illuminating Visual Identity in Universal Multimodal Embeddings Ilias: Instance-level image retrieval at scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.326487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:16.217922Z digest=sha256:00aea07235a35c5c90203d1e94789ad512af900d58c809daae6b15b29075e27d

Observation 73ae2735-f4c0-4927-ba64-3ea619c019e7 · outbound

This paper cites jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.254863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.254863Z digest=sha256:327b9325701a2c25374537ee8d0b0bd90eebbeb351a45ca31dcf22c95c86c407

Observation a65a327a-6dda-4007-b95c-ae815c88f5d0 · outbound

This paper cites 3d object representations for fine-grained categorization.

Illuminating Visual Identity in Universal Multimodal Embeddings 3d object representations for fine-grained categorization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.112459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:16.403310Z digest=sha256:8acca06a42e4905086d518e9e9a8f2f8fa54ba2bcf4e1ed4a493e16011502a5c

Observation 4d378f39-698e-463e-b460-d32df3b19ecd · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.514852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.514852Z digest=sha256:f6dfd3bc88a648eee0c324f4122bccd46dd2d07c5f8686ae1b08812bf265e2dd

Observation 8fdaf9ed-6939-417d-b030-3aef457a9870 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Illuminating Visual Identity in Universal Multimodal Embeddings NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.588928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.588928Z digest=sha256:038965e42f7d146c3b4fb567a75428122ea18b6508326fd909334ac823d6ee70

Observation c5aa0e80-d869-447d-9cf1-3cc72948598d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Illuminating Visual Identity in Universal Multimodal Embeddings LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.658249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.658249Z digest=sha256:be4b5a9849d9a985b3b0c5e0fe4e5af7dffaed08b44d55700f2da906c4bfb1e4

Observation b1c7a46d-5e3a-4422-ba4e-2b86744e564a · outbound

This paper cites MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization.

Illuminating Visual Identity in Universal Multimodal Embeddings MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:43:17.744970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:16.771390Z digest=sha256:37a1393c61e3171556da3c525ac09753d62930a1ef77ff8330bdf6074fa20b92

Observation 6dc4ce12-4e7a-495b-8c25-13ae83d0a9db · outbound

This paper cites Personalvideo: High id-fidelity video customization without dynamic and semantic degradation.

Illuminating Visual Identity in Universal Multimodal Embeddings Personalvideo: High id-fidelity video customization without dynamic and semantic degradation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.903587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:16.813054Z digest=sha256:dda65f27b46933f98a07157b428c9b330370243e9f1335462fe765125173b32a

Observation 77dbe72f-dc1e-4ac4-ae64-923d92c34748 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Illuminating Visual Identity in Universal Multimodal Embeddings Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.848115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.848115Z digest=sha256:96b99db4c2b63829e5cf1b450c21b73539d96ebc920036d58c1b745c0a8842de

Observation 20c6ace6-e135-4bb5-a263-204f179e002a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Illuminating Visual Identity in Universal Multimodal Embeddings Open-Sora Plan: Open-Source Large Video Generation Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.957129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.957129Z digest=sha256:e388147ed1cc4cd6640a29b85ad1a0914aaecb523e90226c5c93aca2f21cdad9

Observation f6134962-d0e8-4c11-8465-dabdb4092ea0 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Illuminating Visual Identity in Universal Multimodal Embeddings UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.032612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.032612Z digest=sha256:f85806c4528314c4a6e3e2e1e3dc6ae0987468f092e69151dadab2384a972653

Observation 55759e70-7d51-48d1-ae24-01e839854cf1 · outbound

This paper cites Mm-embed: Universal multimodal retrieval with multimodal llms.

Illuminating Visual Identity in Universal Multimodal Embeddings Mm-embed: Universal multimodal retrieval with multimodal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.674734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.102764Z digest=sha256:b542eb67c3bf58ace13d40f93f5eb20330406dd270ed2cb673916e66b53e2fdd

Observation 97524d2e-81b4-4eaf-b53f-630d677d6522 · outbound

This paper cites IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.179548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.179548Z digest=sha256:30e0ed0c6d173232e1df4842fb65319321d9513a69c06f3f4b4a49001d0744ae

Observation 2c0a19dd-6584-43e3-8e78-f2306bde46d5 · outbound

This paper cites Automatic synthetic data and fine- grained adaptive feature alignment for composed person re- trieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Automatic synthetic data and fine- grained adaptive feature alignment for composed person re- trieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.464572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.257562Z digest=sha256:4c38b2f9529d055ae0913543c7f049089b450667f97cb198d32c2d2711214da4

Observation 9f181751-f6ec-46ee-8eeb-b9440fc26489 · outbound

This paper cites Sphereface: Deep hypersphere embedding for face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Sphereface: Deep hypersphere embedding for face recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.260553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.260553Z digest=sha256:327673f810730ff950e5d393dbbd381e86e78b32611e2d03afc9e8b55aa0de78

Observation 775382e5-b8cc-4a14-a21b-127769462100 · outbound

This paper cites Large- scale vehicle re-identification in urban surveillance videos.

Illuminating Visual Identity in Universal Multimodal Embeddings Large- scale vehicle re-identification in urban surveillance videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.260878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.264145Z digest=sha256:400b3e7df8b6d8f3d447f2a3b14dc19793288935536f726d46a7e7b41760a19b

Observation bcfdc93d-0712-4382-8ac6-8d4129837f68 · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.

Illuminating Visual Identity in Universal Multimodal Embeddings Lamra: Large multimodal model as your advanced retrieval assistant

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.013085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.267112Z digest=sha256:eaf33841d744c2901b2abb7c630b78fc949df6d6ac38d7ceac9adce184287c7c

Observation 45cf1584-619d-4bbf-bf45-0b149e00568f · outbound

This paper cites Deepfashion: Powering robust clothes recognition and retrieval with rich annotations.

Illuminating Visual Identity in Universal Multimodal Embeddings Deepfashion: Powering robust clothes recognition and retrieval with rich annotations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.814995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.270047Z digest=sha256:2548a3fc42694c32912d15957d30ac93506a750cc28ae9d40304ebb1e23b55eb

Observation 8091a9df-dbd6-42d8-81e4-7712ab156d92 · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

Illuminating Visual Identity in Universal Multimodal Embeddings VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.273022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.273022Z digest=sha256:135bf32f38103964e487f0aaa60cf691eb6adc3a67b44362d56eccbf0c79e723

Observation d6eef85d-e283-41e8-8d54-c2db472faef7 · outbound

This paper cites A culturally-aware benchmark for person re-identification in modest attire.Engineering Ap- plications of Artificial Intelligence, 158:111494, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings A culturally-aware benchmark for person re-identification in modest attire.Engineering Ap- plications of Artificial Intelligence, 158:111494, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.638218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.276038Z digest=sha256:6fdf37f4574ae3aec7688a2921c4a63cba01daaaa832a1851317ee7b590a1e61

Observation 421feeaf-c31b-42ed-a906-f54525ffb7e8 · outbound

This paper cites Deep metric learning via lifted structured fea- ture embedding.

Illuminating Visual Identity in Universal Multimodal Embeddings Deep metric learning via lifted structured fea- ture embedding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.392243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.281806Z digest=sha256:98e59779c877ef84d37c1a17756d3850fdf5a08fcf8747b9839f9909d5c6dd08

Observation 5898f116-c7ce-4c52-b079-08ad080bd6a7 · outbound

This paper cites RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification.

Illuminating Visual Identity in Universal Multimodal Embeddings RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.284647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.284647Z digest=sha256:9d3dd3b21b4d7b6c3c325d9ef8e6345fb4c8558f22f1c7cbce5399dc07994b9e

Observation 4eaf0447-583f-46d7-8672-60cb21b66218 · outbound

This paper cites Revisiting oxford and paris: Large-scale image retrieval benchmarking.

Illuminating Visual Identity in Universal Multimodal Embeddings Revisiting oxford and paris: Large-scale image retrieval benchmarking

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.266199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.288092Z digest=sha256:112f1868f6e6172422c38ab41f7e7e0caeb18f626a2d230da664355a19780448

Observation 624f29d7-4559-42a1-9e5c-1285ade0939e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Illuminating Visual Identity in Universal Multimodal Embeddings Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.290997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.290997Z digest=sha256:34f0f73207f53b1fa1ad83125eaf5c805b04215f5d4f4681b7b919d07295712d

Observation e24db9c0-0477-404c-b96c-4fa04bdd4ca2 · outbound

This paper cites Performance measures and a data set for multi-target, multi-camera tracking.

Illuminating Visual Identity in Universal Multimodal Embeddings Performance measures and a data set for multi-target, multi-camera tracking

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.078583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.293921Z digest=sha256:fa23931f1f80fdd6a7d005740de78fcb0404fe053eb223eb035b57df8280f12d

Observation 3aa65bd2-642e-44e9-9f56-f9c3284cf5c8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Illuminating Visual Identity in Universal Multimodal Embeddings EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.297183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.297183Z digest=sha256:3d3920646e1adc9401dd0971aeda2c3b9659cdfb011e6ad8d6124673d6595c90

Observation 51c43246-bcb5-42d8-85e2-138c2169789b · outbound

This paper cites Visual named entity linking: A new dataset and a baseline.

Illuminating Visual Identity in Universal Multimodal Embeddings Visual named entity linking: A new dataset and a baseline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.950019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.300330Z digest=sha256:20b098a90a21e845e6c35841f3d7316742ad352625bb4cb66e33d3cc3896734a

Observation 9297161b-4cfe-43c4-8db0-90ab72749cd6 · outbound

This paper cites Breaking the batch barrier (b3) of contrastive learning via smart batch mining.arXiv preprint arXiv:2505.11293, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Breaking the batch barrier (b3) of contrastive learning via smart batch mining.arXiv preprint arXiv:2505.11293, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.303217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.303217Z digest=sha256:55cdb99f67d2690950f69d3dfb2c3ef8d3ca51783c87de28088e7e47252df33e

Observation 7b175f18-ce65-49f6-b80d-f97439b0ad29 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Illuminating Visual Identity in Universal Multimodal Embeddings SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.306059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.306059Z digest=sha256:89e6048f6e90709c45cae8cb97fb3f0e726ef7a0bf1d2077619a6e95685be131

Observation 610fae48-b01c-4259-8dc5-e769b16bfb37 · outbound

This paper cites The inaturalist species classification and de- tection dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings The inaturalist species classification and de- tection dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.309133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.309133Z digest=sha256:4dd2934184a26df151639b15433c3266bb8057433be4d111453ce777dc38f69b

Observation f4055274-555e-4706-a121-2c528e67943e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.312239Z digest=sha256:eeb78d5ae134c40c437ea05940bb7da217bf4633d50aba539e530648ccdc633f

Observation 60351181-dd15-4ad5-9198-04f7cc56be39 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Illuminating Visual Identity in Universal Multimodal Embeddings InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.315397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.315397Z digest=sha256:92ac38d351b5bc89d30b5b2e5916ee0f7bcd914e0543dd1f0481bf1d785cd695

Observation e875a0a6-b8b5-442d-8947-11f2f9821327 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.318655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.318655Z digest=sha256:437567545efd34233e1ecab7563fb65526f6fd41f0f32a8ddbd2bcda7be30eef

Observation 932e528f-e291-479a-afd3-bda7ba484c26 · outbound

This paper cites Uniir: Train- ing and benchmarking universal multimodal information re- trievers.

Illuminating Visual Identity in Universal Multimodal Embeddings Uniir: Train- ing and benchmarking universal multimodal information re- trievers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.787105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.322048Z digest=sha256:158b790498b9875b168d463b511c85429d33c3d6f39d0792e1ec8f28ccb09a19

Observation 6755346c-bcb1-4235-8fe7-e6c08722c841 · outbound

This paper cites Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.613019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.324853Z digest=sha256:ea6535c6c124fadbc1788b06dfe3a594ea9a6d1df4ec403cf96422d05312f975

Observation 99e89c04-a22a-4938-bea5-2df5c91abf9e · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

Illuminating Visual Identity in Universal Multimodal Embeddings Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.467311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.327787Z digest=sha256:043146cc77ffd7c8504644bc2c60d1925163dfc2b6cd039c08790f891e4c0743

Observation d03ce222-b0c5-458c-aa79-c6143ef507e5 · outbound

This paper cites Forb: a flat object retrieval benchmark for universal im- age embedding.Advances in Neural Information Processing Systems, 36:25448–25460, 2023.

Illuminating Visual Identity in Universal Multimodal Embeddings Forb: a flat object retrieval benchmark for universal im- age embedding.Advances in Neural Information Processing Systems, 36:25448–25460, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.366074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.330652Z digest=sha256:eb59a0133c6b22a98440049251a874021d6c2424c5f01e1d775248ebe7c30406

Observation 46449807-f27c-469a-8b8c-f762153cc6e0 · outbound

This paper cites Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.333489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.333489Z digest=sha256:5eb327fa2834373e46e3af99f6128ec04796873e8c90d55713e9d66bacdf3ede

Observation d2e52caf-a117-4d20-864f-ee1847e6c1db · outbound

This paper cites Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying.

Illuminating Visual Identity in Universal Multimodal Embeddings Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.336608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.336608Z digest=sha256:13356f1effa91f0bd6087daf8f986dce67aae6970943b23d88c986137f17b70d

Observation 87e9efb2-e9f0-4622-8961-558d1fdb1567 · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Illuminating Visual Identity in Universal Multimodal Embeddings Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.339889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.339889Z digest=sha256:8e7e17ba33a0989293367fdcfc1c45c527efe1108938b3d5609289d941b1f667

Observation 7af4341d-dbac-4eda-b7b4-20a07289fc47 · outbound

This paper cites A large-scale car dataset for fine-grained categorization and verification.

Illuminating Visual Identity in Universal Multimodal Embeddings A large-scale car dataset for fine-grained categorization and verification

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.229304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.342861Z digest=sha256:fb23f3108e57b6508d4bdfeb228439a64e584f4abab3daee5516172f2dbf19c7

Observation e90943c0-354b-48fb-82e6-fcba2ef6a499 · outbound

This paper cites Deep learning for person re- identification: A survey and outlook.IEEE transactions on pattern analysis and machine intelligence, 44(6):2872–2893,.

Illuminating Visual Identity in Universal Multimodal Embeddings Deep learning for person re- identification: A survey and outlook.IEEE transactions on pattern analysis and machine intelligence, 44(6):2872–2893,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.085925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.345722Z digest=sha256:7981f324557066b5e140cd266ef3d066a067d32e18bf6c5b3a046e2e25b6c74a

Observation 7cd99d99-95d9-4e89-a6ef-3f036daaf510 · outbound

This paper cites Learning Face Representation from Scratch.

Illuminating Visual Identity in Universal Multimodal Embeddings Learning Face Representation from Scratch

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.348652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.348652Z digest=sha256:911a174b30cab632b1368014a3ee0e7c7817d9e1ae413f829a4eda92e358d0fb

Observation ac953921-ea05-4cee-a1ac-eed9274e9d03 · outbound

This paper cites The met dataset: Instance-level recognition for artworks.

Illuminating Visual Identity in Universal Multimodal Embeddings The met dataset: Instance-level recognition for artworks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.941829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.351877Z digest=sha256:3d90074b3b1d717b41fa1f10763010ab86489b0952b9f567581b079fd97670e0

Observation 2bd6fb3a-74be-4cc9-a278-c3e40dd7f15b · outbound

This paper cites Towards universal image embed- dings: A large-scale dataset and challenge for generic image representations.

Illuminating Visual Identity in Universal Multimodal Embeddings Towards universal image embed- dings: A large-scale dataset and challenge for generic image representations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.824469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.354680Z digest=sha256:ae2d1041e0b7b200aa5952b15f255c6bc887b9d893ac694f8cd9403958d47cb1

Observation 9cbd0a3e-e878-45fd-857a-ffbc63cac070 · outbound

This paper cites Udon: Universal dynamic online distilla- tion for generic image representations.Advances in Neural Information Processing Systems, 37:86836–86859, 2024.

Illuminating Visual Identity in Universal Multimodal Embeddings Udon: Universal dynamic online distilla- tion for generic image representations.Advances in Neural Information Processing Systems, 37:86836–86859, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.694822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.357587Z digest=sha256:2b956d9337d350e8ba4bf4db4dacbc32f685651a8b940ff50277c317a787901a

Observation 931fdf07-6b8d-4b55-98f8-1f24ea3029e9 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Illuminating Visual Identity in Universal Multimodal Embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.360998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.360998Z digest=sha256:2c7536075827aa1bdb3cf7560641213f417e052c3ec9ea9171dd13242e9c9fc9

Observation 8178b8ee-26cc-4913-9162-c5f2d8ad71e5 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

Illuminating Visual Identity in Universal Multimodal Embeddings VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.364610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.364610Z digest=sha256:80cea5e9b9b92240e6fb11d92ba06f6bfca28a2da3d2fa65044d0473d9017c0b

Observation 6c629151-d8cf-4adc-8784-379eba4fed65 · outbound

This paper cites Identity- preserving text-to-video generation by frequency decompo- sition.

Illuminating Visual Identity in Universal Multimodal Embeddings Identity- preserving text-to-video generation by frequency decompo- sition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.558554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.367779Z digest=sha256:490d7b1357ba6f3ae0e392eeea625a17f5c10bc8762a48aab585a004a5e001e5

Observation 40690f5a-428d-4b23-8919-477e13ee1eb5 · outbound

This paper cites Sigmoid loss for language image pre-training.

Illuminating Visual Identity in Universal Multimodal Embeddings Sigmoid loss for language image pre-training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.370533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.370533Z digest=sha256:5d3056074211a8a4f32b154cc8f8982930d44e3293444dbdfff91781f196f7c3

Observation b18f833e-7fe4-4b36-9661-ab9b29429545 · outbound

This paper cites Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining.

Illuminating Visual Identity in Universal Multimodal Embeddings Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.413590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.373366Z digest=sha256:1db7b76d2879451fd963926d9082f58a40f171bc34fd1cbb504a3444795cd99a

Observation 5fd02a67-b5ed-461a-99b7-0bc1bf3d7fdd · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

Illuminating Visual Identity in Universal Multimodal Embeddings MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.376305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.376305Z digest=sha256:44b55d5be5f0dde31e90e81c3b4564f17a5ce40b49b875a1133db0243cd223ee

Observation 9fd7f93b-194a-406e-b802-404beba054a2 · outbound

This paper cites Assess- ing and learning alignment of unimodal vision and language models.

Illuminating Visual Identity in Universal Multimodal Embeddings Assess- ing and learning alignment of unimodal vision and language models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.379454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.379454Z digest=sha256:c1cc50d608bd38ef6ed1b46a6fe88f5f307c8f12c2734fdf438ec348438fa284

Observation 1dde37f5-d3b3-4440-9a43-fc552cb4ac59 · outbound

This paper cites Beyond frontal faces: Improving per- son recognition using multiple cues.

Illuminating Visual Identity in Universal Multimodal Embeddings Beyond frontal faces: Improving per- son recognition using multiple cues

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.275526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.382385Z digest=sha256:b7131e442be21dc1c648e5eded4541bd4c7c68ce66cc39119430a632dffbced4

Observation a68bac3d-d1c0-410e-b9bb-f932ca113872 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

Illuminating Visual Identity in Universal Multimodal Embeddings GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.485401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.485401Z digest=sha256:9ec9cd281e54894143ff4332339590e983da8b0d48049314d093fbc4da78f606

Observation 1ea36a60-e6aa-4864-a5f7-160629ce6583 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.489048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.489048Z digest=sha256:ec2c66ca37c751b28b215dc3daf9dec7861031525eb48d763f9646c91ab84a3a

Observation e18f69cb-8fdc-4673-917f-1354f340451b · outbound

This paper cites Magicmirror: Id-preserved video generation in video diffusion transformers.

Illuminating Visual Identity in Universal Multimodal Embeddings Magicmirror: Id-preserved video generation in video diffusion transformers

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.119414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.492298Z digest=sha256:fd7ba4baa2d3308a282e63ee771689a2eaa7943a970688ccb60ff64038f6feff

Observation 0d85fe8b-7bc2-406a-87ae-95816f61eebe · outbound

This paper cites Scalable person re-identification: A benchmark.

Illuminating Visual Identity in Universal Multimodal Embeddings Scalable person re-identification: A benchmark

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.035923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.495323Z digest=sha256:b1b6aa5fe865b27d73d6f900c1c957d3b090eac79e24fb6313289917f585fa22

Observation da79ef23-b212-4b4f-999e-de3b37100501 · outbound

This paper cites Scalable person re-identification: A benchmark.

Illuminating Visual Identity in Universal Multimodal Embeddings Scalable person re-identification: A benchmark

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.010772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.498322Z digest=sha256:f6ada5668b2c3fcf6960ad8f5ba8b0ab91c8df719473e502d2cd7eb7e96097db

Observation 10ef6a1b-c816-419f-b476-c14682660d44 · outbound

This paper cites Cartoon face recognition: A bench- mark dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings Cartoon face recognition: A bench- mark dataset

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:17.983167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.501286Z digest=sha256:3be13da48c7f6ad6f619dd526ff165aae861f9ba5332ee4c4c23b67f8bf1f829

Observation 47e166c9-8827-4454-9a0a-97499ad8221c · outbound

This paper cites Megapairs: Massive data synthesis for universal multi- modal retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Megapairs: Massive data synthesis for universal multi- modal retrieval

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:17.963805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:43:17.504442Z digest=sha256:3e8c55efdd9b86135fff8d6223ac8ea4f5b442fd347d0676e6093dfd072eb92b

Observation 7f78a490-0d0c-41d6-bd8c-2321a9113e37 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Illuminating Visual Identity in Universal Multimodal Embeddings InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 89

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:43:17.507439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.507439Z digest=sha256:c6cc67b66dd333f6366211e44f0a85376eef2169e88186dbee4d37af824e015b

Pith citing papers

No inbound Pith citation observations are available.