Pith. sign in

Paper Citation Record · LEDGER

PLUME: Latent Reasoning Based Universal Multimodal Embedding

As of 4 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 3 inbound Pith citation observations for arXiv:2604.02073.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.02073 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T21:48:40.722921Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:35:21.372512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T18:03:36.933948Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact30
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 027c6af3-c73b-4517-b32d-bf11d5ae2e4d · outbound

This paper cites Qwen2.5-VL Technical Report.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:20.018342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:bd7bfe1de6aba560128f9ff64559d089792c4b2e23291d33d324d6c6159a4225

Observation 461f082c-7575-4ae2-890b-803131c07d9e · outbound

This paper cites Curriculum learning.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Curriculum learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.153236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:3da4d482b0221f389574097b22c8649feb7981a893b9d5dbad3f9c0974a85df9

Observation 2346178e-b092-41c7-a44a-ce6802d66eb7 · outbound

This paper cites MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings.

PLUME: Latent Reasoning Based Universal Multimodal Embedding MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.962706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:1e99426d9335b2947e790f5c239844666a86ab8fa48ff02354dcd639d9ef0e68

Observation badc0a05-4088-4815-9d78-5fb4439f7e6d · outbound

This paper cites Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025a.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025a

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.013498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:5c45a7b41c8156bc8c28922952b66d0ac2eaf2f92538e07983fd0adaf6cf744e

Observation 6b92bb5f-be49-41bb-96d6-f23ce6f2d9ab · outbound

This paper cites arXiv preprint arXiv:2510.05014 , year=.

PLUME: Latent Reasoning Based Universal Multimodal Embedding arXiv preprint arXiv:2510.05014 , year=

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.034261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:afb06de46b67d19b4e27997b63fbb7a8b419394f3c14b887365b7b325df9fcda

Observation d1329f7f-cdc7-4682-8bdd-f69b80cb800f · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding ColPali: Efficient Document Retrieval with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:35:41.289440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:5d6d4bcd03b6d7ee66ac843abedb610bba9e024dbe63d90bc640bead8f08fc7a

Observation 6f5c6208-db78-4bb5-a047-88c8e29a820a · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:19.965479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:cfbd1864a1af860af99ee883009340dca6d36cc27f8545e5e0e8f4d1dabebc96

Observation 11d34cdd-0d3b-4a67-9fcb-9d87f1ea196b · outbound

This paper cites Think before you speak: Training language models with pause tokens.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Think before you speak: Training language models with pause tokens

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.975640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:7670b8bcffcabf66ef0825238d575f8926513b5baaa81c57af6ef817b44405af

Observation da4e3547-a919-4072-bd7f-2672385a0c53 · outbound

This paper cites The Llama 3 Herd of Models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding The Llama 3 Herd of Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:20.005863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:70dadd8a22504bbedcbe39abdd72ec77e7750cec4354b604db9832bb63a6d15b

Observation be72931d-abf2-427c-8c25-2e00efb2ca3b · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Breaking the modality barrier: Universal embedding learning with multimodal llms

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.978870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:5419cab2cdf52601f6c7d7777886b1004fc129e7528d586c8db5d7a5599cffa9

Observation 5f274357-1d70-4366-a59e-914aa6d48500 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PLUME: Latent Reasoning Based Universal Multimodal Embedding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:20.029063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:018f855ce9444cdede9da0158cf1eb17572b344299644df27b50fdee79549398

Observation 7d433eff-2e84-47db-bd70-d835ff7855cf · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Training Large Language Models to Reason in a Continuous Latent Space

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:19.991578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:024986b9a5063ad93060a84e7561eaa445dbe2836569c629f2248b07674240c7

Observation bc4003e2-f3be-4ac7-9fb1-0897c9becd2b · outbound

This paper cites Trace: Task-adaptive rea- soning and representation learning for universal multimodal retrieval.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Trace: Task-adaptive rea- soning and representation learning for universal multimodal retrieval

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.029287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:ffb065bf43602e09b1ea7bef8e9a890d1992baa501dccb5ba4088cc384c285e8

Observation 01334dc6-7c3e-47c8-b4cb-0a71141a0ef7 · outbound

This paper cites Gaussian error linear units (gelus).arXiv: Learning.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Gaussian error linear units (gelus).arXiv: Learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.017891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:eac4da400efe1ba3c05b1f9d35c87b7abdd848d3a22ea0a258f0cdeb30c371d2

Observation d2208a64-fef1-408b-a043-349785fa0c5e · outbound

This paper cites Unsupervised dense information retrieval with con- trastive learning.Trans.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Unsupervised dense information retrieval with con- trastive learning.Trans

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.003060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:2f05f3c727a4202f11834753e9d0fd2952e37245a190cf8d0456d376f912ef73

Observation 94535ed8-055f-41d5-9f02-69dee1b07f38 · outbound

This paper cites Cumulated gain- based evaluation of ir techniques.ACM Transactions on In- formation Systems (TOIS), 20(4):422–446.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Cumulated gain- based evaluation of ir techniques.ACM Transactions on In- formation Systems (TOIS), 20(4):422–446

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.161916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:bd0328e231857f2e0bf134d86d8e7f552839281cc57f010eaef5c0b2c8f0e334

Observation 7eab02af-d4fb-432b-ac93-dcb59c52d2ba · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.994968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:d3cc81323c303240518b6af4bdd0aaf2a27412a6c360cff9f203528042595447

Observation f2d9fa5f-eebd-470e-befa-e22263348b26 · outbound

This paper cites Embed- rl: Reinforcement learning for reasoning-driven multimodal embeddings.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Embed- rl: Reinforcement learning for reasoning-driven multimodal embeddings

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.008281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:74bdb1876844d25ea14579ac309bdaa06bffdbb2aa9e86717d64c1558c367e2d

Observation 5e6b8e40-9c2e-4d3d-b3cb-d30bf3819630 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.013841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:a015b4a798f2f09811eb5d5213e86ddb8ef292aea18bc23676094782b093f84e

Observation 2cea8a44-9ef6-40d3-b0bd-81c2db8ba60c · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

PLUME: Latent Reasoning Based Universal Multimodal Embedding VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:19:44.112009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:7b9d36aa2d7304467bd1c1062adaa8242278239006c1aa564f4810dc5227377f

Observation e11658cd-707b-4a21-b7c9-75a16ef72353 · outbound

This paper cites Laser: Internal- izing explicit reasoning into latent space for dense retrieval.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Laser: Internal- izing explicit reasoning into latent space for dense retrieval

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.990190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:e0bb84956430d437ae6c4f29bd4687118a719b61d6c5f692cf2cda14458f4895

Observation cd622f0b-be9e-4732-aaf8-3d66fc25472b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Efficient memory management for large language model serving with pagedattention

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.010643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:679c90b983df9ddd2e35391ba2bd4dc85f37ef4b1dce47f45935945463f798d2

Observation 19d27ce0-432e-4653-bc68-7d16213f1e9e · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Llave: Large language and vision embedding models with hardness-weighted contrastive learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.959643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:693f8cfeec744bde780e880d3e3e3c7f7170640d3fa2055c6cc4498189abe4e5

Observation e2c51b19-46bf-4d5d-87cc-1a8a8cd0105c · outbound

This paper cites arXiv preprint arXiv:2511.00405 , year=.

PLUME: Latent Reasoning Based Universal Multimodal Embedding arXiv preprint arXiv:2511.00405 , year=

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.971113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:82f1c484e227a16b1bfd4824126ffdb69263ef4cb974a2387e92010e424d084f

Observation 86d43a64-3273-4782-aa5e-0deb8234e33c · outbound

This paper cites Nv- embed: Improved techniques for training llms as generalist embedding models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Nv- embed: Improved techniques for training llms as generalist embedding models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.972017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:8d031e08e1a0355eab01c0e71411d903b757228de580315dc5c2934ad715e969

Observation b17b24d4-f844-4c54-b861-5183de808c80 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

PLUME: Latent Reasoning Based Universal Multimodal Embedding LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:19.998828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:60e8631bb4f8d4b440c026133624c45f59de10e40f915b47da90dca313a61fa8

Observation 40444dc4-293a-46fe-8067-b6bc3e284249 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.967900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:3a0bd77537877e12b6a7801b07d856b423a4ed58d9104451d05a9640c3c7f9e0

Observation bdbc48e4-9afb-4b04-b7c2-854017371b57 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

PLUME: Latent Reasoning Based Universal Multimodal Embedding MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.976365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:be42987c0f9e4903d67f1c51713b44a2248eefa69b58e207e627bccd01f548ee

Observation 1d54ffc9-2123-44d5-bd4c-c384c449e324 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.155843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:37babba2815d05b4a1d044003158ccab523a50917cbe8876747c20f6465c943f

Observation aa58d8bf-3d54-4b85-8766-aeaebbb968cf · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Lamra: Large multimodal model as your advanced retrieval assistant

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.170958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:ce65b9e39f211f7abc057026556a4253c6d2d7c343259e8210db3d4995dba305

Observation 9bd35991-65a9-421c-a6b2-b0fc6ee7ec3b · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

PLUME: Latent Reasoning Based Universal Multimodal Embedding VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.258990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:3fad35b25502952b2f9faf061601d737867d3922779dad60fa8b748921c9817e

Observation b32160fd-d626-4731-9e46-3cca74647fe2 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:20.026513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:fd3e19df5679f0160f4eec34f6f84328817573689b7d9de25a02c7bffe44eb9d

Observation 40a95346-5555-4576-ad42-fa5c1ddb6c42 · outbound

This paper cites Let's Think Dot by Dot: Hidden Computation in Transformer Language Models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Let's Think Dot by Dot: Hidden Computation in Transformer Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.031713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:959b584546eb7e26294cf27a057cde11a3e162dccb7edf0d591bb79d8a587da8

Observation e9afa19c-b5b2-4b05-b7ee-471e604c67ac · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.164788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:5ba2b09dcd50a03f6024a5d9d343ceb39535898e557982d5acfe844ba47e9de5

Observation 6111300e-d583-4b53-a2c6-2b00dc9f8001 · outbound

This paper cites Outra- geously large neural networks: The sparsely-gated mixture- of-experts layer.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Outra- geously large neural networks: The sparsely-gated mixture- of-experts layer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.167944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:4abaf65f1f793e00831ce5445b987de5e6ecd42c423da47255661689a5c8feeb

Observation cd332ca2-3c3e-45b2-a30e-08d422e79fef · outbound

This paper cites CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation.

PLUME: Latent Reasoning Based Universal Multimodal Embedding CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:14:57.978007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:96da2eee292794d70f597293a99d6ce3871a466ec1a81c1e7dd7c9a52eadc193

Observation 77212cd5-ae69-4cb0-9a01-6a7d2a844b71 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:13:25.159072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:70b604afafa13a70346eedf423b945b536ce2faa01560027a232e9185e87c4cd

Observation 5a9fb2b3-d68f-4582-9f6e-32c418029bf0 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:20.003704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:b607b265c428123d6fe50611d59b24a33826396e09ce48f2cd15a5e5c262c6db

Observation 86b69184-0572-4f49-8617-1c6017aae286 · outbound

This paper cites Improving text embeddings with large language models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Improving text embeddings with large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.998957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:9df110e61062edb8decc9b88d30555b5027f2578dcd9eb8d6a38ff5f3bf3fa3c

Observation 2d867f4c-e269-4853-acf3-3280717a0efa · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:20.010512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:28d4fd4cc9dc1b5e92078743d5aa2ed60419d192dd77074e7fa37ff1218f6722

Observation 9f6a6085-fbab-4bbb-9f40-a1a8d5d6bdc1 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:53:19.973440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:823c2deb59dbc1b5b93458e97c24c781af0c8f2c4b9a9e8af63e2d90ba296ce6

Observation 452ed2c9-4662-4d7f-8a6d-01b28701feac · outbound

This paper cites Uniir: Train- ing and benchmarking universal multimodal information re- trievers.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Uniir: Train- ing and benchmarking universal multimodal information re- trievers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.979084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:a3851bd364bdc4c7cb693e5741b19960fa6487863741cb80ee73dd268846ec0e

Observation 7b94c3f6-c6a6-4154-8353-24563eb8f58e · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.021582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:b62d8d6bf8456983812593dbf41786d6905c5510a9b6b402ee675611bed9bcee

Observation 985941c3-3489-431a-a2b2-aff6f97d0890 · outbound

This paper cites C-pack: Packed resources for general chinese embeddings.

PLUME: Latent Reasoning Based Universal Multimodal Embedding C-pack: Packed resources for general chinese embeddings

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.006690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:3414aaee5bd75890811bfbe9f092cbed608066be67e36985979a168b414b467e

Observation 138efd70-7aac-470f-a770-a1d132aa6eb0 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Llava-cot: Let vision language models reason step-by-step

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.037610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:43e2e905a48c29a8fece1b5a87d535e3b855e8e2e876fc7adb6f7a9ff2bd9d4a

Observation 25e8e9b0-f326-4f07-9957-616063eba2f1 · outbound

This paper cites Recall: Recalibrating capability degradation for mllm-based composed image retrieval.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Recall: Recalibrating capability degradation for mllm-based composed image retrieval

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.036789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:8349bfe7e3188ff231e2401ea0bc66d0eb587fd3405b656e38ad40baa16e5c48

Observation c688f3f5-316d-4691-92e5-8e7f4ffb708e · outbound

This paper cites Recall: Recalibrating capability degradation for mllm-based composed image retrieval.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Recall: Recalibrating capability degradation for mllm-based composed image retrieval

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.014233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:d6379df9f1b2998d7b18293e27006728a96e2932b998e53bd913dd1d196d766f

Observation 9d2a5160-8c06-406c-8ca9-295e6b704ad0 · outbound

This paper cites CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning.

PLUME: Latent Reasoning Based Universal Multimodal Embedding CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.015992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:6b5b06bb7b32440e535f5ce5ae8ade2618d3e8d37cd5371681b1e00727d4d07a

Observation 1cca916c-89f8-47da-92af-6b7aeef10780 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

PLUME: Latent Reasoning Based Universal Multimodal Embedding VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:37:25.955223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:3a334e0ac9d7fd53c5814d2fb6cbbcfa68b062644be3532460c8f720050da1ba

Observation e8e0c4f7-d7c0-46f7-acb2-5b266bafbf33 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:45:05.263509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:87897a74c1f925e1dcb3a3741f590ad80bbc8e51b67f3f00c1fde3834be6460c

Observation d64212e5-1af5-4a74-9017-ff7f63611ef9 · outbound

This paper cites Sigmoid loss for language image pre-training.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Sigmoid loss for language image pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.982538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:db34e62d4f80f25074487bbf25703e5a5b3e3ee3c1758419e3ee467846fc7c83

Observation b799b7ee-9621-48ef-802b-1c468197ca96 · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

PLUME: Latent Reasoning Based Universal Multimodal Embedding MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.981605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:244f844a934d5a0e85bc0e66d4e33a499d55c43e139be0c1c25b3c048c8f4ff2

Observation a9a6f59f-192a-4e7b-a711-4202275ec05e · outbound

This paper cites Chain of preference optimization: Improving chain-of-thought reasoning in llms.Advances in Neural In- formation Processing Systems, 37:333–356.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Chain of preference optimization: Improving chain-of-thought reasoning in llms.Advances in Neural In- formation Processing Systems, 37:333–356

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:27.033448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:301d1bc4f8c1e72f93c2b225f76896028dd49c007f2c086ba64e74cbedc51bde

Observation 05cbe39e-e60d-454f-a13c-600656786bb5 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

PLUME: Latent Reasoning Based Universal Multimodal Embedding GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:35:22.964550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:d43a9f71b194597ec2171ca5b863269e34ed6eff8d4cdbcd3bfd6f2f55e6e662

Observation b788fd08-06a3-4eb5-8cf1-a9e2bca17276 · outbound

This paper cites Bridging modalities: Improving universal mul- timodal retrieval by multimodal large language models.

PLUME: Latent Reasoning Based Universal Multimodal Embedding Bridging modalities: Improving universal mul- timodal retrieval by multimodal large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:08:26.986324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:d4dd2ade77ea30637dbe3e58e09b969cf4a8c6b1dc91537aba45a35bf00d94c8

Observation 4b2daacc-64ab-4298-8848-2183d22a76d8 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

PLUME: Latent Reasoning Based Universal Multimodal Embedding MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:20.023927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:1396377a0c9abe25f827af10e9cb13f01b74e29d4f52471b273636570eb46618

Observation e55ba35f-04fb-4fec-85c2-97846305af55 · outbound

This paper cites ARMYJUNK.

PLUME: Latent Reasoning Based Universal Multimodal Embedding ARMYJUNK

Reference 57

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T23:08:27.025345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:4e411fe2b7126765d53b0cf9932dfe375419c0497c6527292e42612ed1691e5a

Pith citing papers

Observation 2ebf4763-57f7-40cb-87b9-1abda9c2abb7 · inbound

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? cites this paper.

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? PLUME: Latent Reasoning Based Universal Multimodal Embedding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:47:04.370388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:42:54.802658Z digest=sha256:c3d904dca3128a6acaca215d13cea5730923a2e7c7fd977466a76d4939d8081f

Observation dc1ef913-e2cb-41a7-933c-4f1ebe1bad2f · inbound

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens cites this paper.

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens PLUME: Latent Reasoning Based Universal Multimodal Embedding

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T18:03:36.935342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T18:00:18.315737Z digest=sha256:ee1f333cdaa532312ec9b0de8c84de2467cde245b39950b394f1abe45327d498

Observation cf4d0f5b-d526-4072-afe1-b2879d7b72e1 · inbound

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding cites this paper.

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding PLUME: Latent Reasoning Based Universal Multimodal Embedding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:35:21.372512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:35:21.372512Z digest=sha256:bec30041d1fad83daceb285e2ad087247278b724619457c64f49671486e6dee1