Pith. sign in

Paper Citation Record · LEDGER

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2605.14448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14448 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T02:51:39.142437Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:23.733753Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:43:28.407453Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact11
  • verified fuzzy31
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 94adbe60-0729-4d73-8a99-83e7efbbafc4 · outbound

This paper cites Qwen3-VL Technical Report.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.620043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:509f5166fc3add98619674b98efda57683f11b05fe60cbdc7c4d85564db8057b

Observation d3eb9b61-dd0a-413b-b66d-f29fe0365f3b · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.616969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:1e25a239033726e5f69afbfaa193983e845f78d8750f38bc7d212cd7d81cb9e7

Observation 42fd8a27-34c3-4abc-9e9a-53c10427595e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.298030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:7f145cc03e3eb1d56f0987c8af2749bb93cf9312c0502135e2e2661de94ce429

Observation 25ce55d7-42cb-4c1f-848d-445effd6ea1d · outbound

This paper cites arXiv preprint arXiv:2510.05014 , year=.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture arXiv preprint arXiv:2510.05014 , year=

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.574647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:0656ec99f7c380dfc6611cd7708ded77de4345f5a52e6254711d935023cb71d0

Observation 40922353-9b79-4575-8926-64edb62500c5 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.274547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e4fdbbe3400d78a371fa568c2feaf138bba6e4beeefa3d6f6d6086eebcebed49

Observation f90352f8-3ccd-499d-a2cd-d0e67a4075f8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.607157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e8b3b380e72d1d446f933071cd5dcb8be3c1d1af0b957f2191a9ecd20da2bc6f

Observation 7b41eda3-caf2-4586-bef9-e41f5f07707f · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Colpali: Efficient document retrieval with vision language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.278740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:d16d5c7e7a76f0e132fd77bcedced17cb98b89fab5c1b873d5c5966c7f869939

Observation 3af8c85c-284f-4954-9b19-3eddcf13b494 · outbound

This paper cites Scaling deep contrastive learning batch size under memory limited setup.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Scaling deep contrastive learning batch size under memory limited setup

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.288208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:1351d89d36ecd883b24f0506d9121ba69fa6f4922278be4299de28c81ca04a8b

Observation 2206bcc1-fde0-43e8-aab8-39fefd06315b · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Breaking the modality barrier: Universal embedding learning with multimodal llms

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.644603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:9d5d5b9e2cc8b3af38fda351d3567433d423243517321966882bd1dda2066fcc

Observation 77ac9636-68f6-4152-aa96-d27a2844eb3f · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Lora: Low-rank adaptation of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.625790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:1ec8f6c00d439c38df0c6acfe2973f01129e58bbd1e5c9d94c697d9b776779c3

Observation 726c1a68-4710-4e0f-8e58-62d996a9e19c · outbound

This paper cites Cumulated gain-based evaluation of IR techniques.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Cumulated gain-based evaluation of IR techniques

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.632221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:5e3c822e09156591cf83565e3c7f2d4333c8a2e333c75afe67c9cedd95be2cb1

Observation eea7682d-c7ae-40c2-afb1-4ad1dc97be29 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.638328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:06a645c6a756b5ff9fb0a40e4dabb998f178faa536125e97ba3035b604475390

Observation f298b2ba-57b1-4556-9b70-3f0eb4cef87a · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.013841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:205d08ea2d7dd4da29b0f7f8eadadc42196df340e84bfa21d6d3560425fef46a

Observation 8b067369-f4a1-41b0-8240-1d9f39b73375 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.650503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:6104d46a06cf070a3fde3bb0df2ef7490872de9cc9e045663ee3b2e4813de5ae

Observation 53ea582e-0153-4b8c-ad25-f86a72722762 · outbound

This paper cites Large language models are zero-shot reasoners.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Large language models are zero-shot reasoners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.598543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:56a5c35f687a23edeb6599dcf8c7515f80df1e5403b107a41e07a2035ee13dca

Observation 172de98b-cec0-4fb6-a8ab-e5474d3aaed5 · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Llave: Large language and vision embedding models with hardness-weighted contrastive learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.564439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:968fa0e611d65fae9ea31a0d6eb83af13ac85872bdf7901c14774569cc36e08a

Observation 2d20370a-6f3c-4ca1-b10f-0a4df1dc1923 · outbound

This paper cites arXiv preprint arXiv:2511.00405 , year=.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture arXiv preprint arXiv:2511.00405 , year=

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.610275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:207115605f6f1ea06b27004c3d76f906c1a5e7291ad950cc80a3c3e877bd27fe

Observation 27772d70-64f0-4b92-b333-4f37453b81e0 · outbound

This paper cites Nv-embed: Improved techniques for training llms as generalist embedding models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Nv-embed: Improved techniques for training llms as generalist embedding models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.570183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:0f3fec75718d29521ada96a1f0d61111094b61d54e0a866166cc6cdf77bb2e40

Observation 7c55fd41-6d7d-4067-b70b-f541a8ba9599 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.663216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:b06a8355488ec6678bc8f38f2ddcd2fafab083c0f6bea29d73716cdab83e4a66

Observation b8ec17a7-4ed5-45de-8d2d-951bcfe8d1ef · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.656404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:ff9580777ad3b5da02abb686cd00d81aa27ea3029b7a0ba99af03a2208c55512

Observation 7820e316-f812-4cb0-af06-1b8e88d78800 · outbound

This paper cites Mm-embed: Universal multimodal retrieval with multimodal llms.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Mm-embed: Universal multimodal retrieval with multimodal llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.293336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:a0c1864515ce44d502e902a8194b65fe2a70b66154f7cc4eaef0c683f8289305

Observation 90618f00-1002-4306-8f97-d57f97f28423 · outbound

This paper cites Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval.arXiv preprint arXiv:2511.16150, 2025a.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval.arXiv preprint arXiv:2511.16150, 2025a

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.613684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:ff805082deb7bffb998432ceb7e24fd7ff0ba48541fb5ec284b224a04e60c6fe

Observation 22cc78b9-ade5-430f-8089-07b9cd8217b0 · outbound

This paper cites Visual instruction tuning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.521297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:82aecdfa2eca8d7a681fad422d62a176d3e9302f9e89e3b329e6e4d33fdb12f9

Observation f2b3302f-c83e-4bb9-9bb9-e1470ed0f58d · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Lamra: Large multimodal model as your advanced retrieval assistant

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.546804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:0bf46eee3011732045d074d600102bfba3b9086e46b02642e863cfc2107907f8

Observation bf140399-66ac-4dd8-b1e3-dc555b67bee1 · outbound

This paper cites Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents.Transactions on Machine Learning Research.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents.Transactions on Machine Learning Research

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.533510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:a365e15e106b6a79fe53746cf5194e8ae93ab9e8311b2080415280274fcfd8b2

Observation 7dec594e-9711-436f-9a7a-fc61c71f2de6 · outbound

This paper cites Mteb: Massive text embedding benchmark.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Mteb: Massive text embedding benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.488914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:fc18075062b5016c250bbccffdf008eab3c0a81c12074688978bbe8c92c266de

Observation f775eba8-d4f0-4432-a150-44943d056710 · outbound

This paper cites Vladva: Discriminative fine-tuning of lvlms.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Vladva: Discriminative fine-tuning of lvlms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.494114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:a15ce9b83aae4c783f650aa92259dfda87a201d7f754f5bb670c969ffa37fdd0

Observation bb206138-b18b-4a66-8360-34b9b6206aea · outbound

This paper cites Learning transferable visual models from natural language supervision.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Learning transferable visual models from natural language supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.529126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e7009ed2bd6a2a27d0b508bce4b6a5d5a8a7d5511362cfae1489f7b87536e9ac

Observation 3d8133bb-2658-49a1-b1a8-686d3b852be5 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.479475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:b9e7c741d88758b6725000ab86789847c3f5820c4383e85d7dc88a34ea0fd057

Observation 757f8968-967f-4c29-a2a3-252cf2bd40c1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.603874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:6f04a87d932e9d1e14f7ee65f1bd615689766d7ce403ffbea57f762a0e659082

Observation 85f17632-76b2-4724-93ea-09451092bffc · outbound

This paper cites Smith, Luke Zettlemoyer, and Tao Yu.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Smith, Luke Zettlemoyer, and Tao Yu

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.282994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:801de242f657273c40d670d7ee273ad4ca2247ae2603adf3226441242503a678

Observation 8bc26bd1-8c86-4ce7-b1a9-0b05849ae7f8 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.580487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:7d5523c17b989f105d67403e1ae71849583fa47f7bf329f8790c64dab35c4b93

Observation a06c2e40-b8e6-41f8-934d-a382f565b79a · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.600217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:cd395f1f2258e66795c8a92ba92c1f862556875cb8468c2eadd28b70005898d1

Observation 3e391284-0470-486c-bc79-a579880df57c · outbound

This paper cites Improving text embeddings with large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Improving text embeddings with large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.266541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:d312094674199f0ae6d50d514b0f9d5432451be967f6ca8012cf09633309a85a

Observation ebe631c6-5aa0-40e7-afea-907fc4ea5f7d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.577662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:c6ae428d36b4b5150336b67cb1925b6e7a1752eddcd03a8cdaf300313243d5af

Observation 34e63aaa-94ee-4414-a698-afd43ccfc5ed · outbound

This paper cites Chi, Fei Xia, Quoc Le, and Denny Zhou.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Chi, Fei Xia, Quoc Le, and Denny Zhou

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.270504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:a53afe9c1672e14500d1580620f440b33fdfcfeb540ada665733062c7676b7a1

Observation e7129446-75d7-4c23-9a26-1a6bb699f922 · outbound

This paper cites Cafe: Unifying representation and generation with contrastive-autoregressive finetuning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Cafe: Unifying representation and generation with contrastive-autoregressive finetuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.461203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:006948dff8a81353dca3fa8cb8ac05717792dc1f01339a08e49c5886708220a1

Observation 4f3b7b09-ed5a-4b8d-af1f-4f3417d2b966 · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.466795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:481620844bb67b9fd5728d5222746fd8a77d491ee513ebfd7514102092df0f60

Observation 637e55fb-c38b-4171-9d01-648dc1e5cd25 · outbound

This paper cites Gradient surgery for multi-task learning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Gradient surgery for multi-task learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.473299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:bd9dc75dc8dacb1bf48ca98357ab5886904474489604bcb1da64505cc582df5e

Observation 33c1b78a-bf01-4b73-9cc9-8f5f9d76b865 · outbound

This paper cites Sigmoid loss for language image pre-training.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Sigmoid loss for language image pre-training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.809767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:5c7d80c841ae165aeff432a4a1ddb011469553d7b2b262bc8625052e6044841a

Observation 1c230e21-1155-4bc8-99b6-5a18b2dfed93 · outbound

This paper cites Hauptmann, Yonatan Bisk, and Yiming Yang.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Hauptmann, Yonatan Bisk, and Yiming Yang

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.483846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:8c839a34d28481f7ca6f31455881cdafbd3cbfa7ec294b656032cf47ef06fdad

Observation 4ed72dd8-0c61-4dbe-9ffd-e1f44cc25ef9 · outbound

This paper cites Bridging modalities: Improving universal multimodal retrieval by multimodal large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Bridging modalities: Improving universal multimodal retrieval by multimodal large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.551142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:b4bcf0053c0ede79fabb095742e75b9880962ef03456ab2c419cc0e8d8a05365

Observation 56230514-a7eb-4d62-9193-f769326d580a · outbound

This paper cites think" field completely empty (.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture think" field completely empty (

Reference 43

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T03:59:47.455293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:06a8d45863d5eaf0091edc2accf18a3de9074dd370dba5e9f69a0e6148a5d0be

Pith citing papers

Observation c1fef859-a397-4ace-b531-59308763c5da · inbound

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding cites this paper.

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:43:28.454660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T11:43:23.733753Z digest=sha256:c017b8320297976979af45e3b7616cf849b5f62e8cccfcce1d13e82b46cd61a5