Pith. sign in

Paper Citation Record · LEDGER

Unifying Multimodal Retrieval via Document Screenshot Embedding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.11251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11251 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:54:30.078700Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T20:35:34.389743Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0a8c74e-3b6a-4a59-b860-720164b12216 · inbound

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks cites this paper.

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:19:44.003434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:19:43.882232Z digest=sha256:7dfe9fe5c49b145940a0c541de4528382811c31d62f1bd69f18f1e0fbb1f52b8

Observation ee7345cb-6457-452b-8d99-139940b2767b · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:37:25.914858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:46c82ea99c17d7e4398f854bf86ccd0cd12893dddd1b317c687920be48b50f60

Observation 52ff63e4-f94b-4245-9fff-2ff617cce2dc · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.878444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.878444Z digest=sha256:1777cd45dd103f24ca487914fc4b862624f02cfca1e19c0d10191e83a9d4c0b0

Observation ddbb6ba8-2596-4272-ad8a-bc4bc0c14887 · inbound

MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark cites this paper.

MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:54:30.078700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:54:30.078700Z digest=sha256:72132c5bf2ca234b79191b47c46198a047d2a7957b2f1ecba69ea041511565a4

Observation 34bee540-80c6-443c-8c28-4cfe1de2c402 · inbound

DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers cites this paper.

DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:54.715670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:54.715670Z digest=sha256:2aef3922c9d0a48d2a1a060b32abf16f07784769c9548744e63ff00513e8348c

Observation c9498298-07a1-4917-9acb-0a3c6fc677da · inbound

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings cites this paper.

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:38.087216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:38.087216Z digest=sha256:d40d20e30e147b6ab42d0fa8a67bd57bc4b8f3ec85ce0e6b91098c2d16c78bff

Observation f1ee92c4-9e5a-47d5-aa40-e5628d48190b · inbound

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval cites this paper.

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:02.852824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:02.852824Z digest=sha256:d46f26471a92a594ab892bda36735092b6e4b8ab3639066adaa1292489bd6ed0

Observation 12aa11bf-4c35-4687-8124-4d932a64a8bf · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:03.049531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:03.049531Z digest=sha256:71bf9f0b3eab8d3acd31c4400b9e352b531e63c13e00eefe4cb27a92455e3d26

Observation 8e2e274e-6fc6-4b8d-b793-bdaac0a7bc63 · inbound

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding cites this paper.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.163447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.163447Z digest=sha256:f7a80ce46edce2dd2f35d85d7620a77b23b3f1cc49a2effcd85804ecd26c900f

Observation 6c075382-8f53-4aca-af63-26429245d348 · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:44.667001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:ad923b05e095b1c7ab11f68c6127982aac9bd3b1908c6328ff4e250724c77d48

Observation 58ad22d6-b653-4236-9404-83176f58b5ae · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.730245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.730245Z digest=sha256:8d45a9ac35d60f9caff508c1a1e70775ab7360472db2eb39c615f7c6cacde60c

Observation dc5dfa3c-7e2e-4a50-a259-76316b58efe9 · inbound

UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates cites this paper.

UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T03:30:18.920992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:30:18.920992Z digest=sha256:ec391e72371602555559c64cc2b934cd159928b1f10ade5b69ef8e4697957e2f

Observation 9d64daca-7ed4-473c-9939-efbbead07a87 · inbound

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL cites this paper.

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:00:58.257323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:51:16.675856Z digest=sha256:f84a2fe33ba0f70e646713719983b2a6f8933b26156ac415ff35c9aa0409eb8c

Observation 1ebebf6e-4a94-4a38-ae5d-354e7e95c0ab · inbound

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment cites this paper.

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:57.258511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:28:59.838565Z digest=sha256:79ed8934394a8d670d9cffa27f99f11ddde68dfcea0d252212f1e7f90b2b8150

Observation 61a2e927-ece4-4e15-b5d4-550b9650d052 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.417608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:937e049cfe7740b6ced09caaa0eb10e22350991d266bcc0646bc9b15da1c732c

Observation 113d6025-1d2b-45fe-8ee0-40e2765d67c2 · inbound

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark cites this paper.

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.122467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T08:09:41.068000Z digest=sha256:1920f4e412714fd018c3c2741425be247eba9b276842b091bdaf116d98a59156

Observation 1491a0ee-69d4-4915-84df-6cdec3795c12 · inbound

CMDR: Contextual Multimodal Document Retrieval cites this paper.

CMDR: Contextual Multimodal Document Retrieval Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T20:35:34.390920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T20:30:59.126121Z digest=sha256:fe642d883e2da71665a47bd0cf7731beb021a585045cb0f160a3c3e865f38ae2