Pith. sign in

Paper Citation Record · LEDGER

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 3 inbound Pith citation observations for arXiv:2508.07493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07493 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:10:08.326585Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T20:30:59.126121Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T20:35:34.394180Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact2
  • verified fuzzy31
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cef3d718-2b83-49af-b4eb-40d8a75fdd03 · outbound

This paper cites Ask in any modality: A comprehensive survey on multimodal retrieval-augmented generation, 2025.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Ask in any modality: A comprehensive survey on multimodal retrieval-augmented generation, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.400345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.070614Z digest=sha256:94e28262706c10ce47acf5dd2fd3dbab1a273eca3572275d333094ccc432190d

Observation 49a2697d-b88d-4b04-961d-8c0987481054 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.075964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.075964Z digest=sha256:18c2bc1f7362debf1e83d3d05d382d713d757bad8cb7f86bddbeb4feb60fee8c

Observation de2ffbca-14ff-4c85-b1c6-71f3de933b71 · outbound

This paper cites Scene text visual question answering.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Scene text visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.387813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.080601Z digest=sha256:98820d5f9d24bac6895838d06f2246449d48bc251d87d7154b5a27ac4316f670

Observation 236003d3-acec-4c0a-b34f-932a000559ec · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.085053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.085053Z digest=sha256:463d4b7d636cfa29a20b8cc858670ecb7e9d8e03a8901fe079af8e2650c15c01

Observation 17ee1c91-bd37-4bdd-b3c7-811c5e3a18e8 · outbound

This paper cites MMR: Evaluating Reading Ability of Large Multimodal Models.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding MMR: Evaluating Reading Ability of Large Multimodal Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:10:08.828773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.089589Z digest=sha256:a597d0957796b03573e3d86aaccefb5e1f636ae9570e4d05ba6029a864f7555a

Observation 427bba16-f186-4e88-8af2-fbe1d210e983 · outbound

This paper cites SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.093349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.093349Z digest=sha256:e0a55440d53e1f8b88696aef5325bdc78ec54941a638d29edd2adc75213d6763

Observation 686aaee8-751d-4f80-adcb-3b263e763ec3 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.097589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.097589Z digest=sha256:7cf10754dd8dbb2db15c2ce3e6eaafbd73af2ce7ac23ab9ecf31d39f3d954e66

Observation fbe50c63-4ed7-409d-a45e-8fe5940338bb · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.101248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.101248Z digest=sha256:b78bcc2a326b6c26f70e046997bf46d87a07ec739b8abdbce98acd956e73939f

Observation ba546d15-0083-4c4f-bd2f-46129c89b4e3 · outbound

This paper cites Mmvqa: A comprehensive dataset for investigating multipage multimodal information retrieval in pdf-based visual question answering.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Mmvqa: A comprehensive dataset for investigating multipage multimodal information retrieval in pdf-based visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.376320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.104968Z digest=sha256:c0170ce760a552eadd1eea7c3b1906c3eca6e078d4958e4aa6939750e513fe8a

Observation bc034659-1053-4925-b18d-9e75c27f4a8b · outbound

This paper cites Mmdocir: Benchmarking multi-modal retrieval for long documents.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Mmdocir: Benchmarking multi-modal retrieval for long documents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.108542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.108542Z digest=sha256:7e21b14fbc4496d859eff6aff8a663bab8d656d11e3631d1caa6ec4a46eed029

Observation 7e66d650-7e94-4262-9dba-76ccc06a12bb · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding PP-OCR: A Practical Ultra Lightweight OCR System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.112765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.112765Z digest=sha256:7085436397bda9ff4a4d0d98496d0d23c82c08c587cb7bc9cd3fbee46e00cd22

Observation eb663534-4708-4585-b950-ec34a7f6396c · outbound

This paper cites Colpali: Efficient document retrieval with vision language models, 2024.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Colpali: Efficient document retrieval with vision language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.365473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.116239Z digest=sha256:522f3c51ad4e36c4c6a079abd9498ee615fd6724538e2abbd3216a69e5d8425d

Observation 3839b108-bf82-4209-a682-6e971dacc5a3 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future direc- tions.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Exploring the frontier of vision-language models: A survey of current methodologies and future direc- tions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.120202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.120202Z digest=sha256:4589afed7c1e34c944e00fc921a34d3e61e25dd9f3d347904eaff5f1da89e618

Observation 67c4d1d6-05bb-4042-900a-2df0fe228000 · outbound

This paper cites GPT-4o System Card.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.124378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.124378Z digest=sha256:e26f66c1f06f7b5e31713f3d5e6efbbcfd1f627e8e9e54170328885775297e04

Observation 25facea2-f7f6-4d1d-ae90-eec21fdc58f4 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding FinanceBench: A New Benchmark for Financial Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.128979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.128979Z digest=sha256:c168ae2b78f1b02b7266ca66de4ba409baf0c11b234b351be1c5ed5a8f7d29e9

Observation ce12ed8f-f79e-4f24-874a-71968cbf950d · outbound

This paper cites Mistral 7B.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.132923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.132923Z digest=sha256:4d06b1f2fb7dd7f7c19a59fd77667610beea443bd34bd6d270124e412619f6a5

Observation 4401a79d-ceba-45eb-bd92-96292f0a2121 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.136811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.136811Z digest=sha256:f5ffe24946897b91e013e93a9a7e55c285cc0e0ad122126bc5b1e27657d76991

Observation 8178269d-3ae0-44d5-9ba9-589d7024731a · outbound

This paper cites Colbert: Efficient and effective passage search via contextualized late interaction over bert.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Colbert: Efficient and effective passage search via contextualized late interaction over bert

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.352324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.141567Z digest=sha256:3cc0a122da41917aaafbc6c0d482b91033eb19d785b17a5b6c720510b34f4961

Observation e254639f-2e33-484c-87c7-28b6ff9e545a · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Building and better understanding vision-language models: insights and future directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.145820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.145820Z digest=sha256:d5bb7b7942a44161a42b74ce5ab2bf5fc4a14fbc325e72ea8859a634b04f155b

Observation e85c21d5-6ce7-49c9-a880-20486e152a65 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.149428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.149428Z digest=sha256:1493e5d0adff19cf25b3db9279b844265a271c06ae488e37a49ecb0d022970a4

Observation c6a6873c-73af-4d37-b112-bbf862166eb0 · outbound

This paper cites CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:10:08.541500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.153107Z digest=sha256:2bca1d40de821ff41af34f4a180c805fecef5ab8b54354ab955cac2d1d3e5317

Observation f41e5b77-b01d-417b-9af1-f9c64d773573 · outbound

This paper cites Towards visual text grounding of multimodal large language model.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Towards visual text grounding of multimodal large language model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.156626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.156626Z digest=sha256:5df8f617593e986494ad98651ac634a8218ff7bd4515cbea0773618a1ee31311

Observation 3c52f217-b39f-4eae-8729-03018f9ad76e · outbound

This paper cites Chatqa: Building gpt-4 level conversational qa models.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Chatqa: Building gpt-4 level conversational qa models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.340001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.160227Z digest=sha256:5c47949a82cd7ac85ac7bc8af5424fef93edeab759114d3c54acff9dc1d3e33f

Observation 8e2e274e-6fc6-4b8d-b793-bdaac0a7bc63 · outbound

This paper cites Unifying Multimodal Retrieval via Document Screenshot Embedding.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.163447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.163447Z digest=sha256:1594a45a46f824067b8cfb7ad07dc711d7396bd91b2109c702bcba090b55cba4

Observation 539bd716-b703-4f35-bd20-4a0174b3bf81 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.166980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.166980Z digest=sha256:bc5e8a92e9f64a6ac4de29804a2b0fab4490a3bf6e5a8f96481224470a615e44

Observation 7c855fdf-38c4-4dc0-989b-5f88e97bf2fe · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.171236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.171236Z digest=sha256:d0d97e21427a0154a86778b2d09ac44617b1fb9c10067a5aee0f0af4d81777f3

Observation dab151a3-4749-4307-ac39-70c52a018b2e · outbound

This paper cites V Jawahar.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding V Jawahar

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.322398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.174784Z digest=sha256:b59189770e4121fff56c438fe542d9b33a4cac951406d9c4235859f26b7c30ac

Observation 7d6ba0b1-06cd-4a3d-bcee-703b04a1b437 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Docvqa: A dataset for vqa on document images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.310288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.178898Z digest=sha256:ae1f26079a213f5cdd6f91da51b7a74f6b0364325e4d5e871d23c5dfb930d18e

Observation 6a8bb147-62c3-4461-aaa6-2ede29c9886a · outbound

This paper cites an unresolved cited work.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:10:09.297590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.182279Z digest=sha256:2af0429f90159610649adfa270d39c17907f038e67e55f148dad876ddc3fd40f

Observation 0991974f-67fc-4a92-8ac9-52703c17303d · outbound

This paper cites Infographicvqa.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Infographicvqa

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.287121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.185760Z digest=sha256:b085e4303260627553885fc88905c907118f1d14270e2520520a49dee3e85de8

Observation 1cbcec69-b295-4927-9eea-018fcc8cdba1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.188996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.188996Z digest=sha256:1d382b3c666c9e76cce45efc4dd8620372c3b49b704b3c4053de4a74233da7bf

Observation 5a57f2eb-4f56-45bb-bff0-972b5bc44ece · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.192384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.192384Z digest=sha256:a05da0f10737d209cb1a76b8c7ab27fce2ef4cb262c339957c3a223ebf7969ec

Observation 298f9ed9-afbc-4c2c-9288-e10c7e74780f · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding The probabilistic relevance framework: Bm25 and beyond

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.197362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.197362Z digest=sha256:3fb5d6b802b4ade63ea65726f496da97aa6f1dcd674b6cd778bffe727dcf7603

Observation 428e2e7d-8199-4a49-9c58-b2f70941b2a0 · outbound

This paper cites CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.200780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.200780Z digest=sha256:c93ceb029873fe6f34653af4138e4f66a25606f6e0345c8e682e8c4d8e874fdd

Observation 1fbde0bf-ea6f-4e3a-a360-1a19eabb338c · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.263242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.204634Z digest=sha256:22d7e9b6e73b2e84f3219f761e16e4c224b2b598e33dc3f9fe40e3b542a3e9a8

Observation d1815816-bec4-4e11-acf0-527dac0a6236 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.208206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.208206Z digest=sha256:2bb3fc8ad107884ed9b6ee89269bae9edf370ca4789eb2a10736d427e111c1b8

Observation b81e0c3e-2dac-4d10-911c-bcf57a3e07fe · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Slidevqa: A dataset for document visual question answering on multiple images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.249944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.212586Z digest=sha256:f494329bc79f573c46ddbbb698fbe2bf2322356496d1e3d4f7a33e076dec0d55

Observation bbcbc31f-1536-48bd-b1f5-f15156d55e78 · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Hi- erarchical multimodal transformers for multipage docvqa

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.236420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.216257Z digest=sha256:b454c170e9b9981c3b9a06f37f6c197d9c708a9c56aadb6852b871c2f2108e1c

Observation f00acd49-5a06-4e0f-9e0a-198588d902a5 · outbound

This paper cites Ccpdf: Building a high quality corpus for visually rich documents from web crawl data.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Ccpdf: Building a high quality corpus for visually rich documents from web crawl data

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.224271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.220070Z digest=sha256:d8702caacf8288d8a7f0d816630b0bc0c7a915ae6d57cc5511ac6a941032fee6

Observation 695fa568-86b0-4283-b7cd-fd17737a10ce · outbound

This paper cites Document understanding dataset and evaluation (dude).

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Document understanding dataset and evaluation (dude)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.212398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.223643Z digest=sha256:c463ae7e94d081d5d05e801cff8a669a921e1cc6f8cde66ee16b4480d4198a33

Observation 04ba4415-aea9-419e-be8c-4388905b652e · outbound

This paper cites Needle in a multimodal haystack.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Needle in a multimodal haystack

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.200594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.227655Z digest=sha256:c80c301b812b69968c4c72a204a9b7e2790d6ba81fe1df0e959b30232d876bb6

Observation 062f3806-bfe6-4314-9aa3-6e8a2fea737c · outbound

This paper cites SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.231807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.231807Z digest=sha256:eee73148f336ddbe23d545941d3b5de00ccbfe746d4ca45f278fddcb089ebc7d

Observation 82d4f2a7-320c-4b62-802a-370574976ebb · outbound

This paper cites C-Pack: Packed Resources For General Chinese Embeddings.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding C-Pack: Packed Resources For General Chinese Embeddings

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.236037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.236037Z digest=sha256:097f89c27f50e4f9d89753358ecb305887e6c47ddb7578be9035403b10a3bd28

Observation fb41c3c8-b199-4735-9d7c-a93153a199cf · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.240158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.240158Z digest=sha256:6c2db645179c1d6dfba5c88ba6ce45b3730c139bd538320a54db572e398446ac

Observation 112787aa-b36f-40d4-8470-18808813eaa1 · outbound

This paper cites Sigmoid loss for language image pre-training.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Sigmoid loss for language image pre-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.190017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.244392Z digest=sha256:618c63fe4e86191090d6d455c074b48b57d81b90e65f0e4196d41277853d12d5

Observation 174ac9e7-7817-48c3-b7a4-a7056ac16699 · outbound

This paper cites Vision-language models for vision tasks: A survey.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Vision-language models for vision tasks: A survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.247951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.247951Z digest=sha256:70139643dd628dfdb9e9951ce72f1135982fcdadfacb816795dcf1d095e40671

Observation e84236d5-00f3-42ac-ba1c-8225e3b66a8b · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.251950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.251950Z digest=sha256:e7e682b8bbf61d41f21e650ef35cab941b42df5e1b98f5f9301555c1b799f6b8

Observation dccb1653-7e37-4aff-8044-5805b4604376 · outbound

This paper cites Multilingual Large Language Models: A Systematic Survey.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Multilingual Large Language Models: A Systematic Survey

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.255901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.255901Z digest=sha256:e523802f45221c48f12fc198a3cc7def3d3c91d276d2c246d452df28b79e8ea0

Observation 1d9415f7-ecb0-4587-bb78-a591530a7ddc · outbound

This paper cites Figure 3.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Figure 3

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.170235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.260035Z digest=sha256:63e75bf2d54a1d00e7396fa32c82732e64195c2ff399cc1e872e3cb81024afea

Observation b9f4a5a2-f0a5-4814-9487-29d892cd3f1a · outbound

This paper cites • If multiple figures or tables exist, distinguish them based on their title, description, or content.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding • If multiple figures or tables exist, distinguish them based on their title, description, or content

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.158776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.264001Z digest=sha256:4a9bb5aa5b90355e79b13cd01f2865a89e6bd958b2996b5f44600d160899a3bc

Observation b9ebb48a-fa49-40ae-858f-0310d05fdaa1 · outbound

This paper cites detected_language.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding detected_language

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.146526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.267582Z digest=sha256:391b6f0fd4752af5e2cbe79ce6a817c768d73f1eb634908a4af1deffc8e36b56

Observation 5d51418f-cb26-46cf-8bc7-4349fd89146d · outbound

This paper cites an unresolved cited work.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:10:09.134902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.271444Z digest=sha256:55eccc758f5b32e231da80f78638acf6293b366a66d8e56037125844046841a5

Observation fd00dd00-3833-40e2-bddd-9af86d00cae3 · outbound

This paper cites Now, please revise the given question pair according to these guidelines.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Now, please revise the given question pair according to these guidelines

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.123366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.274994Z digest=sha256:26f8e525281e850346ec9e6718d5fe6fa34147d67ff2b51516ce1d1a7f5b1c7a

Observation 34021eb1-0a6a-4f6c-8f98-691c52fda117 · outbound

This paper cites The question should be about the subject of the page, and the answer needs to be found in the page.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding The question should be about the subject of the page, and the answer needs to be found in the page

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.112226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.278528Z digest=sha256:1e0168e154ccea1ae27dd983cf135fcb1e2abd5ed04173b05d2f236f39af1c0f

Observation eb5dcbb4-ca4c-4c51-97f0-c48a0e464fa9 · outbound

This paper cites Generate a question that could be asked by provided infomation in the given page.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Generate a question that could be asked by provided infomation in the given page

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.099368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.282492Z digest=sha256:f532c558c3f77eb59575a7279b8ea527b991a7979b48ebffaa7e5d93a903e2a9

Observation e23b6640-6fd2-4a3b-baaa-3d32f995ec08 · outbound

This paper cites Please do not generate:.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Please do not generate:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:09.088464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.285879Z digest=sha256:83a8a04aa9c82089208b23fcb932545109ed7e4a797fc7fb0a5f788db25a03b3

Observation 7eec0267-3ef8-4805-a2e7-f668a91285d7 · outbound

This paper cites What is the page number?.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding What is the page number?

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.965761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.294022Z digest=sha256:97febcf4df483d92c73ed89c011ff8e164faf60cacc9656f4e5e98359566f4b7

Observation 22e9eb5a-af8e-4ac2-823f-e10107f08c61 · outbound

This paper cites detected_language.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding detected_language

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.954615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.297781Z digest=sha256:1aed3bcefe8a824357e0709c91618aa3cb7a57635af267567723da32bb3c16ad

Observation e62c39ae-2e71-46df-9c17-3d99cbfca7c3 · outbound

This paper cites an unresolved cited work.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:10:08.942238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.301458Z digest=sha256:c36c26540ef64ed1e2815499a65164288f5e9b15992db026654a48a6a6996e07

Observation b18e7707-7ee6-4b3d-9796-675a1ad0538c · outbound

This paper cites How has X changed over the last five years?.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding How has X changed over the last five years?

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.927041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.305054Z digest=sha256:5f73954d344c170e182e5a2292a5d09121c58d750ef59f9c589c3953038e158e

Observation b538a815-d213-435b-84ec-e0949ce91b60 · outbound

This paper cites The question should be about the subject of the page, and the answer needs to be found in the page.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding The question should be about the subject of the page, and the answer needs to be found in the page

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.913564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.309386Z digest=sha256:5b2ae8b678baf3b2299256fca82dc05fa6c44a9bd3c54d4a1e12914a0612f8c6

Observation 6306bdc0-9729-4865-ac33-a4ced277e330 · outbound

This paper cites Generate a question that could be asked by a user without knowing the existence and the content of the corpus.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Generate a question that could be asked by a user without knowing the existence and the content of the corpus

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.901102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.312602Z digest=sha256:c2bb4c83bc81373d2450a680e8992f952620f5629f62d96b4377ffc00689440b

Observation 264377b4-8d4b-4ea4-b12f-189f7c8d60d8 · outbound

This paper cites And the format of the answer should be a list of words answering the question.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding And the format of the answer should be a list of words answering the question

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.889450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.316031Z digest=sha256:983d8c71bc917d8ffcace74049cf1f28a2e3558cb79a931e2d4f4d69b30ac42a

Observation 9aa26f0b-a847-42a6-8f8a-d692e855d0f6 · outbound

This paper cites an unresolved cited work.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:10:09.077964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.319626Z digest=sha256:9682f1ab7c1e3b3fd891ef893d900a33c236f2a24d818ffb47a66cf3f22cb899

Observation 7603cd4e-66a0-410e-9040-c94df6f0891e · outbound

This paper cites What is the page number?.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding What is the page number?

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.876958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.323166Z digest=sha256:843355921c0f2f986b7e0cca7e79f3101bddacde1cd30a08c60966af06ef42a2

Observation 6911ac89-10ed-4726-a925-e34863514e1a · outbound

This paper cites detected_language.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding detected_language

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:10:08.863023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:10:08.326585Z digest=sha256:9ad084c00635f63de9816d70575e637cdec7d60b5fe19847021ea92a7adc51d7

Pith citing papers

Observation e0b5a8c2-a50f-4d35-8fd4-0f029e609d28 · inbound

MINER: Mining Multimodal Internal Representation for Efficient Retrieval cites this paper.

MINER: Mining Multimodal Internal Representation for Efficient Retrieval VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.643227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:40:33.437364Z digest=sha256:d7e017709d3f7c181d180b996c6ed55075b2941e39c0713f2cf7da2aa106150a

Observation 5dc0e3c7-5291-48ee-a3e6-dd165098b149 · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.555944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:44738e4b26c3cf4f0d7fbf4c53fb204b7460adc7d72d502cdf4887e42394e038

Observation 98c99682-c945-41b9-b38b-4462c8e29e6a · inbound

CMDR: Contextual Multimodal Document Retrieval cites this paper.

CMDR: Contextual Multimodal Document Retrieval VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T20:35:34.395414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T20:30:59.126121Z digest=sha256:bda2219dc6c048873102a2fdd86b99a6af4d819dd5f20bb0a936f6ac0a57a6e6