Pith. sign in

Paper Citation Record · LEDGER

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2509.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07538 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:14.726012Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:12.301504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:06:14.962988Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 606b6d8b-d461-4960-af5b-173171b3b925 · outbound

This paper cites TextlessRAG: End-to-End Visual Document RAG by Speech Without Text.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:06:15.028832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.301504Z digest=sha256:ec0bc09b25b873d1302e7d7ea85e49a899d7cd847a3969113929cadf5c12a3f4

Observation 600a6e06-3e4b-4834-b9ad-dbfe3a638a6e · outbound

This paper cites an unresolved cited work.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:06:17.673515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.385386Z digest=sha256:1a209b0630a51b1f9d142bef996b6d673cb31cb9a4ec16bcc3f68a3fad1507dd

Observation 47eab786-962c-4950-91ee-8c1849fab26a · outbound

This paper cites QA” and “Pool.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text QA” and “Pool

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.545134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.476150Z digest=sha256:dd83481fd46bc798783cca05f575433d14d2157579f3592e554b9ca083e9e788

Observation 1bacaeaf-5a0f-4afa-abb0-3abaca15851a · outbound

This paper cites T”, “I”, and “A.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text T”, “I”, and “A

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-04T22:06:17.451643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.554832Z digest=sha256:5c512a8db1a2961f962edba40b970b2367e275448220a207c01cc41675f46a43

Observation 3440a9c0-d2b6-40a1-a534-d80f8976fdb2 · outbound

This paper cites We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.317046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.640657Z digest=sha256:099bf15845c3daaba3a32d2598b4702e7908f74325dafd26d92111f3baadbbcb

Observation 1423b882-83c0-40a3-8715-4419875f8612 · outbound

This paper cites Qwen2.5-VL Technical Report.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:12.726138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:12.726138Z digest=sha256:5e4e1fc5dc93aeb6b8f9b148fa56b6827fd550843d33f7466f9033158756b510

Observation 0f0f9cdd-5cb6-45df-a3b4-b80af33a98cd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:12.859876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:12.859876Z digest=sha256:149cc64ae8539564cdc693f4d1d6a974590e66c19501fa71e5e03250d5cae134

Observation 871b664e-c294-4bcb-b8d9-0f3f6e942b40 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.214049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.949476Z digest=sha256:e46b3d24633153542a152afd72e807146f6c1da3603ebd633d74d347fc91d48a

Observation d34ea977-7fb2-4015-9f58-671224254057 · outbound

This paper cites Qwen2.5-Omni Technical Report.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-Omni Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:13.082058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:13.082058Z digest=sha256:3fccba8693f49379f8c1d9079abe205fecbe67f395485475c18d84a1946a7d20

Observation 32547a26-8c4f-457a-99e8-5ac6ce85f793 · outbound

This paper cites Multimodal large language models for text-rich image understanding: A comprehensive review,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Multimodal large language models for text-rich image understanding: A comprehensive review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.120494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:13.228865Z digest=sha256:2fc59094d368eae9311dca91082f2ac24d036d4eac5a2a2526d35bcf3d5ba6be

Observation 95a827ec-44a4-4fe2-bbe2-090653ed6d23 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.990451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:13.373610Z digest=sha256:a28c5eece3a5dd59344891356aff4afb4e89c4ebd7f9a0934b4c9135be1209bc

Observation 562fc4ec-99ff-4236-b3c1-ac4a53790dd8 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple im- ages,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Slidevqa: A dataset for document visual question answering on multiple im- ages,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.851276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:13.458860Z digest=sha256:e8d5fc71ecd86f4a263da43b7c1d92bb0e57af42909c1a795a9f7c7c9ad3c1cf

Observation 480d65f9-44d7-4389-96f5-54c4f5ff02ee · outbound

This paper cites Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.748304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:13.685022Z digest=sha256:9441040b8e8d88836e99ff113e7c3b9f2c51d5bbf8f9c5e7ec8fedfb98cd1fa8

Observation cf69189f-c193-4101-ac81-a1f6822cfb10 · outbound

This paper cites Vdocrag: Retrieval-augmented generation over visually-rich documents,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Vdocrag: Retrieval-augmented generation over visually-rich documents,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.643795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:13.763678Z digest=sha256:4f170673bc9636e56676a241d5191a5952f706f4659258724bfe04acc9a160bf

Observation 7cbf5f85-c8b1-4fce-a48e-2d61d87c3c8c · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:13.857070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:13.857070Z digest=sha256:e5a9498a1572c67037f3405b8f7a3433b9ed80d0c5fd9193295d595f3e49af61

Observation c36e5c04-4cf5-4a1e-b3f3-c55cc9e501dd · outbound

This paper cites Towards multilingual spoken visual question answering system using cross-attention,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Towards multilingual spoken visual question answering system using cross-attention,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.545105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:13.956121Z digest=sha256:46130e24f1e12110f94889d8800113090bc91feb5693aeab12e467a251033b15

Observation 3629ed68-c9d4-4e57-899b-ccbee08d4cd5 · outbound

This paper cites Spoken question answering for visual queries,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Spoken question answering for visual queries,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.411574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.052491Z digest=sha256:5e2f55595eda38514e7c13c07f5fff052209ebfda3df63316c2b074289e6c03d

Observation 7cd47da3-cd8c-4f36-9db2-675f04a48a10 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.316678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.163239Z digest=sha256:a701c35bb8350970298eb47652cffbdc6efe96bd7bedf74c0de19fd41920940b

Observation 643ea55e-f9f3-40c9-a69c-72a08ff89757 · outbound

This paper cites Infographicvqa,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Infographicvqa,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.186822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.279547Z digest=sha256:e21f5d11c687c6192d77a4940852fc605267989d720abaaff43953227bf8cc9f

Observation b53390ec-f5f9-4d6a-9ad6-698038c1d832 · outbound

This paper cites ICDAR 2023 competition on document understanding of everything (dude),.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ICDAR 2023 competition on document understanding of everything (dude),

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.083954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.365990Z digest=sha256:7022878b83b39005a230c06d11119bdf279134165ad604a80c814aeb4ec04147

Observation 451b199f-c2f9-4bf3-b9a5-4bffa7b36efa · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text The probabilistic relevance framework: Bm25 and beyond,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.934745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.495250Z digest=sha256:1f38880b865a678423107931980e98f3d2a2c8da084e615ae8aa66b22b2c3c2d

Observation 4298f707-c53d-4f71-9f43-94128b5750ac · outbound

This paper cites Text embeddings by weakly-supervised contrastive pre-training,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Text embeddings by weakly-supervised contrastive pre-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.783246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.561399Z digest=sha256:7d3aca312e2fb71786592016f1015ee91f26547499adf925e720350db9102a0b

Observation 61a5874e-f27b-4497-a356-f2e1dfdda83f · outbound

This paper cites Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.560850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.635933Z digest=sha256:46fa01eaa5c1b9ab021e54f0db21d526488b811449f6dcf09424dab4ca08e715

Observation a17f529f-8511-49a9-9759-ce5e9170135b · outbound

This paper cites Learning transferable vi- sual models from natural language supervision,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Learning transferable vi- sual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.372478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.705712Z digest=sha256:95b5a3cfb6a07e79a8283de9972ec6b2d1dc2f7f34df81f4b7d41dcd23be4eb7

Observation 79f12e5e-06ae-4220-b8e0-5c89fe9d0eca · outbound

This paper cites Uni- fying multimodal retrieval via document screenshot em- bedding,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Uni- fying multimodal retrieval via document screenshot em- bedding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.231025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:14.726012Z digest=sha256:02ea8b1156a563b660e18a4bc589b5b2857b065d0536e6c2b35671f76d24dc60

Pith citing papers

Observation 606b6d8b-d461-4960-af5b-173171b3b925 · inbound

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text cites this paper.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:06:15.028832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-04T22:06:12.301504Z digest=sha256:ec0bc09b25b873d1302e7d7ea85e49a899d7cd847a3969113929cadf5c12a3f4