Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:14.726012Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2509.07538.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:14.726012Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:12.301504Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T22:06:14.962988Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 606b6d8b-d461-4960-af5b-173171b3b925 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 600a6e06-3e4b-4834-b9ad-dbfe3a638a6e · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47eab786-962c-4950-91ee-8c1849fab26a · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text QA” and “Pool
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bacaeaf-5a0f-4afa-abb0-3abaca15851a · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text T”, “I”, and “A
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3440a9c0-d2b6-40a1-a534-d80f8976fdb2 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1423b882-83c0-40a3-8715-4419875f8612 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0f9cdd-5cb6-45df-a3b4-b80af33a98cd · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871b664e-c294-4bcb-b8d9-0f3f6e942b40 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d34ea977-7fb2-4015-9f58-671224254057 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-Omni Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32547a26-8c4f-457a-99e8-5ac6ce85f793 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Multimodal large language models for text-rich image understanding: A comprehensive review,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95a827ec-44a4-4fe2-bbe2-090653ed6d23 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 562fc4ec-99ff-4236-b3c1-ac4a53790dd8 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Slidevqa: A dataset for document visual question answering on multiple im- ages,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 480d65f9-44d7-4389-96f5-54c4f5ff02ee · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf69189f-c193-4101-ac81-a1f6822cfb10 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Vdocrag: Retrieval-augmented generation over visually-rich documents,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7cbf5f85-c8b1-4fce-a48e-2d61d87c3c8c · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36e5c04-4cf5-4a1e-b3f3-c55cc9e501dd · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Towards multilingual spoken visual question answering system using cross-attention,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3629ed68-c9d4-4e57-899b-ccbee08d4cd5 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Spoken question answering for visual queries,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7cd47da3-cd8c-4f36-9db2-675f04a48a10 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 643ea55e-f9f3-40c9-a69c-72a08ff89757 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Infographicvqa,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b53390ec-f5f9-4d6a-9ad6-698038c1d832 · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ICDAR 2023 competition on document understanding of everything (dude),
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 451b199f-c2f9-4bf3-b9a5-4bffa7b36efa · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text The probabilistic relevance framework: Bm25 and beyond,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4298f707-c53d-4f71-9f43-94128b5750ac · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Text embeddings by weakly-supervised contrastive pre-training,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61a5874e-f27b-4497-a356-f2e1dfdda83f · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a17f529f-8511-49a9-9759-ce5e9170135b · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Learning transferable vi- sual models from natural language supervision,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79f12e5e-06ae-4220-b8e0-5c89fe9d0eca · outbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Uni- fying multimodal retrieval via document screenshot em- bedding,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 606b6d8b-d461-4960-af5b-173171b3b925 · inbound
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.