Pith. sign in

Paper Citation Record · LEDGER

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

As of 6 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 5 inbound Pith citation observations for arXiv:2410.05970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05970 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T19:39:35.147671Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T16:02:00.920066Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-23T19:15:47.037493Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact30
  • verified fuzzy39
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59ec3b1c-bf5c-4af9-b826-5472323df84d · outbound

This paper cites PDFTriage: Question Answering over Long, Structured Documents.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling PDFTriage: Question Answering over Long, Structured Documents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.789541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:4f0c35b161aa249944835f0c4a49b0cfe6ed862a5fd078ea6266c1efc103a81d

Observation 51befefd-d04f-48bf-8eb2-5208b0d833d9 · outbound

This paper cites Prem Jacob, Beatriz Lucia Salvador Bizotto, and Mithi- leysh Sathiyanarayanan.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Prem Jacob, Beatriz Lucia Salvador Bizotto, and Mithi- leysh Sathiyanarayanan

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.128531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:df00371e85b7ddfefc7a36f6e6ee1e466af9c538f0a46793c32caad89396e427

Observation c04e0b88-dd43-4063-9a38-fe2543e1dca6 · outbound

This paper cites YaRN: Efficient context window extension of large language models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling YaRN: Efficient context window extension of large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.124847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:294a60f631888b9dadb8c359e258e203c7acbe4cfff7281a0cbfa488713f74ef

Observation 62073c19-7e96-4787-84d1-0c74ef39b940 · outbound

This paper cites LongloRA: Efficient fine-tuning of long-context large language models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling LongloRA: Efficient fine-tuning of long-context large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.168960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:26dfa948e4dfc6c1fc9aae3f2045563dd787a797209d9a2cbb2b911e378ad097

Observation 9396d137-9a68-4429-a27a-cc64ca13ab7c · outbound

This paper cites Fo- cused transformer: Contrastive training for context scaling.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Fo- cused transformer: Contrastive training for context scaling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.149159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:d462a058fd40a3ac3ef4ef2120b8e9b50f6252893a6ddc1bba2587ec5cb7356c

Observation 93cd4be0-0cf7-438e-b8c0-6dc22e1c589f · outbound

This paper cites DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Services.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Services

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.690395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:6fb60454c4fe786e2c5fbe4d639759081d4217a45b2e9e96085f47103d666319

Observation cfcaecbc-6ccb-485d-80cb-492b650e2912 · outbound

This paper cites DISC-FinLLM: A Chinese Financial Large Language Model based on Multiple Experts Fine-tuning.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling DISC-FinLLM: A Chinese Financial Large Language Model based on Multiple Experts Fine-tuning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.672488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:11ab306093524636a6b39e8b47c97213a400a520e763669856de69b381d45822

Observation 44f5dced-529e-4bf6-8c7c-58543959a0ed · outbound

This paper cites MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.850554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:fa89760e78df15add44b4b534bc64cbc9dfdb6e02c1cd5770c6e3ea87e89af77

Observation 7a9822f6-5e06-4260-80e4-9860f2969627 · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.858023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:16cb3ac89b55c1870d510eb299daf9d6826de6e3820fc9b1e72f569fe32981fd

Observation 8a11461b-8a19-4f64-b893-9cb44861c518 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.776809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:eeea753060fdba80e5c63639c269c75d3f6e67dbfd5e5d4cbd949f2b86d99260

Observation ea104539-ffd2-4e9a-8df9-1f0db8dcca74 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.837243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:0ec31c033cae3c3da1b4b5a62bba08cd01ac76582625c4e41ec813db5e479258

Observation 4bfd8580-e6f3-445c-a550-23a40ebe2a42 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Vary: Scaling up the vision vocabulary for large vision-language model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.005008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:0ae0ea62aa5115d7df54f6f4315972123ab8670bf93e128340b339305b4f0206

Observation 24203791-56ad-4de6-a38b-b4b5b4e634cd · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.844054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:f2373eaee75e1f6d72dc938ccee3e1da1a3acd9d9c1db46e84e7ecb7d80a1dad

Observation 2787279b-24bf-4388-ba86-2c12f396c7fe · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Hi- erarchical multimodal transformers for multipage docvqa

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.011488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:4916ed950f6611e41501aa59d8de46d26a5e6c29297e1228ad9ac9da0082f424

Observation ec48b73a-84f4-4f5f-8f27-b883ea709152 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Slidevqa: A dataset for document visual question answering on multiple images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.152428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:3102d182e57c574e4876de39a6f9b3ceeae3d464834dc40c0b59f6f95fca02c7

Observation 4e3ff36b-f80b-416a-ab1e-3fa47497083c · outbound

This paper cites Gram: Global reasoning for multi-page vqa.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Gram: Global reasoning for multi-page vqa

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.155764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:1068d12201d66344f29ec09b2fdee626eee86340ddcc8714d37fd997ded01084

Observation 1e4a60e8-f72c-49be-a5a3-4d04c3c4507c · outbound

This paper cites Document understanding dataset and evaluation (dude).

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Document understanding dataset and evaluation (dude)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.138882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:ffa0a031aa5f34b2957946203831858cfbf769b329fdb3e23abb94b71f1eb552

Observation 14cdd7a0-dc02-4776-bf0d-ddf33a65acab · outbound

This paper cites Needle In A Multimodal Haystack.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Needle In A Multimodal Haystack

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.758786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:80e23a91a508fa3a0c3ebc364af033fe24aeae471f18436b5aae201bdecc89f4

Observation 8a0fa990-2ef3-42cf-b723-f59f72c3c5ce · outbound

This paper cites RAPTOR: Re- cursive abstractive processing for tree-organized retrieval.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling RAPTOR: Re- cursive abstractive processing for tree-organized retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.142523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:c68e1541f6c864ae26bba33895d5e1e00c3366f013ef3fe128e9bf15b8a1b897

Observation 5c4d44eb-128a-43f5-8845-ecf38bff1e41 · outbound

This paper cites Unidoc: A univer- sal large multimodal model for simultaneous text detection, recognition, spotting and understanding.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unidoc: A univer- sal large multimodal model for simultaneous text detection, recognition, spotting and understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.132108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:27ad2e941016905e7624338cf5f9ebac04816d7d135353139aa05d5be41e1db0

Observation cd5b616a-9db4-4433-8e40-f861a0281528 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.728217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:0f81cb0b8ed0bcecf269c6e00646204d3d195647269ca18fdbbe16c4c208cf8f

Observation dc054965-971e-4bc6-872d-ed373c5e9791 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.831310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:3f238899f83356c8a6fc2b36a740fa196d6e5f68d5af1fe95880df3a336cfe6c

Observation 0a4f7277-ce3d-4289-a2a8-045c1ca16437 · outbound

This paper cites Llava-next: Im- 9 proved reasoning, ocr, and world knowledge, January 2024.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Llava-next: Im- 9 proved reasoning, ocr, and world knowledge, January 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.135201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:15bf52069142340aa1fea2a7426ac8e4b3c7034378e572f013c6dd549bcaaea6

Observation 3e634f4d-3aed-4300-ae02-6f829d41dc71 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.812998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:138963c39b5d1945d21f4fff5cbf922ca473fb22f09a4945bb93afdb85db16d0

Observation 0811788f-b25a-4749-996f-216d0da37f59 · outbound

This paper cites mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.145975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:72ca850f73560249e17de9abba31a2aa93833c585b40ffae256f01d8d7e2f653

Observation 36e89ca5-4f32-40b2-854b-82a831c01e0e · outbound

This paper cites Cream: Coarse-to- fine retrieval and multi-modal efficient tuning for document vqa.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Cream: Coarse-to- fine retrieval and multi-modal efficient tuning for document vqa

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.159121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:f3d0f11d7c18da50ce89455b46f44bb4fe19bf528a8724b5d8d17b7132878fb9

Observation 844dce37-954d-4756-9ac1-5a32abaeb84d · outbound

This paper cites Efficient attentions for long document summa- rization.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Efficient attentions for long document summa- rization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.162272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:452c0b3d20416109a749346cbf6931e9e82ee37026263cda202304f8fe1f3448

Observation 088c87cb-557d-4c1c-9bde-ecb7754e71ee · outbound

This paper cites A dataset of information-seeking questions and answers anchored in research papers.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling A dataset of information-seeking questions and answers anchored in research papers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.121380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:63fcf2c6d2e225000ff21dbccd5031cd026c61ad3b12ddbf379b4d9cb326bc24

Observation 1b7d594a-7d53-41b2-933e-d1442fd41f19 · outbound

This paper cites Pub- laynet: largest dataset ever for document layout analysis.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Pub- laynet: largest dataset ever for document layout analysis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.114634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:967f7f189404a6f1a07e751e0bd3e1c0d672d44d01bbfa53fba6a0216083e235

Observation cbe6757a-b898-4a41-8f09-6c9fe10f03f3 · outbound

This paper cites DocBank: A Benchmark Dataset for Document Layout Analysis.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling DocBank: A Benchmark Dataset for Document Layout Analysis

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.753536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:4ff96c9cc16a3aef71b6453de002a7a34315b57e4226f408f044ed073cbd3bbe

Observation fcaa281b-b571-42b3-8630-eefd8e131e73 · outbound

This paper cites Doclaynet: a large human-annotated dataset for document-layout segmentation.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Doclaynet: a large human-annotated dataset for document-layout segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.106182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:76e53605a369641cf1f2d8c42060ceec88caffd2c8bc6b2b10234254bb8f0bd6

Observation 702c1049-5825-4624-a60c-fcc7d759a399 · outbound

This paper cites Docile benchmark for document information localization and extraction.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Docile benchmark for document information localization and extraction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.187571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:e6e8afcbc9e4dadffada07ff002e64bc050ad1e640bc099cb042e9782716f32a

Observation e152d2b3-749e-4433-ad17-e80cf27fbcc1 · outbound

This paper cites Cord: A con- solidated receipt dataset for post-ocr parsing.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Cord: A con- solidated receipt dataset for post-ocr parsing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.165556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:815a1e702f392e9a5c1d24dacc4f618bce0ab35dad9c8d941ce97f19e675415e

Observation 1bdb4dcc-044e-4e9a-9466-00e5bcfb68fd · outbound

This paper cites Icdar2019 com- petition on scanned receipt ocr and information extraction.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Icdar2019 com- petition on scanned receipt ocr and information extraction

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.183595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:3fba1ee606f1ada70cb699c88e72a822e0ffe247368ca76d7150ba91a1f55396

Observation 9e40d379-dbc8-4ae5-af87-86a7724d0416 · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.109798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:6842e525f8635d0e5419de887405cc9d6eb4660b517b7348092a833bfdf8a587

Observation bb13bd1e-9a8b-4239-a274-0d231e470592 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Ocr-vqa: Visual question answering by reading text in images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.099494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:e5a03431c8bca01fbead7106e4d2836109eeec0cea362c99e11f09ec5e70dbba

Observation a5cc83a2-013a-4aa5-992d-55e3800a39fb · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.715241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:990f84bfad3f0bfdad2aa24e9468a6d2d4f50bc06050bc5a8e70c5d58d2cf670

Observation b138e72b-54d3-4a0e-bed2-1ff8e8c9b124 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.678200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:d8ddbbac0d2bc69ef798cdecf0a6e7b30e20ddc783fa4f490df4af9b0dde877d

Observation d7c5483a-670c-4b63-bdef-68f7bf107a24 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.783118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:032797d85e8fe3e43fa3398062dc6338ad66af2cf83410cbf4fe12e4ffc913e8

Observation cc4a16dc-a7af-4172-bbd0-fecab8d8d01c · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.096550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:2ec8d2f4827f8a010de0453641d8943d58b292d2a265f58c40d993582b4a1d98

Observation 9cfde41b-d470-4329-91f8-7a527814720c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.721597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:5ce0e0a0dbd4a106b2c807ad7cb63d3b112f411d16f2fba510389775bfcc4e14

Observation 3e3b8f40-4ef9-4520-b393-42e1a00eca51 · outbound

This paper cites Docgenome: An open large- scale scientific document benchmark for training and test- ing multi-modal large language models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Docgenome: An open large- scale scientific document benchmark for training and test- ing multi-modal large language models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.740574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:a0ccca0a0e9df0b1352e799d58d0ea61399fdbeaa194357b58610f6d5fba56c6

Observation d7c11ba4-317b-46ab-a8ea-093ef6059d9c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.806608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:db21d0d957998f49933a1eb8515d44609754beb50a6c28086480dbcf8359fb83

Observation ba688170-4cc9-41be-9210-1982b6be366e · outbound

This paper cites Gpt-4v(ision) system card.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Gpt-4v(ision) system card

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.102491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:9938f7acf35f198a4efe7c7b28b7c0337c70ac4432cdb650ba792a8454d0ee1e

Observation af80f1fe-895f-4d73-b21f-44ec17f08942 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.770260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:6f94595c034edf73ee0c8877dedd4f7538c03ade12df1480783edefca9d363cc

Observation d9c5e129-8a67-4f14-86f2-1161c8ecb1e5 · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Hi- erarchical multimodal transformers for multipage docvqa

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.117609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:b83db08812330c1de9ef15c9ab2c090e858c893a3257e6b6574387d4981a0f77

Observation 0f711f0f-29ac-4e59-8891-f33b65bb980d · outbound

This paper cites https://github.com/kermitt2/grobid, 2008–2024.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling https://github.com/kermitt2/grobid, 2008–2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.081836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:e71b15af994a22fc108593c4b28998f800976ae5dd74156b1f399768601c6d18

Observation 0abb7f0e-61cc-4a4b-83f6-8a7587c6d067 · outbound

This paper cites Mineru: An open-source solution for precise document content extrac- tion.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Mineru: An open-source solution for precise document content extrac- tion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.085541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:924413f09dc1e094a2cdddcba4ba9139d8af5c938cc411f6fabbbe9b1bff9cdb

Observation f0b63aec-932e-42ec-ab88-2ce839bfe3de · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.764585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:1f8029d22a482bbe216ed7f11667e3b6eeb583f5166d1ee0ba65a558baa58c5f

Observation 2361ccb1-027f-4f04-b526-c759b826729e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Gemini: A Family of Highly Capable Multimodal Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.824722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:53322f2846f9f5a9f73ae4947c27d9f3f491cc395d3d5ff9661dc1d80fc41dce

Observation e2029b65-e72e-4344-a28f-76d5685b4cf4 · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.092881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:46adf3ae5de4761e250a336eace69a5b611ac0ec0b032b8909b59806c76a3f26

Observation b4e3fcab-1bab-4762-942f-5647ee71f516 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.795542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:5b49429e614d283518357bdf19628f6352425ee4db4c039f404f891f4cd85859

Observation 22bffec0-53ee-4851-aa26-0bffa56f0da8 · outbound

This paper cites Qwen2 Technical Report.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Qwen2 Technical Report

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.800975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:36490bc754961c55efd286bc45c5e4864237863312422ba6fe4a522ab95ee2f4

Observation 83bfedeb-186b-410e-9eb8-7d66e48f2710 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.078658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:e8ee0290ea0e69ac7f6fb0fead9e6387bf769c8fa7425ad7e6b048c2c4c978a1

Observation c31e0714-37f5-4aed-ab58-a05416c3cda6 · outbound

This paper cites Longformer: The Long-Document Transformer.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Longformer: The Long-Document Transformer

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.818907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:246c8017f2f909da87309cc7c456f1d26e63b2664e7b0dc5e098fe0fb19b4c90

Observation 3d381cd5-5034-4818-bef2-63117802afe3 · outbound

This paper cites Big bird: Transformers for longer sequences.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Big bird: Transformers for longer sequences

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.072366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:7bc265aa847ca87f6cda21fd9daebd561f79af67e8104d6d733268287bbf6fc2

Observation f5878661-be7c-44c6-b3da-a0f878c8a855 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.709537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:bfb6d43a52a6d1d46ff0af1cd57e9aac778796bc3bf2aff75359de4eb2f2d053

Observation 9d3a7fc1-83a7-4d6c-9880-70a9f42332fa · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.703956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:25f194f535a8be5c35d5b50b2414c44a0b2439442b13228b9b92fb33a9f8ce21

Observation 5471ee76-3e2b-4c07-9a6d-32c2bc3c6d0a · outbound

This paper cites Generative multimodal mod- els are in-context learners.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Generative multimodal mod- els are in-context learners

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.075483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:12d69885959bfe21b016c9a8007293694c2a017b98d2cdfbc2e056a6921631ef

Observation 955db559-e0d1-44d1-8684-22cd722d8b81 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.746793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:5bdc20792a2b4cb2acd21da6c47a985ec7b01f952e797841df17d6a7c5f9e751

Observation eb35f914-a198-4dbf-9193-7bd2b7f97aef · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling CogVLM2: Visual Language Models for Image and Video Understanding

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.734159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:57db6e2b55becf813913b76f177d57bcec120c968fa232c216898aae2f96522a

Observation 03c3bda6-d15d-4433-b521-9eefbf563a73 · outbound

This paper cites TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.684066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:733a131db3c35d1ffba5ab20d5dd50d84e76faeda5d0435ec6ab1cfae8ed6c11

Observation 34c9f4d1-ecf7-46ff-838a-a7714a670e9a · outbound

This paper cites Docformerv2: Local features for document understanding.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Docformerv2: Local features for document understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.088832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:71be8bf58df64e254b679eff495f9194f2bfbe662df9cef50b9227a5bb655347

Observation a1f363ee-fd9c-4f4b-9f63-943b70c3ffba · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.175909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:26e231a253ec7373b214c1d90f577a3f89063114e814db999514bf76dded6152

Observation 42287f0c-861a-4685-9553-82626b020e64 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Vila: On pre-training for vi- sual language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.179100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:4498337288073d65146cbef4ba1a172fd7725a10688dfd6fa26bab2e6ac24953

Observation b9ae670d-aa01-42ae-9bc7-828c9f64695f · outbound

This paper cites https://tongyi.aliyun.com/qianwen/.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling https://tongyi.aliyun.com/qianwen/

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.069034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:5ca481bf56e85233f5c7b3ae98e54b8de00c92e04c53c23a8dee0ec00afc11c5

Observation 2531f673-1e30-463a-94cd-f3a3a7929421 · outbound

This paper cites https://chatglm.cn/.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling https://chatglm.cn/

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.045505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:270009bb2dfd90194d29585c8f6fb26f3433bb3945cc098fda7c7052f147f62a

Observation 73ab3032-ac13-4cd4-b822-26e03913d7db · outbound

This paper cites https://kimi.moonshot.cn/.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling https://kimi.moonshot.cn/

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.062593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:48b171818d7fe344ac46bbabb92802cbb5f1b9a07cbe92cabbfc1d0fbb0aa81f

Observation a7cca632-fd65-482c-a8ff-0cbcde187078 · outbound

This paper cites https://gemini.google.com/.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling https://gemini.google.com/

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.057031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:42afe07be98f687b8498412f42695def1a0ce63eda32a9fdefd3d8952ba1e41e

Observation ce7cd5e9-0ee2-45a3-a858-405325b6361d · outbound

This paper cites 用于控制电机,实现循迹与避障。 Evidence 1.底层运动系统的软件设计如图5所示,控制核 心是STM32单片机…进入程序后…信息采集完成 后进行数据处理,控制电机相应转动… 2.图5.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling 用于控制电机,实现循迹与避障。 Evidence 1.底层运动系统的软件设计如图5所示,控制核 心是STM32单片机…进入程序后…信息采集完成 后进行数据处理,控制电机相应转动… 2.图5

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.029488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:bf81bc85936debc6a6926d29bdcc1f4600bf5633e83547ebea6700c2401e5462

Observation 6ebc067a-4801-4ac7-abbc-d74500a3668c · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.033155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:b044a78bb1e216e74a83f86c843e7e560bb0c3d8454c00eeab3977d61c7a26ff

Observation 3e17deaf-cca7-4c39-b34c-7af6ee997b8f · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.036969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:9a9dab0b81e4a703d8707645ab39f90f840df6daf1e9711a8c34d5859a805bfc

Observation c518d369-3b1a-44ff-8a42-f9dc42854ce1 · outbound

This paper cites 2.图1显示了W-Cu二元相图。在1084°C时,相区标记为“W+Cu”。这意味着在这个 温度下,钨和铜是以各自的固相形式存在的。这一点可以通过浏览图中1084°C线 下的相区标记确认。 Evidence.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling 2.图1显示了W-Cu二元相图。在1084°C时,相区标记为“W+Cu”。这意味着在这个 温度下,钨和铜是以各自的固相形式存在的。这一点可以通过浏览图中1084°C线 下的相区标记确认。 Evidence

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.021696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:b0a27e36d0d157338e7cf9b045e1db6e876ed00ae3e5ee496c127c422cb2481b

Observation a84c9a6e-bc19-4f07-a9a5-adbd8812c2ff · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.014653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:a74a37cb3dbdf0f0d0868b0bf1a007dc67c6674018b8515c82074df19f2f4d60

Observation 6803476d-7402-4c38-a96e-f04dbec34176 · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.018246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:d52474a45cc493c41a114cc576f24863fd5f6d11e8c3179be90bf430eb9e0b58

Observation 707bda7d-c98f-4915-b4ce-1381ddd174e1 · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.025433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:75e5ea6436cf75875ec7b45bf0e053f5735eb8fed24fca7d5e6520d79235ec86

Observation 8426d14b-11f2-44a3-b39d-d4fff2a42f3c · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.040226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:c1da8893fbf8e22ed938020be29249aef0b4f54fe595dd396138f5977f51fef6

Observation 2586812c-8d1c-460b-92f0-28fb5b57bb7a · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.065892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:1a9e50b8afd382578d7e5a3d7e805a469e9512075c64af4e35d49d25e8c2655b

Observation 52878216-8501-4bef-a831-087ccb9a52d7 · outbound

This paper cites an unresolved cited work.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-23T19:45:48.172434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:c82e9682607b02ef73f49d8014113d3303cf813fbb101d4d5a7bd213c0a2f7fa

Observation 61bfed9c-e1c8-4e4f-a8c6-9cc0d91bf538 · outbound

This paper cites thought chain.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling thought chain

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T19:45:48.008226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:bad69c0331b239ccffd0f93fa38fa088dd2bd5b8d65c85380b5e4545103781fa

Pith citing papers

Observation 67f46f78-935c-4fff-a7d4-fdb834cb75da · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 267

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:15:47.039772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:3f300e4d4534ff07557ba1d3961a4175dc9bb666d629052a5b28f217a27c6965

Observation 856f71ad-cd05-454d-bcee-156a91a6b51c · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:42:04.371481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:d669577cccea60dfdf93af8840215a48093c9be2875631cd9f2023ce71fd9656

Observation e51baeb6-2042-4ee9-a0e3-a0293a368abd · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.504487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:c29e82b2dafd1e6d6f679a78a9ebfb3fdf37b4574f7d0b529da8084ff7a0ac5a

Observation 9ec09e0c-fea9-4777-8354-bbaacdb71405 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:26:24.394356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:f13189509ce5c658f5904af6cc2a708435f7d56dde0b51771b6f9e952252fe74

Observation f98e27b1-3b68-4450-b015-796f12fdaf20 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:91e75dc56b2d477b8a66a779ae8d193d5c333a625696498282314ac58e71f670