Pith. sign in

Paper Citation Record · LEDGER

KOSMOS-2.5: A Multimodal Literate Model

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2309.11419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.11419 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:52:10.280261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T13:27:05.552868Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d71fdbcf-0d36-4319-b658-ea545a8cba30 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites KOSMOS-2.5: A Multimodal Literate Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.153006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:795d0628fb960e561223116cd46a0b9ebbc5bd4b38765b94633717cb7234222f

Observation 01e0e38b-26fe-4082-b818-ccec1f9638a5 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction KOSMOS-2.5: A Multimodal Literate Model

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:00:25.685643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:f33d9701046a93e8788dd47a0541339887cdd1ace73aa568179dd8dd03f5e620

Observation 227ab008-0fa0-484c-aa81-4b0e5b2a84ea · inbound

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE cites this paper.

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE KOSMOS-2.5: A Multimodal Literate Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:10.280261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:10.280261Z digest=sha256:17023fba30d142184add80f4c6e0693d521a27689165a9812cd6057a09a3172f

Observation 55824c0c-9478-463f-8219-ebe694609b52 · inbound

DOGR: Towards Versatile Visual Document Grounding and Referring cites this paper.

DOGR: Towards Versatile Visual Document Grounding and Referring KOSMOS-2.5: A Multimodal Literate Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.768116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.768116Z digest=sha256:886af7b0fa24e91bc104affd9470f22e4ef596847cc8a0736f0a87f9538b50a3

Observation ec05580b-55f9-47b3-8f33-4751a25e6efa · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding KOSMOS-2.5: A Multimodal Literate Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.711726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.711726Z digest=sha256:6b8f7bef0383288bf3625dae32bd890212af879bf88fef978cc8d5f779fd2926

Observation 92df3bc4-0195-4ce9-bdc4-f6412729c6ed · inbound

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy cites this paper.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy KOSMOS-2.5: A Multimodal Literate Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.158358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.158358Z digest=sha256:e5c603cd86861baa0c2ee7ee8ed2a7e89844d9164f6cfb32668eda20e0da36cb

Observation d14e1eef-9b97-4868-8462-b5c5073baecc · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? KOSMOS-2.5: A Multimodal Literate Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.587776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.587776Z digest=sha256:44ce5092a0d5689e7ba3d97d436f785526a189d4dd25fe7d016b127d9fa8614c

Observation 6c523022-cb1e-44f8-b80a-60d4e97f94e9 · inbound

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding cites this paper.

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding KOSMOS-2.5: A Multimodal Literate Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:03.255104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:03.255104Z digest=sha256:5d0cb7264ad741896f9a9c72b4a7db97108b41b2edada55c4cd4a895da2cc1e4

Observation ee1b0aa8-656e-4374-b20a-32bffae4b191 · inbound

InstructOCR: Instruction Boosting Scene Text Spotting cites this paper.

InstructOCR: Instruction Boosting Scene Text Spotting KOSMOS-2.5: A Multimodal Literate Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.070213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.070213Z digest=sha256:92a172b509cf7a5b1325cd7bf78a1e2df58a5a1b47e52329e14bc7ac9b257afb

Observation 9eca4f22-337d-46d1-8a28-0836cec76d99 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey KOSMOS-2.5: A Multimodal Literate Model

Reference 285

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.395890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.395890Z digest=sha256:b37c42e0423a4d9c9ac44e48bec2d1134a117c98015189f6b56f39e71d9a82b9

Observation 46f1f4dd-0b04-4f43-867f-71dad31aed6a · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends KOSMOS-2.5: A Multimodal Literate Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.865103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.865103Z digest=sha256:245c184f582d42390e560f4be98b1ef8e7a05664bd94055b989d3f99dfe49977

Observation f092e221-6002-4759-95c0-d245e880bfbc · inbound

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents cites this paper.

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents KOSMOS-2.5: A Multimodal Literate Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T23:13:13.922039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:13:13.922039Z digest=sha256:68312243dd5eb9824ffcc7337bc034f07b99629735abcb1ab41e6e928b5b13f2

Observation c6d922dc-292b-4119-a33a-29d6dcdd10b1 · inbound

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System cites this paper.

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System KOSMOS-2.5: A Multimodal Literate Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T20:20:13.222393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:20:13.222393Z digest=sha256:65eb07e3a8e276c884c7adfe8e98c00ba0ccf30744bfbdcfd8fed25bfeade0a8

Observation 875dff2a-34fb-4356-b4d9-97215fd81879 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting KOSMOS-2.5: A Multimodal Literate Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.594364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:26.594364Z digest=sha256:fde479003e327028356ec5863f515d46560f0ab4ac973a396d78c8222a8bdf2a

Observation e759238f-dab3-4c82-8923-2bcb3f9e021a · inbound

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling cites this paper.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling KOSMOS-2.5: A Multimodal Literate Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.299853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.299853Z digest=sha256:03f001dc61f17c0351a69800e37b42377cf85b3e0337b18ecaf6cda52027e8f3

Observation 195cbc1e-4303-4036-aab0-dde0d903b056 · inbound

DREAM: Document Reconstruction via End-to-end Autoregressive Model cites this paper.

DREAM: Document Reconstruction via End-to-end Autoregressive Model KOSMOS-2.5: A Multimodal Literate Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.336576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.336576Z digest=sha256:04e4d4be01426121724da51fc2f3b81f846fbee46c80cbe1231560eac6f23312

Observation 89cb16e3-c619-4a7a-a2b7-0f6800ab48c9 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends KOSMOS-2.5: A Multimodal Literate Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.410171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:5d717994054bae69df0fbe18fa5ec8e0b6f69c301ea0e19798133384a7b0b812

Observation 93b1e719-603b-4d90-b458-ff54a1b5a326 · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models KOSMOS-2.5: A Multimodal Literate Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:55:19.445174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:3108c16b51792fa6da2912cc4ff3153bec913312d729e765d72f3af1e0fdf3b1

Observation ac398883-1f0d-41be-9a61-1671c3d694c7 · inbound

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction cites this paper.

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction KOSMOS-2.5: A Multimodal Literate Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:14.448490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T06:34:56.032634Z digest=sha256:7605633d31b988135cb73667016e147c8b2a6368504571e4980fc816b36552b8

Observation 5e2fee03-ac49-4edd-807d-a17ef55f27ec · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing KOSMOS-2.5: A Multimodal Literate Model

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:48.986357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:ce6c15f70a767b8a3516780d4858aac5bd270bc69f48370c5adab28ea7c78bb1

Observation ccfa463d-4a18-4714-8688-5414e90469d6 · inbound

Towards Characterizing Scientific Image Utility and Upgradability cites this paper.

Towards Characterizing Scientific Image Utility and Upgradability KOSMOS-2.5: A Multimodal Literate Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.869035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T11:10:37.254428Z digest=sha256:16c9bad668e81f34863c25e3cc2698294558b89b2288fc54ff88a3a836c7d577

Observation c7a1dff8-f4d8-44f0-9374-2abae61392f0 · inbound

Vision Language Model Helps Private Information De-Identification in Vision Data cites this paper.

Vision Language Model Helps Private Information De-Identification in Vision Data KOSMOS-2.5: A Multimodal Literate Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.600332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T16:29:33.294960Z digest=sha256:5c3ee7ce5d9b5f4a36dccd08ca1ce928d48da892af0fde2673a725cea532f908

Observation 7ef6317b-b9f2-48b7-93ea-d1db3b4d17dc · inbound

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis cites this paper.

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis KOSMOS-2.5: A Multimodal Literate Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:44:22.134761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T07:37:05.806389Z digest=sha256:ee054361972f70120f0a5220328da48f99b6cc57ce6c3ec9b0f3cd8d6703726c

Observation aa2516aa-7cdf-4590-b65a-4475a4517604 · inbound

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs cites this paper.

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs KOSMOS-2.5: A Multimodal Literate Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:57:03.927131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T14:47:26.322568Z digest=sha256:fde293414a155d2cffbd7a2a5352b41017d0c5c9c2728b7e8a3eff39a39fd0c8

Observation a1da6840-be94-461b-babc-ab510dd557ab · inbound

MORE: A Multilingual Document Parsing Benchmark and Evaluation cites this paper.

MORE: A Multilingual Document Parsing Benchmark and Evaluation KOSMOS-2.5: A Multimodal Literate Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T05:51:08.895411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:51:08.895411Z digest=sha256:5bb474184f489e0d8cb7958120206fdfb5dba1a8ae41254b04f57f4a3b15f748

Observation 9127f1fe-4245-49f1-9952-7e6a7468c280 · inbound

Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment cites this paper.

Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment KOSMOS-2.5: A Multimodal Literate Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:27:05.554903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T13:26:41.210163Z digest=sha256:5cc94391e66d67ca6a2df752a9d757647017569b4ecad0c04387925986b39721

Observation ad987fdd-86ef-40e7-a4a4-e14abcc0413e · inbound

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study cites this paper.

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study KOSMOS-2.5: A Multimodal Literate Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T02:58:30.989588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:58:30.989588Z digest=sha256:fb42b28c72bff6bccc794f29ae645a98a3dc5a902f45a052b3633f9698a7491d

Observation af234a1e-faf5-40b5-8d18-930e6cf4014f · inbound

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment cites this paper.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment KOSMOS-2.5: A Multimodal Literate Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.575255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.575255Z digest=sha256:f868505fd298b20eb72fc94cde896f5ced8052652070a869a86ac28c617204e1

Observation 6fcf668d-273b-4f2b-a57d-2c8e85e97bd3 · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth KOSMOS-2.5: A Multimodal Literate Model

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:38.554048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:38.554048Z digest=sha256:b6934ee1330b797f8b68099e1b3341dab5b51daa5d12050f977a68d5cdfc741b

Observation daa62925-ab67-491f-b611-66f55496c011 · inbound

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards cites this paper.

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards KOSMOS-2.5: A Multimodal Literate Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:40.903096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:46:40.903096Z digest=sha256:ea05713de430b22430f6fa5b6c574ea34883b0ef8ed2b4cdd0b26e10fda20358

Observation 9f2100cf-6c2e-4745-baf1-9640a5a0ec1d · inbound

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards cites this paper.

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards KOSMOS-2.5: A Multimodal Literate Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:55.351147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:55.351147Z digest=sha256:05aa34051a3ca8d21b0050c3118f3a90201ffb4f72b8299f2e269d5cdbab694c