Pith. sign in

Paper Citation Record · LEDGER

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.11576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11576 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.990881Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:37:30.147692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4f8d34df-db08-49f0-9a62-eb1f29194697 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.990881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:26.990881Z digest=sha256:eb80d3b12a007c9c4b2a852a39d8c31fa444bbd2d89aa774aed5e0a6f7a1294d

Observation dc6b1429-fcc0-4f58-87ac-341cefcf2677 · inbound

PaddleOCR 3.0 Technical Report cites this paper.

PaddleOCR 3.0 Technical Report SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:24:28.327571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T23:24:28.094041Z digest=sha256:a771d443600d03d595659dcdcaa5674a85af7f57d57add92550aabb16189b816

Observation 02a6f8ee-4043-4f86-b4ad-8467fcba4c50 · inbound

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation cites this paper.

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:18.495715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:10:18.495715Z digest=sha256:8869b0920cbbeafafde08e49ecb4f1fd50259e7d6782a93a3f70bd6584f5e447

Observation 059b3d4f-e097-489e-adee-4c8cced1226e · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.940165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.940165Z digest=sha256:d6676b82c67916ced57c16d0fba595dc62cc571ffe40bb265b7c5659becf5653

Observation 0b8afe44-8a1a-46f0-bb21-6192ce1ac89a · inbound

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing cites this paper.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.170342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.170342Z digest=sha256:6bf53f1f6935ec90a76e80b87edbb968baadea505908506f18d3bb9dd5789e70

Observation 2243c99a-0e45-49f1-8fff-ddf6d83300a6 · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:32.028685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:3553849ffbf82733cadf06fe40854b4061df6af5f873d43907f02e6836127461

Observation fd30d2fc-68a4-447b-be80-4c3bb7adff4b · inbound

DeepSeek-OCR: Contexts Optical Compression cites this paper.

DeepSeek-OCR: Contexts Optical Compression SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:35:50.751259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:35:50.647950Z digest=sha256:bd209b94dbed797068763c32dd1642fbf713abf0ea34cda4e1e917f67a2633e2

Observation a8e5ba99-f8e0-4699-9f45-d7b49523c9b2 · inbound

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction cites this paper.

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:35.295207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:35.295207Z digest=sha256:657d2b1ab575829dacda796d65644c1ed59b204d31c8072a7c8ae6ccca9b9f3a

Observation 1a8d3d5e-c9d9-4741-bad0-dea1ffb02914 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.805641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.805641Z digest=sha256:30397cc37ff1529d9d68c9716dbf6c780730db9161ca15ea989c8f4726158991

Observation 52ed26b1-3a17-4224-8537-68037e5b214d · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:42.627628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:42.627628Z digest=sha256:60262dec67b78efc0013292b86451b53ee8f52a1508bf140cf078f5b332f077b

Observation 914ea71a-7da4-48d3-9742-7e72075b1ed7 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.638872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:241f2201be15ccfdaebef344c03b0241c241d6b20dec508c5b49ffa5353ece68

Observation 433cff05-eac4-4d32-b500-173a5f898e30 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:02:58.627073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:2c32087ea7ebe18fcbcc5faab87abd16156dbee4c4a6b87d315f67ce73a38e9a

Observation 6db6664a-9a7e-45e2-93d5-55f7e1410f98 · inbound

TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications? cites this paper.

TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications? SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.663288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:23:34.256279Z digest=sha256:6f2bf64b991ef5cdea8450608c59932fd2503b424d69c23f3e173f9e3c938c48

Observation 8ddd3b6c-0e70-4645-914d-c65499a583f2 · inbound

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding cites this paper.

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.101834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:48:34.771799Z digest=sha256:446a3489ecc9569db4e4cd183cf7cd1597a94e3f7969d69f8122331c5be036b9

Observation 2cf726a6-1fe4-4d53-b916-8553cd54bc25 · inbound

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction cites this paper.

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:37:30.149616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:06:33.518620Z digest=sha256:4ac9034dda87eb7f841193c2908402475bd364ae08a536263a60fdae6b4d851c