Pith. sign in

Paper Citation Record · LEDGER

DocFusion: A Unified Framework for Document Parsing Tasks

As of 20 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2412.12505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12505 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:51.807410Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:27.675801Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:42:32.291738Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9219257-a460-45f9-8cf5-633234105552 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

DocFusion: A Unified Framework for Document Parsing Tasks Nougat: Neural Optical Understanding for Academic Documents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.712943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.712943Z digest=sha256:8fe9c593bfbfc06610e5e50693ecea529b8bdeaa239b26c0e7331da14de4deef

Observation 7735b225-1c7f-43af-a4ab-9a9bf45db3dc · outbound

This paper cites DaViT: Dual Attention Vision Transformers.

DocFusion: A Unified Framework for Document Parsing Tasks DaViT: Dual Attention Vision Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.729549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.729549Z digest=sha256:1d210746d2d8d2d730e0aa72c8c05c72b3591a6c0ea70876e654a2bc955db365

Observation f0e9cf88-61f0-4e5c-aec6-84e7e3da4ffd · outbound

This paper cites LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking.

DocFusion: A Unified Framework for Document Parsing Tasks LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.738129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.738129Z digest=sha256:f34ea36aed0ee445fc60fa21adc401c82f69d62e4c4403ac2659fb5baae9112d

Observation ddb24156-520e-48da-87e5-5a8b84046204 · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

DocFusion: A Unified Framework for Document Parsing Tasks YOLOv11: An Overview of the Key Architectural Enhancements

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.742600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.742600Z digest=sha256:49662fc7eff534af7d74804b08bdd8caca7a53d9d21bcd9993cf717c4b73c52d

Observation 07c475fb-f1fd-4ec2-835c-2a6c7b53af61 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

DocFusion: A Unified Framework for Document Parsing Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.746770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.746770Z digest=sha256:5a90b5ceef4b0744aa76716d51b34109b3f96bf41a9f02a8f18aa8340c24669d

Observation d95ab6b9-88fb-4454-992b-9a894289b9da · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

DocFusion: A Unified Framework for Document Parsing Tasks TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.751083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.751083Z digest=sha256:a240b242ff70218898801c39db2df548e853f45eb69baab477c4ad85441bb15e

Observation 6672cdf6-4fef-4255-9a2b-d12bd3b4d655 · outbound

This paper cites Preprint, arXiv:2406.17148.

DocFusion: A Unified Framework for Document Parsing Tasks Preprint, arXiv:2406.17148

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:52.069796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.754902Z digest=sha256:2855d969ba22bdd0a59c547948321b2ef0dd4ea0b62cd92eb2f08e2cb740e6fd

Observation f14083fb-eb4d-4f45-984a-5e3af3918694 · outbound

This paper cites Visually Guided Generative Text-Layout Pre-training for Document Intelligence.

DocFusion: A Unified Framework for Document Parsing Tasks Visually Guided Generative Text-Layout Pre-training for Document Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.758577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.758577Z digest=sha256:b30c6596cc6d7536e4b72b2ab5b8b098a980bb955efa05321f7e4882b45649b9

Observation b29c159f-46eb-4936-809a-6c78bc6dacd7 · outbound

This paper cites Accessed: 2024-02-29, cited in pages 1, 2, 4, 6,.

DocFusion: A Unified Framework for Document Parsing Tasks Accessed: 2024-02-29, cited in pages 1, 2, 4, 6,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.190983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.762483Z digest=sha256:e5c2d986c09711415ea8c55cd55fa071463705ebd865358730874882d8dd59ff

Observation be84449a-d9f5-4460-b4f2-16ba2cfc86a7 · outbound

This paper cites RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking.

DocFusion: A Unified Framework for Document Parsing Tasks RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.769933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.769933Z digest=sha256:876ad412bd570498d621a5c5996bab055b331dc4f9f1de3d7b2b0be3cb7a3d7e

Observation 3a165f9f-8af0-404f-9027-4c9205e448a9 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

DocFusion: A Unified Framework for Document Parsing Tasks General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.777751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.777751Z digest=sha256:d4cd6b062b4224ce994550923e4d054df2df858863a8afd001f20fbfe0fe07b3

Observation c52db8e3-b3d1-4815-8219-9ecf91af8376 · outbound

This paper cites arXiv preprint arXiv:2406.11633.

DocFusion: A Unified Framework for Document Parsing Tasks arXiv preprint arXiv:2406.11633

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.785069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.785069Z digest=sha256:aac3db03ce642f8f5c5c31fb8abcabfaf4ccd45b5815df8caf6a9a3b6bc7b9e0

Observation 114ebe78-1c71-4478-bba7-2c83e5ff44dc · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks.

DocFusion: A Unified Framework for Document Parsing Tasks Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.788539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.788539Z digest=sha256:a91afa92fbd5f105b1962e97809527414c6a271e6ec950dff26170957414a317

Observation eba26595-eb38-46da-afe9-ed752e090211 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

DocFusion: A Unified Framework for Document Parsing Tasks LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.792406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.792406Z digest=sha256:93afbfce121272353a364d99250b544c952d5bc073e6e6a8714fee37b919872f

Observation eea32079-4f42-49f9-8e56-720633a0d8b3 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

DocFusion: A Unified Framework for Document Parsing Tasks UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.796087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.796087Z digest=sha256:2f7e1a824efee93b3dd05a284e7693add675aed6f0b06fb86c15bfc9e8d99de0

Observation 2ec46c1f-b940-4f02-a4a2-341146ca7eb1 · outbound

This paper cites Syntax-Aware Network for Handwritten Mathematical Expression Recognition.

DocFusion: A Unified Framework for Document Parsing Tasks Syntax-Aware Network for Handwritten Mathematical Expression Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.799769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.799769Z digest=sha256:b7c798287c505252803b40492f5d373d8c7e0850bf39c3040547a510d11e67ec

Observation 3930debd-acc2-46a0-9c65-71547acf2492 · outbound

This paper cites Adversarial Retriever-Ranker for dense text retrieval.

DocFusion: A Unified Framework for Document Parsing Tasks Adversarial Retriever-Ranker for dense text retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.803689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.803689Z digest=sha256:9aefa93df7b057b7d1355ccff3c043ca2321aac63ad41713ccc8caeea8744a9a

Observation c693a7b4-4e5e-4b55-b153-e28abb9ce2f4 · outbound

This paper cites However, Latex has not been the mainstream approach for TR tasks in recent times, resulting in a limited number of TR models available for comparison in the main experiment.

DocFusion: A Unified Framework for Document Parsing Tasks However, Latex has not been the mainstream approach for TR tasks in recent times, resulting in a limited number of TR models available for comparison in the main experiment

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T14:04:52.158585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.807410Z digest=sha256:5465e21156fa13f06a9689e6cae504872c1c2219a5a1ddbacafee996e8e1b1e3

Observation ca7fc0be-3469-4690-afdd-b1d369141bcf · outbound

This paper cites YOLOv10: Real-Time End-to-End Object Detection.

DocFusion: A Unified Framework for Document Parsing Tasks YOLOv10: Real-Time End-to-End Object Detection

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.773913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.773913Z digest=sha256:6223af8480ceb8608ec42b1578c22359bf25f577acb584bab1e3252222e425d3

Observation ec6acdb3-6e39-4036-88a0-2c5cd12af7cf · outbound

This paper cites In 2017 IEEE International Conference on Computer Vision (ICCV), pages 2223–2231.

DocFusion: A Unified Framework for Document Parsing Tasks In 2017 IEEE International Conference on Computer Vision (ICCV), pages 2223–2231

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.211337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.717239Z digest=sha256:d72e86ef935c77d34ef47a0b5fae6c6878b42510b4578b920b22acbf19569858

Observation 0d2cea3a-35c7-470b-8072-70fc008ad379 · outbound

This paper cites In 2018 13th IAPR International Workshop on Document Analysis Systems (DAS).

DocFusion: A Unified Framework for Document Parsing Tasks In 2018 13th IAPR International Workshop on Document Analysis Systems (DAS)

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.169695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.781419Z digest=sha256:bc6d48419448d64bae7371db5fc4dab53a0064350eda575c6ba31281bcf69918

Observation 0533ed28-c0ca-424c-ae7a-c7841daa0374 · outbound

This paper cites In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 142–147.

DocFusion: A Unified Framework for Document Parsing Tasks In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 142–147

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.179841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.766461Z digest=sha256:ef8430e4d0bee23cef0ff8550d0daa7fa914eb213327ee6b302cdacc3cff59ef

Observation f6f59ad0-6610-4672-af1b-dc68a737765c · outbound

This paper cites End-to-End Object Detection with Transformers.

DocFusion: A Unified Framework for Document Parsing Tasks End-to-End Object Detection with Transformers

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.720965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.720965Z digest=sha256:f69ca2a6f0eb38d82b4112cc5ed40d4606d473d10d4135dff3e8b81121d44423

Observation 6054a1b2-20e5-4d23-b4e1-69d19817b9ca · outbound

This paper cites In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

DocFusion: A Unified Framework for Document Parsing Tasks In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.200948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.725583Z digest=sha256:ec52d567a9a833093da40381a652c893557518ab87490658a845633f6885d6c1

Observation 90205521-6102-4b43-93cc-802a336fa179 · outbound

This paper cites Accessed: 2024-02-29, cited in pages 1, 2, 3, 7, 10,.

DocFusion: A Unified Framework for Document Parsing Tasks Accessed: 2024-02-29, cited in pages 1, 2, 3, 7, 10,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.222424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:04:51.709022Z digest=sha256:1c139c0f2bc4da6ed08a088a2a9532a5474e6dd04b023c0137724cae7d4cbcfb

Observation 6c35c1eb-d621-4119-963a-855ec8ee9c95 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DocFusion: A Unified Framework for Document Parsing Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.704202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.704202Z digest=sha256:3a1953ff9a9ad203eb3bab2121eebd16361419cd6ea8840892ae8171e96555d2

Observation 335f972b-dc07-42f6-8aa4-5404dc73ac78 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

DocFusion: A Unified Framework for Document Parsing Tasks Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.733707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.733707Z digest=sha256:0eb2e488dc20c9e81b732ade04168d09ef070aacc8479d4f9098bef41401b815

Pith citing papers

Observation d20b8933-9a52-46b1-bfc6-e781c9329f0a · inbound

Improving RL Exploration for LLM Reasoning through Retrospective Replay cites this paper.

Improving RL Exploration for LLM Reasoning through Retrospective Replay DocFusion: A Unified Framework for Document Parsing Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:27.675801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:27.675801Z digest=sha256:9b1d4f31db07396b62f856fdaaf3c442cf57a85aee4c750b9a1f50bb233941df

Observation d0af0262-56e8-4e72-8c9b-7a894f0d0fd4 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting DocFusion: A Unified Framework for Document Parsing Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:42:32.401854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T15:42:25.216463Z digest=sha256:84110827dc49b6b49058ad41d512502947bc779acc997bbe37de20a3f738bb9c