Pith. sign in

Paper Citation Record · LEDGER

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2412.12902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12902 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:40:39.718350Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66ec3495-180b-4d7b-a51c-83e6a5ce15bc · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Docformer: End-to-end transformer for document understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.484541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.484541Z digest=sha256:0d7809b70d04adf5b4cd7412072daec93c146d829d33b4ef7af04e87cb30e857

Observation 8fd7490b-cab5-4011-a31a-31d79b7f9cbf · outbound

This paper cites Visual and textual deep feature fusion for document image classification.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Visual and textual deep feature fusion for document image classification

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.434122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.490094Z digest=sha256:6737e2ad9b800c379b24d0b27dc31262ffca1e6debf66264fbddea0ca578d38b

Observation 9bea97bc-1f37-4130-a7d9-2aaf2f2a2339 · outbound

This paper cites Eaml: Ensemble self-attention-based mu- tual learning network for document image classification,.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Eaml: Ensemble self-attention-based mu- tual learning network for document image classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.420940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.494541Z digest=sha256:89ab0f31ef58c2f9280260f65a147ff107d70ba94b2d944dcff37b2d84bde1e5

Observation 0333c493-8e97-4c62-9425-584319afcd0e · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment BEiT: BERT Pre-Training of Image Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.499670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.499670Z digest=sha256:cb3024b69548c60548949d34f2242bb8e1e7a3a194dd2bb0f04132bf521cdb9d

Observation e9c3414a-0dbb-4003-9c61-65d09ea10892 · outbound

This paper cites Gritsenko, Matthias Minderer, Charles Blundell, Razvan Pascanu, and Jovana Mitrovi´c.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Gritsenko, Matthias Minderer, Charles Blundell, Razvan Pascanu, and Jovana Mitrovi´c

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.407110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.505431Z digest=sha256:7ac695553d69a009394625e3d794a8595ce87cf0e9762f677090b3980aee2ab8

Observation b65dd56a-4296-4227-854b-d3697cd3fa76 · outbound

This paper cites Cascade r-cnn: High quality object detection and instance segmentation.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Cascade r-cnn: High quality object detection and instance segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.391229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.510451Z digest=sha256:a4fa0c3408c55907465e6214d5e0bdd826d0d716df91f94a2475e46a8bc57415

Observation 682cc618-ae98-41f9-bca2-cc528aa83daf · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.515084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.515084Z digest=sha256:5dd42215f8200e944cf32c5351bc7fc8ad4c70feebd8e4151ebc0016701f1049

Observation 8400b773-7954-4b58-8269-14422081365a · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.361518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.520754Z digest=sha256:2e3b671d516df8a8ee117a6ba673a2421974c8afacc77ca68d49b2ae04a6bed3

Observation 49f50b91-07a4-4063-9c8f-30f4706b3949 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment A simple framework for contrastive learning of visual representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.525119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.525119Z digest=sha256:76e23423d655278d61a02854a85708b79b4a842f111e4ea431953c9e18aaeb54

Observation 89747ae7-cac3-433a-8592-12cfd1b1d7e9 · outbound

This paper cites Big self-supervised mod- els are strong semi-supervised learners.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Big self-supervised mod- els are strong semi-supervised learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.327008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.529747Z digest=sha256:177e418badea7492a8aa5e7289dcead73a77f8ec296d434f39320ee7fee812a9

Observation f32e35b1-f074-45e7-a5b6-6eb998f3ce3a · outbound

This paper cites Uniter: Universal image-text representation learning.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Uniter: Universal image-text representation learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.534111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.534111Z digest=sha256:065208c7a4e8baa7bbde02776fa2b5f5639134e4f9bca413c63ff8ae8147c68a

Observation 67578e09-632f-4fae-a273-f183d55d4f1f · outbound

This paper cites Vision grid transformer for document layout analysis.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Vision grid transformer for document layout analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.304685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.538610Z digest=sha256:94a89adbf212ae46dd21dc78b303a486d2efe5be4a10fbaa845f1efb6fc1dbe4

Observation 70680e1c-a978-45ea-84b2-2887f04cb98f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.291726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.542811Z digest=sha256:8cd8f59107aba36fe57add1d72496cb8d55c7b1721d50e15bd14ef9c86c07eb6

Observation b16bcab3-423f-4d11-bfad-252ed5f3bbe1 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Bootstrap your own latent-a new approach to self-supervised learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.546905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.546905Z digest=sha256:d5c4ec351b8e9ae29353c1d6f9e981c9b438631fb945bb743d21aec7d0e9fdb4

Observation e633ae88-9220-4280-8de8-f86f6947bfe1 · outbound

This paper cites Evaluation of deep convolutional nets for document image classification and retrieval.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Evaluation of deep convolutional nets for document image classification and retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.263374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.551384Z digest=sha256:c7ce9b1208d29b4d8eb2d35bddca278c295afd630e8529bdd041d81142bd7e28

Observation 8d25c0fe-45d6-41a7-afdd-f20c1c32357d · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Momentum contrast for unsupervised visual rep- resentation learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.556434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.556434Z digest=sha256:6fe0d45742cf1f0c90a6887bcbb3bd3efbe5af7ec3e3f9fdf6aeba5555c32d6f

Observation a02407b9-2880-48ab-b9de-61c0d382e819 · outbound

This paper cites Masked autoencoders are scalable vision learners.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Masked autoencoders are scalable vision learners

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.236336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.560702Z digest=sha256:50ad1aeaf3fb90c3aed0ad81f86b90b699d8e9ecd540f056b869eaad85491d18

Observation 3383098b-8d62-40c2-869e-743b5fe97593 · outbound

This paper cites Bros: A pre-trained lan- guage model focusing on text and layout for better key infor- mation extraction from documents.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Bros: A pre-trained lan- guage model focusing on text and layout for better key infor- mation extraction from documents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.222480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.564700Z digest=sha256:34c6a996d50d19177fb262a1266a966930dae3d4cee899abfc2ce41748f79f4e

Observation 60f325e4-3b2f-490f-a045-690fcb2b1e8b · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.569129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.569129Z digest=sha256:5577271d6967210df20f78760fb7d8ffb7ce256e7bce5121ea4cd947bc152f10

Observation 323b5738-e0c2-494a-bf73-211acd795769 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.200010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.573650Z digest=sha256:bce93fa1fbc4740f89f4143e9058d84240a6bc805c9c407dc42317c34d221205

Observation 4b133606-b0da-42ab-a007-0f1a1399d918 · outbound

This paper cites an unresolved cited work.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:40:40.184396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.579360Z digest=sha256:048c5d3ccce693a83b3b3f18b67b2800f129c988237ac1d6c6a8062e83203f2b

Observation 444b683c-04af-42ed-a478-d1265eb77b62 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents, 2019.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Funsd: A dataset for form understanding in noisy scanned documents, 2019

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.170443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.584087Z digest=sha256:44b2bd0ff144a2031a462ec5c5812d50442f75ed1bb17f41f1d2b6e38238d4f4

Observation 4d4173ea-df47-4d37-9d8f-d67018a45527 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.588594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.588594Z digest=sha256:5ed2572ee8a79ed9ac6c725dd1af9b6aa99448b195826cdb0b008ea6e2f21f22

Observation ea22a01c-c224-4ff8-8aa3-aaa0034c5661 · outbound

This paper cites Ocr-free document understanding transformer.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Ocr-free document understanding transformer

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.145285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.592918Z digest=sha256:ac4ccecf34222f802c95a450e00ac8241a03ebb280e65b75e802abcae8c327d8

Observation ea3e7d8c-81de-4470-aa51-c34996a3bc19 · outbound

This paper cites Dit: Self-supervised pre-training for docu- ment image transformer.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Dit: Self-supervised pre-training for docu- ment image transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.130892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.597159Z digest=sha256:6b31e0f260ab092e6d7154b09fa69cd3ded94397eba3ed52d89fc3058eae1d18

Observation c8d01da3-b4e8-4be9-ab56-c66f1ed1f1f8 · outbound

This paper cites Grounded language-image pre-training, 2022.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Grounded language-image pre-training, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.601576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.601576Z digest=sha256:dc1d4a8e49cee2874a9f32b8aa90f60f1b9f01b0510455313cb24581b50b7dd8

Observation bfc282d1-3098-4457-bef1-d54bcbc8b982 · outbound

This paper cites Grounded language-image pre-training.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Grounded language-image pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.108412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.605247Z digest=sha256:ad666b6d3d57c4fcdd040d22022fda69a7549597981ada2923345013e65ff4f7

Observation c4734ac9-db7e-4647-8924-73316f72819b · outbound

This paper cites DocBank: A Benchmark Dataset for Document Layout Analysis.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment DocBank: A Benchmark Dataset for Document Layout Analysis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.609154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.609154Z digest=sha256:ea4fd407ae98e43bb3ae41f2fbb16813b476d7010cf6e8707fd2a1280bb1572c

Observation 43b1ff01-eb7e-4580-8647-fe82aa0ad910 · outbound

This paper cites Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:40:39.806354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.613456Z digest=sha256:4288c4c39f164af5eb401daf305833d80d5a65a0f1ee7184932ac5af438f15c1

Observation 74d9812c-b55e-4de7-ab11-fc769ecdbf2e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Docvqa: A dataset for vqa on document images

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.092914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.617708Z digest=sha256:d1037dcfe1e1f0a8c6cbad664d6c58e7da3336728cb68b850ca127817a4df8a4

Observation 560bf376-4f24-4416-a1b2-1b777202dd5b · outbound

This paper cites Infographicvqa.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Infographicvqa

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.080406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.621716Z digest=sha256:e609692165573fd542a1ac6a64e0436a3a5db76235379180e964adc8fa46905d

Observation 953a4027-a9fa-4bcc-8324-f6f675b6ec19 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.625439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.625439Z digest=sha256:00061c45df38d58f25cc427f4e01902ce707e7ad479e98de7f3b80ef047f0583

Observation d869e613-2f33-43ad-9ea8-26a9094cacf4 · outbound

This paper cites {CORD}: A consolidated receipt dataset for post-{ocr} parsing.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment {CORD}: A consolidated receipt dataset for post-{ocr} parsing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.066575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.630141Z digest=sha256:98d38da2188a0a0d892ab348dd78337a70873408bee098d21205a24a20ede967

Observation 93011028-67ea-4c9f-a3cc-353573530b1f · outbound

This paper cites Doclaynet: A large human- annotated dataset for document-layout segmentation.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Doclaynet: A large human- annotated dataset for document-layout segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.053568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.633975Z digest=sha256:765f26d50145d37775685101a6c5e8aaea2e619918b1826ad90e3b21f456846c

Observation 47f192b0-31d2-4a0b-9bac-37cddd362576 · outbound

This paper cites Going full-tilt boogie on document understanding with text-image-layout transformer.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Going full-tilt boogie on document understanding with text-image-layout transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.040932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.637850Z digest=sha256:9e004adae6487a7e4af8db2858a063d72dd95cec8f05c1e139b52693a302122f

Observation fa9145ef-c848-49f8-9306-68dbc95b562f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.641900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.641900Z digest=sha256:d8202ee787fdd79fc45b15849455b3317bc8f90f0105a9e3d3d95328a8e894ed

Observation 62979f90-35aa-4ffc-9a99-a2bb6acb48b9 · outbound

This paper cites Imagenet large scale visual recognition challenge.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Imagenet large scale visual recognition challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.646043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.646043Z digest=sha256:0bbdc2bb04d3cb5ee66644eac8834591d56f9ce74acf0b86c3c9c4a40a4b84f8

Observation 1e2f8194-9003-49d7-8273-f33e93b20574 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.650138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.650138Z digest=sha256:7c5221c61dc5247882c12e2a5d4641b223fc35abb89be30c479278c187c4cfa1

Observation 65695760-db70-443f-aee2-426a6b7ce92f · outbound

This paper cites Complex document information processing (cdip) dataset, 2022.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Complex document information processing (cdip) dataset, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:40.004602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.654274Z digest=sha256:22e5e11fe1afc77c5111c77e7ad1f3c5a6b24405f94f527c6762a254563035ca

Observation 345b4aa3-8ecc-4782-8dd1-45a40c9bd442 · outbound

This paper cites Kleister: key in- formation extraction datasets involving long documents with complex layouts.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Kleister: key in- formation extraction datasets involving long documents with complex layouts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.990257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.658478Z digest=sha256:836976a3e633bf87f99f45c199efe4137f9fe7af878e91bc907264d136d7d943

Observation 5cabec59-b2ba-4b11-8a65-3bbbc638f022 · outbound

This paper cites Vl-bert: Pre-training of generic visual- linguistic representations.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Vl-bert: Pre-training of generic visual- linguistic representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.977499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.662701Z digest=sha256:2484881fdd587aa32ddbe584a57aa4b08d5194781ba774e960e8606ec9a0ce2c

Observation 2028df8f-6ae4-4a69-aa05-8d38aa65529e · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Revisiting unreasonable effectiveness of data in deep learning era

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.964964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.667525Z digest=sha256:11ac8df1c4eca4b320a74a958a710e405713de26cbfa0d9658f63cd8fbfb9c59

Observation 8e70fc14-c584-4731-89ec-6bc277c3b6d3 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Unifying vision, text, and layout for universal document processing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.952752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.671585Z digest=sha256:8a603da024d128412fec5b47378e959b372b588f9170e94bb6dee59034f277eb

Observation 7bb084c8-b4ae-4cd3-9a6d-86c1bb614771 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Yfcc100m: The new data in multimedia research

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.675226Z digest=sha256:78da255ff0f5db3c29f5f29a927dd6fc7ef7a59afe358c584a51cc75c5a49543

Observation c4e96163-fed6-436d-98d9-d8d38625d646 · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention, 2021.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Training data-efficient image transformers & distillation through at- tention, 2021

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.679383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.679383Z digest=sha256:884f4a1dc6d6196774188f4bd0d81852d2f56e8a777ff0dddbe9695bec70943c

Observation 91bc9e1f-cdd3-455c-81da-290aef275597 · outbound

This paper cites Detectron2.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Detectron2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.683469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.683469Z digest=sha256:98826f1ca7e50ec6cb08619e73e251cebf0da7b415e1b0f924463a17d2b42810

Observation b5a48ba7-8102-4142-8dab-dd75a01539d3 · outbound

This paper cites Aggregated residual transformations for deep neural networks, 2017.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Aggregated residual transformations for deep neural networks, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.917601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.687287Z digest=sha256:41b33a39b44448bd613f7da5609735a8b7cae4a26a36c6ff54c84d74d7ea2659

Observation 90ce9152-23a7-469c-bfe6-609d473f6bb1 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Layoutlm: Pre-training of text and layout for document image understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.691284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.691284Z digest=sha256:fbef95cb5f667567954c29ba3289af6588e454767a8863989d61b8e4be8a41d9

Observation d8d7e7f7-c9b3-4c1b-a072-821c6f2242fb · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.695400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.695400Z digest=sha256:5350c3c4b9011933fc90b8ec2fa5f3005082eaf423cd16e5dba016854c037058

Observation b6ede840-0a43-439c-b526-01ee219b47c5 · outbound

This paper cites FILIP: Fine-grained interactive language- image pre-training.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment FILIP: Fine-grained interactive language- image pre-training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.895575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.700181Z digest=sha256:6a58c412dcc4d952036b92a517ef2c9b58da0b73254a626c5325d5f48beb0b90

Observation 3c7e747a-4390-4ab8-b9ed-b9d9b70a1ea0 · outbound

This paper cites StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.704209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.704209Z digest=sha256:82979e0e551c3a1131ce6c70ba3b4f5b9f54474c69e632ad7b9a31d07685ba26

Observation 97cf0576-2402-4cb0-9cd7-0e7c19cab44c · outbound

This paper cites Sigmoid loss for language image pre-training,.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Sigmoid loss for language image pre-training,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.709000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.709000Z digest=sha256:77606c83d5c514663050621aae7e5b2e30126cd2d18b10461a9f1a37e4353490

Observation 2c154ed6-f2a5-4984-8f95-5f5575dfe960 · outbound

This paper cites Zhang, H.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Zhang, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.868641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.713836Z digest=sha256:b81fd4d447215e0e7eefa3157aff32c586a5ff30f69f5528e9fcf0ff52de84e9

Observation 58d87be3-8112-40c9-9d50-d2d0c2364bd7 · outbound

This paper cites Pub- laynet: largest dataset ever for document layout analysis.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment Pub- laynet: largest dataset ever for document layout analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:40:39.850776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:40:39.718350Z digest=sha256:c9bd1d4502029ea6948ac2e6f7daca22297f9cc7f30d6a1b89082bb32f0b9bc1

Pith citing papers

No inbound Pith citation observations are available.