Pith. sign in

REVIEW 16 cited by

Docling Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.09869 v5 pith:RRBHMRJW submitted 2024-08-19 cs.CL cs.CVcs.SE

classification cs.CLcs.CVcs.SE
keywords doclingeasymodelsreporttechnicaladditionallowsanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This technical report introduces Docling, an easy to use, self-contained, MIT-licensed open-source package for PDF document conversion. It is powered by state-of-the-art specialized AI models for layout analysis (DocLayNet) and table structure recognition (TableFormer), and runs efficiently on commodity hardware in a small resource budget. The code interface allows for easy extensibility and addition of new features and models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Deleting structural announcements makes following prose harder for LLMs to predict, swapping notation does nothing, and the paper proposes a pure-frame format that strips announcements into sidecars.

  2. The Documentation and Traceability Burden of the Indian EV Transition

    cs.CE 2026-07 conditional novelty 7.0 of 10

    The paper systematises India's EV compliance-document lifecycle into a two-layer evidence model, a six-stage lifecycle with four failure loci, an exergy-destruction analytic lens, and a six-problem research agenda.

  3. Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

    cs.IR 2026-07 accept novelty 6.5 of 10

    A multi-domain corpus of 5,264 scientific GitHub repos plus two IR benchmarks (219 expert queries; 117,950 snippets / 119,720 queries) shows large domain- and documentation-driven gaps in scientific code search.

  4. OmniPresent: Generating Coherent Presentation Suites from Scientific Papers

    cs.SE 2026-07 conditional novelty 6.5 of 10

    A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.

  5. Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Analysis of EMNLP 2025 Responsible NLP Checklists shows 44.9% of NO justifications are brief or empty and 15.5% of main-track checklists contain parent-child logical contradictions.

  6. Scout: Scalable Document Extraction via Data Similarity

    cs.DB 2026-08 conditional novelty 6.0 of 10

    Scout generates and refines LLM-written span-locating rules to match full-document agent accuracy on document extraction at orders-of-magnitude lower cost.

  7. Measuring research data reuse in scholarly publications using generative artificial intelligence: Open Science Indicator development and preliminary results

    cs.DL 2026-04 unverdicted novelty 6.0 of 10

    An LLM-powered indicator measures research data reuse at 43% in publications, exceeding rates from established bibliometric methods and implying that benefits of data sharing are currently underestimated.

  8. Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents

    cs.DB 2025-11 conditional novelty 6.0 of 10

    A new benchmark with two new datasets and end-to-end metrics shows that table extraction from PDFs is still unreliable across heterogeneous layouts.

  9. SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    SARA combines natural-language snippets with semantic compression vectors in RAG to improve answer relevance, correctness, and similarity on 9 datasets across 5 LLMs.

  10. The Hidden Threat in Plain Text: Attacking RAG Data Loaders

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Invisible characters and formatting tricks in ingested documents survive popular RAG data loaders and can manipulate end-to-end RAG outputs.

  11. Improving Access to Historical Archives with Real-time RAG-based Systems

    cs.IR 2026-07 conditional novelty 5.0 of 10

    LLM post-OCR cleanup plus dense retrieval and cross-encoder reranking raises NDCG@10 by 31.9% over BM25 on 500k Swiss newspaper segments while cutting CER/WER substantially.

  12. Making Sense of Data in the Wild: Data Analysis Automation at Scale

    cs.IR 2025-01 conditional novelty 5.0 of 10

    A multi-agent LLM system with retrieval-augmented generation automatically curates datasets from Zenodo and Hugging Face, yielding small retrieval gains and a confounded synthetic-data improvement.

  13. Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Typhoon 2 improves Thai LLM performance through continual pre-training on curated Thai data and post-training, releasing text, vision, audio, and safety models.

  14. Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Across three invoice datasets, multimodal LLMs extract fields more accurately from raw images than from markdown converted by a parsing tool, with Gemini 2.5 Pro leading.

  15. Knowledge Graph Fusion with Large Language Models for Accurate, Explainable Manufacturing Process Planning

    cs.AI 2025-06 reject novelty 4.0 of 10

    A retrieval-augmented framework that auto-constructs a CNC machining knowledge graph with GPT-4o and uses graph retrieval to make small LLMs answer process-planning questions accurately.

  16. Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery

    cs.CR 2025-07 reject novelty 3.0 of 10

    Scout applies off-the-shelf LLMs and vision models to triage digital evidence, but only anecdotal examples are shown and accuracy is withheld.

Pith tools