REVIEW 16 cited by
Docling Technical Report
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This technical report introduces Docling, an easy to use, self-contained, MIT-licensed open-source package for PDF document conversion. It is powered by state-of-the-art specialized AI models for layout analysis (DocLayNet) and table structure recognition (TableFormer), and runs efficiently on commodity hardware in a small resource budget. The code interface allows for easy extensibility and addition of new features and models.
Forward citations
Cited by 16 Pith papers
-
The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora
Deleting structural announcements makes following prose harder for LLMs to predict, swapping notation does nothing, and the paper proposes a pure-frame format that strips announcements into sidecars.
-
The Documentation and Traceability Burden of the Indian EV Transition
The paper systematises India's EV compliance-document lifecycle into a two-layer evidence model, a six-stage lifecycle with four failure loci, an exergy-destruction analytic lens, and a six-problem research agenda.
-
Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark
A multi-domain corpus of 5,264 scientific GitHub repos plus two IR benchmarks (219 expert queries; 117,950 snippets / 119,720 queries) shows large domain- and documentation-driven gaps in scientific code search.
-
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers
A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.
-
Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025
Analysis of EMNLP 2025 Responsible NLP Checklists shows 44.9% of NO justifications are brief or empty and 15.5% of main-track checklists contain parent-child logical contradictions.
-
Scout: Scalable Document Extraction via Data Similarity
Scout generates and refines LLM-written span-locating rules to match full-document agent accuracy on document extraction at orders-of-magnitude lower cost.
-
Measuring research data reuse in scholarly publications using generative artificial intelligence: Open Science Indicator development and preliminary results
An LLM-powered indicator measures research data reuse at 43% in publications, exceeding rates from established bibliometric methods and implying that benefits of data sharing are currently underestimated.
-
Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents
A new benchmark with two new datasets and end-to-end metrics shows that table extraction from PDFs is still unreliable across heterogeneous layouts.
-
SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
SARA combines natural-language snippets with semantic compression vectors in RAG to improve answer relevance, correctness, and similarity on 9 datasets across 5 LLMs.
-
The Hidden Threat in Plain Text: Attacking RAG Data Loaders
Invisible characters and formatting tricks in ingested documents survive popular RAG data loaders and can manipulate end-to-end RAG outputs.
-
Improving Access to Historical Archives with Real-time RAG-based Systems
LLM post-OCR cleanup plus dense retrieval and cross-encoder reranking raises NDCG@10 by 31.9% over BM25 on 500k Swiss newspaper segments while cutting CER/WER substantially.
-
Making Sense of Data in the Wild: Data Analysis Automation at Scale
A multi-agent LLM system with retrieval-augmented generation automatically curates datasets from Zenodo and Hugging Face, yielding small retrieval gains and a confounded synthetic-data improvement.
-
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
Typhoon 2 improves Thai LLM performance through continual pre-training on curated Thai data and post-training, releasing text, vision, audio, and safety models.
-
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
Across three invoice datasets, multimodal LLMs extract fields more accurately from raw images than from markdown converted by a parsing tool, with Gemini 2.5 Pro leading.
-
Knowledge Graph Fusion with Large Language Models for Accurate, Explainable Manufacturing Process Planning
A retrieval-augmented framework that auto-constructs a CNC machining knowledge graph with GPT-4o and uses graph retrieval to make small LLMs answer process-planning questions accurately.
-
Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery
Scout applies off-the-shelf LLMs and vision models to triage digital evidence, but only anecdotal examples are shown and accuracy is withheld.
Discussion (0). Continue with ORCID to comment.