Pith. sign in

REVIEW 5 cited by

Docs2KG: Unified Knowledge Graph Construction from Heterogeneous Documents Assisted by Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02962 v1 pith:6QGL7VVT submitted 2024-06-05 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords docs2kgdataheterogeneousknowledgefilesinformationai4wadocument
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Even for a conservative estimate, 80% of enterprise data reside in unstructured files, stored in data lakes that accommodate heterogeneous formats. Classical search engines can no longer meet information seeking needs, especially when the task is to browse and explore for insight formulation. In other words, there are no obvious search keywords to use. Knowledge graphs, due to their natural visual appeals that reduce the human cognitive load, become the winning candidate for heterogeneous data integration and knowledge representation. In this paper, we introduce Docs2KG, a novel framework designed to extract multimodal information from diverse and heterogeneous unstructured documents, including emails, web pages, PDF files, and Excel files. Dynamically generates a unified knowledge graph that represents the extracted key information, Docs2KG enables efficient querying and exploration of document data lakes. Unlike existing approaches that focus on domain-specific data sources or pre-designed schemas, Docs2KG offers a flexible and extensible solution that can adapt to various document structures and content types. The proposed framework unifies data processing supporting a multitude of downstream tasks with improved domain interpretability. Docs2KG is publicly accessible at https://docs2kg.ai4wa.com, and a demonstration video is available at https://docs2kg.ai4wa.com/Video.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering

    cs.IR 2026-03 unverdicted novelty 7.0 of 10

    Docling with hierarchical splitting reaches 94.1% RAG accuracy on domain documents, beating naive PDF loading but trailing manual Markdown curation at 97.1%.

  2. FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Embedding FAIR Digital Objects as graph nodes yields a GraphRAG system that measurably improves accuracy, coverage and explainability on biomedical RNA-seq queries versus a non-FAIR baseline.

  3. Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    NKW is a retrieval-augmented system that combines text, graph, and narrative tools to assemble and audit evidence for long-form story-world question answering.

  4. Guideline2Graph: Profile-Aware Multimodal Parsing for Executable Clinical Decision Graphs

    cs.CV 2026-04 conditional novelty 6.0 of 10

    A decomposition-first pipeline with topology-aware chunking and interface-constrained merging converts full clinical guidelines into executable decision graphs, raising edge precision from 19.6% to 69.0% and triplet r...

  5. From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering

    cs.IR 2026-03 conditional novelty 5.0 of 10

    Docling with hierarchical splitting and image descriptions reaches 94.1% RAG QA accuracy on Portuguese admin PDFs, beating manual Markdown and other open-source converters.

Pith tools