Pith. sign in

REVIEW 3 cited by

GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.05354 v1 pith:W6ANLQHO submitted 2023-07-11 cs.CL

classification cs.CL
keywords modelsancientlanguageprocessingintelligenttaskstextsautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent information processing of ancient texts. These models have been trained on an extensive dataset that encompasses both simplified and traditional Chinese characters, allowing them to effectively handle various natural language processing tasks related to ancient books, including but not limited to automatic sentence segmentation, punctuation, word segmentation, part-of-speech tagging, entity recognition, and automatic translation. Notably, these models have exhibited exceptional performance across a range of validation tasks using publicly available datasets. Our research findings highlight the efficacy of employing self-supervised methods to further train the models using classical text corpora, thus enhancing their capability to tackle downstream tasks. Moreover, it is worth emphasizing that the choice of font, the scale of the corpus, and the initial model selection all exert significant influence over the ultimate experimental outcomes. To cater to the diverse text processing preferences of researchers in digital humanities and linguistics, we have developed three distinct categories comprising a total of nine model variations. We believe that by sharing these foundational language models specialized in the domain of ancient texts, we can facilitate the intelligent processing and scholarly exploration of ancient literary works and, consequently, contribute to the global dissemination of China's rich and esteemed traditional culture in this new era.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AutoHDR restores full pages of damaged historical documents by combining OCR damage detection, LLM-based text prediction, and patch-autoregressive image diffusion, improving OCR accuracy on severely damaged documents ...

  2. RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture

    cs.MM 2025-06 conditional novelty 4.0 of 10

    The authors built a voice-interactive digital human system for ancient Yellow River culture with a curated 20,000-segment knowledge base and showed that RAG improves answer quality.

  3. ICH-Qwen: A Large Language Model Towards Chinese Intangible Cultural Heritage

    cs.CL 2025-05 reject novelty 4.0 of 10

    They fine-tuned Qwen2.5-7B on Chinese intangible cultural heritage texts to build ICH-Qwen, and report n-gram metric wins over general LLMs on 100-sample ICH QA tasks.

Pith tools