Pith. sign in

cs.DL

Digital Libraries

Covers all aspects of the digital library design and document and text creation. Note that there will be some overlap with Information Retrieval (which is a separate subject area). Roughly includes material in ACM Subject Classes H.3.5, H.3.6, H.3.7, I.7.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Co-evolution loop prevents drift between ontology and data

X-DigCheck provides a domain-independent environment for continuous profile–graph alignment across regeneration cycles.

· “X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna”

open re-runnable review →
Figure from the paper

Human outlines and reviews lift automated survey quality above baselines

Multi‑agent system retrieves from multiple databases, re‑ranks by impact, and uses real peer‑review comments to revise drafts, achieving…

· “SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation”

open re-runnable review →
Figure from the paper

MultiGhostBench: no single method wins multilingual LLM attribution

New 928-book benchmark across six languages tests domain, author, and language transfer; Transformer detectors cross languages but with…

· “MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts”

open re-runnable review →
Figure from the paper

NPMI drop metric recovers patent gaps with high specificity

BLANC compares cluster associations before and after keyword filtering, flagging 'established globally, unexplored locally' at ~30%…

· “BLANC: Discovering Patent White Space via Changes in Normalized Pointwise Mutual Information Between Multi-View Clusters”

open re-runnable review →
Figure from the paper

Full OA aligns closer to patents than hybrid OA

Patents cite hybrid models more often, but gold and diamond OA show stronger semantic links, especially inside patent bodies.

· “Discoverability matters: Open access models and the translation of science into patents”

open re-runnable review →

A 1,951-figure benchmark maps how AI should read scientific images

Four expert-scored tasks move from spotting axes to judging evidence, with Bloom-informed questions scoring each level of understanding.

· “A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images”

open re-runnable review →
Figure from the paper

Congress's press releases doubled their em-dashes in 2025

A preregistered scan of 146,239 releases finds the punctuation shift came late and broad, with caveats.

· “The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025”

open re-runnable review →
Figure from the paper

One small set of climate-health pairings dominates 22,695 studies

Extreme heat and flood-hurricane mental health pairs recur far beyond chance, even after adjusting for term popularity.

· “Mapping the Climate-Health Evidence Base (2007-2023): A Bibliometric, Statistical, and NLP Multi-Label Text Analysis of 22,695 Records”

open re-runnable review →
Figure from the paper

BIP! Ranker computes eight citation metrics on billion-edge graphs

One distributed library spans influence, popularity, momentum, and field-normalised scores, running in minutes on 2.1 billion citations.

· “BIP! Ranker: A Software Library for Citation-Based Impact Indicators on Large-Scale Graphs”

open re-runnable review →
Figure from the paper

16 schemas turn lab methods into comparable machine data

Expert-finished field lists for biology, materials, imaging, physics, and psychology—ready for extraction and knowledge graphs.

· “SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions”

open re-runnable review →
Figure from the paper

18 centuries of citation style held steady in Chinese histories

A span-grounded AI extractor found 5,766 Analects reuses across 24 histories—wording drifted while practice held.

· “Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories”

open re-runnable review →
Figure from the paper

ChatGPT ranks articles as reliably as individual reviewers

In a 200-article study, ChatGPT's averaged rankings matched or beat single human experts—but detailed PDF critiques did not improve scores.

· “Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?”

open re-runnable review →

Same field lands in different buckets on every platform

Four bibliometric platforms classify by different mixes of experts, AI, citations, and granularity—know the scheme before you search.

· “Comparative Analysis of Classification Schemes on Major Bibliometric Platforms: A study of Web of Science, Scopus, the Lens, and Dimensions”

open re-runnable review →

Tiny unsupervised training beats labeled baselines on ancient text reuse

Four to eight thousand raw Latin or Greek sentences turn token models into strong reuse detectors and retrievers.

· “From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages”

open re-runnable review →
Figure from the paper

Only 4 of 984 papers report on government software engineering

Mainstream venues published almost nothing on how public agencies build software — a gap for evidence-based practice.

· “A Preliminary Search for Evidence on Government Software Engineering Practices: Results from Three Rapid Reviews”

open re-runnable review →
Figure from the paper

browse all of cs.DL → full archive · search · sub-categories