Pith. sign in

cs.DB

Databases

Covers database management, datamining, and data processing. Roughly includes material in ACM Subject Classes E.2, E.5, H.0, H.2, and J.1.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Object causal model automates ANN bottleneck diagnosis and redesign

Automatically decomposes ANN workflows into eight replaceable phases, then selects and composes optimizations to improve recall and…

· “SmartANN: Object Causal Modeling Boosts Approximate Nearest Neighbor Diagnosis and Auto-Design”

open re-runnable review →
Figure from the paper

16-module reference model for governed analytical data pipelines

Proposal unites eight pipeline steps with eight cross-cutting capabilities to embed quality and traceability from design.

· “UnespDataLens-RM: A Reference Model for Analytical Data Engineering with Governance, Quality, Provenance, and Reproducibility”

open re-runnable review →

Co-evolution loop prevents drift between ontology and data

X-DigCheck provides a domain-independent environment for continuous profile–graph alignment across regeneration cycles.

· “X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna”

open re-runnable review →
Figure from the paper

Model improvements obsolete data agent scaffolds

Experiments show general coding agents outperform specialized pipelines; the lasting challenge is persistent knowledge about the data…

· “What Happens When the Model Eats the Stack? Rethinking the Research Agenda for Data Agents to Withstand the Bitter Lesson”

open re-runnable review →
Figure from the paper

Thermal drone screening achieves 92.5% precision on inert ordnance

But biased validation set means real-world false alarms could be much higher; tool is for prioritization, not clearance.

· “UAV Thermal Imagery for Inert Ordnance Screening: Multi Campaign Dataset Development,Object Detection, and Practical Recommendations”

open re-runnable review →
Figure from the paper

Agentic data cracking halves LLM agent token cost

By adaptively structuring data during reasoning, agents avoid reopening documents for related queries, cutting costs 53% without accuracy…

· “Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data”

open re-runnable review →

MAGE is a database that records what multi-agent AI systems do as a network of…

A heterogeneous temporal hypergraph memory engine for multi-agent systems outperforms eleven memory baselines on 12 QA benchmarks by…

· “Diachronic Hypergraphs for Orchestrated Multi-Agent Multimodal Memory Curation”

open re-runnable review →
Figure from the paper

Static profiles pick SPARQL sources without a single ASK query

RENSA infers variable classes and namespaces from compact profiles, matching live-probing engines at zero runtime cost

· “RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation”

open re-runnable review →
Figure from the paper

18 resources converge on community curation best practices

A workshop report finds proactive outreach, AI-assisted validation, and shared infrastructure are key to sustaining shrinking databases

· “Engaging the scientific community in high-quality biocuration: a report on the International Society for Biocuration workshop, 'Maximizing community curation for the benefit of all'”

open re-runnable review →

Synthetic databases replay slow queries without touching production data

A hybrid solver keeps global statistics intact while forcing exact cardinalities, so offline diagnosis of slow queries becomes practical.

· “DBRepro: Automated Database Synthesis via a Hybrid Constraint-Solving Approach for Reproducing Slow Queries”

open re-runnable review →
Figure from the paper

Bind integrity to version histories to secure disaggregated key-value stores

ANCHOR treats caches as untrusted hints and uses version-bound authentication to prevent rollback and poisoning without blocking async I/O.

· “ANCHOR: A Vision for Secure Persistent Key-Value Stores in Disaggregated Data Centers”

open re-runnable review →
Figure from the paper

Home activity benchmark shows AI question-answering gaps

HOME-KGQA tests multimodal KGQA on daily household tasks, where LLM methods lag behind their encyclopedic results.

· “HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities”

open re-runnable review →
Figure from the paper

eBPF scheduler doubles throughput for time-sensitive DB tasks

UFS limits background tasks to idle CPUs and shares lock hints to halve tail latency versus standard Linux options in mixed workloads.

· “Unfair by design: eBPF-based scheduling of mixed database workloads”

open re-runnable review →
Figure from the paper

browse all of cs.DB → full archive · search · sub-categories