Pith. sign in

Automated Data Readiness for Scientific AI

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenance-instrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.

fields

cs.DL 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

SetGo: Metadata Readiness for Scientific AI Datasets

cs.DL · 2026-07-10 · conditional · novelty 6.0

SetGo assesses and repairs metadata readiness of scientific AI datasets before publication, raising FAIR scores on four benchmark corpora from ~52–57% to ~81–91%.

citing papers explorer

Showing 1 of 1 citing paper.

  • SetGo: Metadata Readiness for Scientific AI Datasets cs.DL · 2026-07-10 · conditional · none · ref 18 · internal anchor

    SetGo assesses and repairs metadata readiness of scientific AI datasets before publication, raising FAIR scores on four benchmark corpora from ~52–57% to ~81–91%.