Maistros 8B is a new state-of-the-art open-weights Greek LLM built via knowledge distillation from large reasoning models on the CulturaQA dataset.
Leakage and the reproducibility crisis in machine- learning-based science.Patterns, 4(9), 2023
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
baseline 1polarities
baseline 1representative citing papers
A patient-disjoint benchmark for chest CT segmentation demonstrates that slice-mixed evaluation inflates foreground Dice by 0.46 absolute (69% relative) compared to proper case-disjoint splits.
xRFM merges kernel-based feature learning with tree structures for scalable, interpretable tabular modeling and reports top performance on 100 regression and competitive results on 200 classification datasets versus 31 baselines including GBDTs and TabPFNv2.
Longitudinal study of 56,800 AI papers finds sixfold increase in code+data sharing from 2014-2024 with inferred reproducibility rising from 28% to 64%.
Introduces a paired one-switch benchmark that quantifies protocol-induced inflation from decision-time leakage in financial ML backtests on equity panels from 2016-2024.
citing papers explorer
-
Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models
Maistros 8B is a new state-of-the-art open-weights Greek LLM built via knowledge distillation from large reasoning models on the CulturaQA dataset.
-
CTSCAN: Evaluation Leakage in Chest CT Segmentation and a Reproducible Patient-Disjoint Benchmark
A patient-disjoint benchmark for chest CT segmentation demonstrates that slice-mixed evaluation inflates foreground Dice by 0.46 absolute (69% relative) compared to proper case-disjoint splits.
-
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
xRFM merges kernel-based feature learning with tree structures for scalable, interpretable tabular modeling and reports top performance on 100 regression and competitive results on 200 classification datasets versus 31 baselines including GBDTs and TabPFNv2.
-
The Shift Toward Open and Reproducible AI Research
Longitudinal study of 56,800 AI papers finds sixfold increase in code+data sharing from 2014-2024 with inferred reproducibility rising from 28% to 64%.
-
When Alpha Disappears: A One-Switch Benchmark for Decision-Time Leakage in Financial Backtests
Introduces a paired one-switch benchmark that quantifies protocol-induced inflation from decision-time leakage in financial ML backtests on equity panels from 2016-2024.