RCT dataset with sequence-preserving splits demonstrates that tactile-to-text models achieve only 25.1% Recall@1 on held-out materials, exposing generalization as the core challenge.
Leakage and the reproducibility crisis inmachine-learning-basedscience.Patterns2023;4:100804–100804
18 Pith papers cite this work, alongside 671 external citations. Polarity classification is still indexing.
representative citing papers
A framework proves that broad recalibrated leakage is undetectable from predictions alone without an external discrimination ceiling, while near-label leaks produce a detectable unit-purity signature yielding a prior-free test.
Simulations show standard CI methods underperform for classifier metrics in small and nested datasets, while Agresti-Coull, Wilson, Clopper-Pearson, and a new pseudo-count regularized bootstrap perform better, with specific adjustments needed for nested structures.
The authors develop an input-schema identifiability certificate for physics-informed surrogates that decomposes lumen velocity in tubular flow into mesh-measurable tangent direction, boundary-condition-dependent magnitude, and signed-orientation ambiguity using a Cosserat-rod reduction.
dille detects silent semantic faults in random forest ML pipelines with 91% precision via data-informed static analysis on Kaggle notebooks, finding 12-18% of scripts affected.
Neural point-forms are introduced as permutation-invariant neural layers that output learned form-comparison matrices for point clouds, with a claimed consistency proof under sampling and manifold assumptions and competitive results on synthetic and biological data.
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
Complex multimodal architectures do not reliably outperform unimodal baselines or a simple multimodal baseline under standardized evaluation.
A single sentence stating the discovery directly. ≤ 300 chars.
BEA-Dialogue+ expands the prior 85-hour Hungarian dialogue corpus to 200 hours via relaxed splits and demonstrates that SOT fine-tuning improves Whisper and FastConformer performance on word and character error metrics.
Object presence alone fails to predict minhwa genres while image-text fusion succeeds, revealing faithful but insufficient object grounding where genre hinges on symbol arrangement rather than inventory.
SurvBench supplies a configurable, open-source preprocessing pipeline that standardizes multi-modal EHR data from four critical-care databases for single-risk and competing-risk survival analysis.
A pathway-constrained autoencoder extended to multi-omics integration improves breast cancer stratification and provides interpretable pathway activity scores.
Subject-disjoint evaluation on C-NMC 2019 reduces AUROC by about 0.04 versus random splits, with EfficientNet-B1 reaching 0.913 AUROC, 0.87 sensitivity, and 0.80 specificity under honest conditions with calibration assessment.
LSTM networks predict HRRR forecast errors with average improvements of 48% for precipitation, 25% for temperature, and 15% for wind using mesonet ground truth.
fastml is an R package that enforces leakage-free preprocessing through guarded resampling and provides a unified interface for safer automated ML including survival analysis.
A survey of generative crystal modeling, multimodal learning, and closed-loop inverse design pipelines for crystalline solids, including failure modes and evaluation practices.
citing papers explorer
-
RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization
RCT dataset with sequence-preserving splits demonstrates that tactile-to-text models achieve only 25.1% Recall@1 on held-out materials, exposing generalization as the core challenge.
-
A prior-free blind detection of information leakage from model predictions
A framework proves that broad recalibrated leakage is undetectable from predictions alone without an external discrimination ceiling, while near-label leaks produce a detectable unit-purity signature yielding a prior-free test.
-
Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
Simulations show standard CI methods underperform for classifier metrics in small and nested datasets, while Agresti-Coull, Wilson, Clopper-Pearson, and a new pseudo-count regularized bootstrap perform better, with specific adjustments needed for nested structures.
-
Input-schema identifiability limits in physics-informed surrogates for mechanics-governed flow
The authors develop an input-schema identifiability certificate for physics-informed surrogates that decomposes lumen velocity in tubular flow into mesh-measurable tangent direction, boundary-condition-dependent magnitude, and signed-orientation ambiguity using a Cosserat-rod reduction.
-
Are We Lost in the Woods? Detecting Silent Semantic Faults for Random Forest Classifiers with Data-informed Static Analysis
dille detects silent semantic faults in random forest ML pipelines with 91% precision via data-informed static analysis on Kaggle notebooks, finding 12-18% of scripts affected.
-
Neural Point-Forms
Neural point-forms are introduced as permutation-invariant neural layers that output learned form-comparison matrices for point clouds, with a claimed consistency proof under sampling and manifold assumptions and competitive results on synthetic and biological data.
-
The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
-
Fusion or Confusion? Multimodal Complexity Is Not All You Need
Complex multimodal architectures do not reliably outperform unimodal baselines or a simple multimodal baseline under standardized evaluation.
-
Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEG
A single sentence stating the discovery directly. ≤ 300 chars.
-
Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus
BEA-Dialogue+ expands the prior 85-hour Hungarian dialogue corpus to 200 hours via relaxed splits and demonstrates that SOT fine-tuning improves Whisper and FastConformer performance on word and character error metrics.
-
MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting
Object presence alone fails to predict minhwa genres while image-text fusion succeeds, revealing faithful but insufficient object grounding where genre hinges on symbol arrangement rather than inventory.
-
SurvBench: A Standardised Preprocessing Pipeline for Multi-Modal Electronic Health Record Survival Analysis
SurvBench supplies a configurable, open-source preprocessing pipeline that standardizes multi-modal EHR data from four critical-care databases for single-risk and competing-risk survival analysis.
-
Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer
A pathway-constrained autoencoder extended to multi-omics integration improves breast cancer stratification and provides interpretable pathway activity scores.
-
A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection
Subject-disjoint evaluation on C-NMC 2019 reduces AUROC by about 0.04 versus random splits, with EfficientNet-B1 reaching 0.913 AUROC, 0.87 sensitivity, and 0.80 specificity under honest conditions with calibration assessment.
-
Predicting Forecast Error for the HRRR Using LSTM Neural Networks: A Comparative Study Using New York and Oklahoma State Mesonets
LSTM networks predict HRRR forecast errors with average improvements of 48% for precipitation, 25% for temperature, and 15% for wind using mesonet ground truth.
-
fastml: Guarded Resampling Workflows for Safer Automated Machine Learning in R
fastml is an R package that enforces leakage-free preprocessing through guarded resampling and provides a unified interface for safer automated ML including survival analysis.
-
Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design
A survey of generative crystal modeling, multimodal learning, and closed-loop inverse design pipelines for crystalline solids, including failure modes and evaluation practices.
- Towards a more realistic evaluation of machine learning models for bearing fault diagnosis