Local privacy mechanisms preserve rate-double-robustness, enabling unbiased and semiparametrically efficient inference on target parameters indexed linearly by infinite-dimensional and nonlinearly by low-dimensional components from noisy private data.
super hub
month = oct, year =
24 Pith papers cite this work, alongside 28,973 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
A new grid of disk models with grain-surface CO chemistry plus an ML inference tool produces gas mass estimates from ALMA observations that match independent dynamical and HD values without requiring extreme elemental depletion.
ScoreStop introduces a functional score test for early stopping in gradient boosting, testing the null that the current predictor minimizes population risk with a scale-invariant statistic of known asymptotic distribution.
SynQL synthesizes diverse, execution-ready SQL workloads by deterministically traversing foreign-key graphs to populate ASTs, yielding high topological entropy and cost-model training data with R² ≥ 0.79 on held-out sets.
MIBoost extends gradient boosting to multiple imputation by defining a single loss function that produces one set of selected variables across all imputed datasets.
Cosmological MBHBs coalesce in ~1 Gyr with high eccentricities; scaling relations from 30 Griffin re-simulations link dynamical-friction, hardening and total times to galaxy and orbital properties.
Derives the stationary distribution and asymptotic scaling O(ε^{-2}) for ensemble size in a Markov chain model of triplet-based plateau tuning for random forests.
Benchmark across 78 endpoint-split entries finds classical ML winning 47.4% of best performances over pretrained models, GNNs, and LLMs, with performance depending on model-task-split fit rather than scale.
ParamBoost improves GAMs by fitting piecewise cubic polynomials via gradient boosting and supports constraints for continuity, monotonicity, convexity, and feature interactions.
Statsformer adaptively integrates LLM semantic priors into a library of predictors via out-of-fold validation, delivering an oracle-style guarantee that the final predictor performs no worse than the best convex combination of its candidates up to statistical error.
On a 10-bearing PHME subset, a residual-calibrated fusion model hits ~0.15 normalized MAE and 0.90 average 90% coverage under leave-regime-out splits, while conditional diagnostics expose 0.666 coverage and raw-channel-loss collapse.
Proves that rescaled deviations of kernel gradient flow and infinitesimal gradient boosting from their deterministic ODE limits converge to a Gaussian process via a general stochastic perturbation analysis of ODEs in Banach spaces.
LHC vector boson fusion searches enhanced by machine learning can probe substantial regions of cosmologically viable parameter space for freeze-in dark matter mediated by a spin-2 particle.
SLV aggregates expected profit from current inventory through its full selling lifecycle to measure long-term opportunity costs within short A/B test windows.
Graph neural network achieves AUC of 0.883 for up versus anti-up quark jet charge discrimination in controlled QCD simulations.
A systematic literature survey that classifies data-driven KPI prediction methods for 6G networks across KPI type, data source, protocol stack layer, horizon, model family, and objective.
CatBoost achieved the highest average R-squared value of about 0.946 in a multi-task regression task for pectin process parameters, with raw material type identified as the most influential input feature.
VIGILant applies tree-based models and a ResNet CNN to classify Virgo O3b glitches with 98% accuracy and has been deployed for daily use with an interactive dashboard.
The paper introduces ClinQueryAgent, a conversational agent that converts natural language queries into database queries for population health management while keeping patient data secure, and reports its use by 128 staff across 15 NHS practices covering 148,319 patients.
Machine learning classifiers using fifteen cluster-level descriptors from time and ADC distributions effectively separate signal from background hits in prototype RPC detectors.
A qualitative-to-quantitative scoring framework is proposed to evaluate how well model-agnostic XAI methods support EU AI Act explainability requirements.
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
STRIKE improves credit default prediction AUC-ROC by training independent models on feature groups and aggregating their outputs via a meta-learner, outperforming tree baselines and conventional stacking on three real datasets.
Comparative study of machine learning models for multiclass classification of cryopathy syndromes from lab data, with best results from a soft-voting ensemble of Random Forest and Gradient Boosted Trees.
citing papers explorer
-
Private Rate-Double-Robust Inference
Local privacy mechanisms preserve rate-double-robustness, enabling unbiased and semiparametrically efficient inference on target parameters indexed linearly by infinite-dimensional and nonlinearly by low-dimensional components from noisy private data.
-
DiskMINT-GARDEN: Self-consistent Models to Estimate Disk Masses
A new grid of disk models with grain-surface CO chemistry plus an ML inference tool produces gas mass estimates from ALMA observations that match independent dynamical and HD values without requiring extreme elemental depletion.
-
ScoreStop: Gradient-based early stopping using functional score tests
ScoreStop introduces a functional score test for early stopping in gradient boosting, testing the null that the current predictor minimizes population risk with a scale-invariant statistic of known asymptotic distribution.
-
SynQL: A Controllable and Scalable Rule-Based Framework for SQL Workload Synthesis for Performance Benchmarking
SynQL synthesizes diverse, execution-ready SQL workloads by deterministically traversing foreign-key graphs to populate ASTs, yielding high topological entropy and cost-model training data with R² ≥ 0.79 on held-out sets.
-
MIBoost: A gradient boosting algorithm for variable selection after multiple imputation
MIBoost extends gradient boosting to multiple imputation by defining a single loss function that produces one set of selected variables across all imputed datasets.
-
Scaling Relations for Binary Black Hole Merger Times from Cosmological Initial Conditions
Cosmological MBHBs coalesce in ~1 Gyr with high eccentricities; scaling relations from 30 Griffin re-simulations link dynamical-friction, hardening and total times to galaxy and orbital properties.
-
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
Derives the stationary distribution and asymptotic scaling O(ε^{-2}) for ensemble size in a Markov chain model of triplet-based plateau tuning for random forests.
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
Benchmark across 78 endpoint-split entries finds classical ML winning 47.4% of best performances over pretrained models, GNNs, and LLMs, with performance depending on model-task-split fit rather than scale.
-
ParamBoost: Gradient Boosted Piecewise Cubic Polynomials
ParamBoost improves GAMs by fitting piecewise cubic polynomials via gradient boosting and supports constraints for continuity, monotonicity, convexity, and feature interactions.
-
Learning When to Trust LLM Priors: A Validated Framework for Semantic Prior Integration
Statsformer adaptively integrates LLM semantic priors into a library of predictors via out-of-fold validation, delivering an oracle-style guarantee that the final predictor performs no worse than the best convex combination of its candidates up to statistical error.
-
Empirical Calibration and Conditional-Reliability Diagnostics for Bearing RUL Prediction under Operating-Regime Shift
On a 10-bearing PHME subset, a residual-calibrated fusion model hits ~0.15 normalized MAE and 0.90 average 90% coverage under leave-regime-out splits, while conditional diagnostics expose 0.666 coverage and raw-channel-loss collapse.
-
A functional central limit theorem for kernel gradient flow and infinitesimal gradient boosting
Proves that rescaled deviations of kernel gradient flow and infinitesimal gradient boosting from their deterministic ODE limits converge to a Gaussian process via a general stochastic perturbation analysis of ODEs in Banach spaces.
-
Probing Freeze-In Dark Matter via a Spin-2 Portal at the LHC with Vector Boson Fusion and Machine Learning
LHC vector boson fusion searches enhanced by machine learning can probe substantial regions of cosmologically viable parameter space for freeze-in dark matter mediated by a spin-2 particle.
-
Measuring Opportunity Cost with Stock Lifetime Value
SLV aggregates expected profit from current inventory through its full selling lifecycle to measure long-term opportunity costs within short A/B test windows.
-
Application of Deep Learning to Jet Charge Discrimination
Graph neural network achieves AUC of 0.883 for up versus anti-up quark jet charge discrimination in controlled QCD simulations.
-
AI-Based KPI Prediction Methods in Future 6G Networks: A Survey
A systematic literature survey that classifies data-driven KPI prediction methods for 6G networks across KPI type, data source, protocol stack layer, horizon, model family, and objective.
-
A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process
CatBoost achieved the highest average R-squared value of about 0.946 in a multi-task regression task for pectin process parameters, with raw material type identified as the most influential input feature.
-
VIGILant: an automatic classification pipeline for glitches in the Virgo detector
VIGILant applies tree-based models and a ResNet CNN to classify Virgo O3b glitches with 98% accuracy and has been deployed for daily use with an interactive dashboard.
-
ClinQueryAgent: A Conversational Agent for Population Health Management
The paper introduces ClinQueryAgent, a conversational agent that converts natural language queries into database queries for population health management while keeping patient data secure, and reports its use by 128 staff across 15 NHS practices covering 148,319 patients.
-
Machine Learning-Based Cluster Classification to Suppress Background in a Prototype RPC Detector
Machine learning classifiers using fifteen cluster-level descriptors from time and ADC distributions effectively separate signal from background hits in prototype RPC detectors.
-
Assessing Model-Agnostic XAI Methods against EU AI Act Explainability Requirements
A qualitative-to-quantitative scoring framework is proposed to evaluate how well model-agnostic XAI methods support EU AI Act explainability requirements.
-
Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
-
STRIKE: Additive Feature-Group-Aware Stacking Framework for Credit Default Prediction
STRIKE improves credit default prediction AUC-ROC by training independent models on feature groups and aggregating their outputs via a meta-learner, outperforming tree baselines and conventional stacking on three real datasets.
-
Machine Learning Classification of Cryopathy Syndromes: A Comprehensive Comparative Study
Comparative study of machine learning models for multiclass classification of cryopathy syndromes from lab data, with best results from a soft-voting ensemble of Random Forest and Gradient Boosted Trees.