Introduces REST-ASMR multimodal dataset of PPG, stimuli, and continuous annotations for ASMR research, validated with 97% responder rate, significant agreement, PPG deceleration, and BiLSTM achieving 75.51% frame-level accuracy under strict subject-video independent 4-fold CV.
Mixed citations
Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , pages =
Mixed citation behavior. Most common role is method (54%).
citation-role summary
citation-polarity summary
co-cited works
representative citing papers
FLOATBench is a tabular benchmark dataset with 582,120 fatigue labels from 19,404 OpenFAST simulations of three 22 MW FOWT towers, featuring alpha-shape regime partitioning and three evaluation protocols for surrogate models.
A new corpus of 108 mixed string-numeric tables shows that advanced tabular learners with basic string embeddings perform well on most real-world data, while large LLM encoders help on free-text heavy tables.
RAVEN proposes a regime-aware MoE architecture with cumulative importance thresholding and correlation-aware weighting to adaptively select temporal context for non-stationary financial forecasting.
Four self-stigma personas identified via LPA on 1,174 Reddit users; persona-conditioned LLMs achieve targeted shifts but experts prefer generic empathy baselines.
Domain-specific macOS features enable an ML detector to reach 98.5% accuracy on 41k samples and 99.5% on 9k fresh samples, beating prior methods by 16-50%.
Strong absolute accuracy on mixture properties often masks poor recovery of non-ideal behavior, with large drops under strict molecule splits, making transfer to unseen molecules the central challenge.
TabOrder learns unsupervised causal variable orderings and enforces them with order-constrained attention for tabular prediction and imputation under distribution shifts.
TIDAL recovers temporal phase signals from LLM-derived semantics of provisioning metadata to enable complementary CVD placement, reducing overload frequency by 79.1% on production traces.
TabPFN-MT is a multitask in-context learner for tabular data that sets a new state-of-the-art on deep multitask learning for datasets under 1000 samples while reducing inference cost from O(T) to O(1) passes.
An agentic AI workflow evolves an adaptive XGBoost quantile regression ensemble that reduces watershed-averaged forecast error by up to 29% versus California's operational forecasts for April-July runoff at 1-6 month leads across 23 Sierra Nevada sites.
A compositional algebraic decision diagram algorithm quantifies sensitivity in decision tree ensembles with certified error and confidence bounds, outperforming model counters on benchmarks.
GHGbench supplies a harmonized dataset and multi-task benchmark for company and building carbon emission prediction, with baselines showing large OOD gaps and benefits from multimodal embeddings.
V4FinBench is a new million-record benchmark where imbalance-aware finetuned TabPFN matches or beats gradient boosting on long-horizon bankruptcy prediction while Llama-3-8B lags, with evidence of transferable patterns to US data.
MulTaBench is a new collection of 40 image-tabular and text-tabular datasets designed to test target-aware representation tuning in multimodal tabular models.
TFM-Retouche is an architecture-agnostic input-space residual adapter that improves tabular foundation model accuracy on 51 datasets by learning input corrections through the frozen backbone, with an identity guard to fall back to the original model.
Introduces Calibrated Size Ratio (CSR) and confidence-weighted metrics to better detect overconfidence risk and calibration issues beyond the limitations of ECE.
Probabilistic PCA latent-space model with Bayesian inference reconstructs TNO near-IR spectra from photometry, achieving 95% credible-interval coverage and supporting taxonomy plus survey optimization.
LP2B encoding converts Lund plane jet representations into Bloch sphere qubit states, enabling a QTTN that matches classical LundNet performance on polarization tagging and W/top tagging with three orders of magnitude fewer parameters and improved low-data regime results.
First observation of electroweak photon plus two jets production yields a cross section of 202 fb consistent with the standard model prediction of 177 fb at greater than 5 sigma significance.
No excess observed; first LHC search sets 95% CL upper limits on H to AA to 4e branching fraction down to 10^{-5} for 10-100 MeV masses and short lifetimes.
FeDa4Fair is a new library and benchmark for creating federated datasets with heterogeneous client-level biases to standardize evaluation of fairness methods in federated learning.
Compares foundation models for probabilistic low-voltage load forecasting on 200 real feeders and introduces a grid-planning metric that scores peak prediction by its effect on asset cost-risk decisions.
Derives the stationary distribution and asymptotic scaling O(ε^{-2}) for ensemble size in a Markov chain model of triplet-based plateau tuning for random forests.
citing papers explorer
-
Lund Plane to Bloch (LP2B) Encoding for Object and Polarization Tagging with Quantum Jet Substructure
LP2B encoding converts Lund plane jet representations into Bloch sphere qubit states, enabling a QTTN that matches classical LundNet performance on polarization tagging and W/top tagging with three orders of magnitude fewer parameters and improved low-data regime results.
-
Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE
Static malware classifiers learn packing artifacts and dataset composition biases rather than malicious semantics, as diagnosed by TRUSTEE interpretability across controlled dataset variations.
-
Predicting co-segregation in alloys with solute-solute interactions
An extended dual-solute segregation model with machine-learned pairwise energies predicts co-segregation bounds in Mg-based ternary alloys, validated by hybrid MD/MC and literature experiments.
-
Highly boosted dielectron identification in proton-proton collisions at $\sqrt{s}$ = 13 TeV
CMS develops two multivariate models for identifying boosted dielectrons with γ_L > 20, reporting 80% efficiency for two-track cases from J/ψ data and 60% for single-track from Z conversions, plus an energy correction.
-
Inferring identified hadron production in $pp$ collisions with physics-informed machine learning at the LHC
A physics-informed neural network infers pT spectra of pi, K, p, Lambda, and Ks in unmeasured rapidity regions from PYTHIA8 pp collisions at 13.6 TeV, achieving 1.5-5.83% yield uncertainties while reproducing yield ratios and freeze-out parameters.
-
Is the `Known' Enough? An Integrated Machine Learning Framework for Eclipsing Binary Classification and Parameter Estimation Based on Well-Characterized Systems
An ensemble ML framework achieves 90.7% morphology classification accuracy and R² values of 0.77–0.92 for key parameters on held-out test data, with external validation against OGLE and Kepler catalogs.
-
Comparing Analytical Approaches for Bike Station Expansion: A Location-Allocation Study in Trondheim, Norway
Three location-allocation models applied to identical spatial data in Trondheim produce distinct station configurations, with consensus identifying 12 priority expansion sites that differ from the existing network.
-
SemiCharmTag: a tool for Semileptonic Charm tagging
SemiCharmTag rejects charm-semileptonic muon background under LHCb dimuon Drell–Yan by secondary-vertex tagging with a hadron track, improving S/B by ~4 at 81% signal efficiency in simulation.
-
Earth Embeddings Reveal Diverse Urban Signals from Space
Earth embeddings from satellite images predict neighborhood-level urban indicators with higher accuracy for built-environment outcomes than for behavior-driven ones, showing city-specific variation but year-to-year stability.
-
Probing Freeze-In Dark Matter via a Spin-2 Portal at the LHC with Vector Boson Fusion and Machine Learning
LHC vector boson fusion searches enhanced by machine learning can probe substantial regions of cosmologically viable parameter space for freeze-in dark matter mediated by a spin-2 particle.
-
Validating a Deep Learning Algorithm to Identify Patients with Glaucoma using Systemic Electronic Health Records
A fine-tuned deep learning model using systemic EHR data achieved AUROC 0.883 and PPV 0.657 for identifying glaucoma in a held-out Stanford cohort of over 20,000 patients.
-
Accelerating the Design of Resorbable Magnesium Alloys: A Machine Learning Approach to Property Prediction
CatBoost and other ensemble ML models achieve R² scores of 0.95, 0.916, and 0.903 on yield strength, ultimate tensile strength, and elongation for resorbable Mg alloys, with SHAP analysis highlighting processing conditions and Zn/Mn/Gd content as key drivers.
-
An Explainable Unsupervised-to-Supervised Machine Learning Framework for Dietary Pattern Discovery Using UK National Dietary Survey Data
An unsupervised-to-supervised ML pipeline on UK NDNS data discovers four dietary patterns, reproduces them with macro-F1 0.963 using a surrogate classifier, and interprets them via SHAP for potential clinical use.