Path patterns weight random forest trees for accuracy gains
Instance-adaptive signals from root-to-leaf paths yield +0.99 pp mean improvement on qualifying binary benchmarks without class trade-offs.
Machine Learning
Covers machine learning papers (supervised, unsupervised, semi-supervised learning, graphical models, reinforcement learning, bandits, high dimensional inference, etc.) with a statistical or theoretical grounding
sort pith recommended most recent
Instance-adaptive signals from root-to-leaf paths yield +0.99 pp mean improvement on qualifying binary benchmarks without class trade-offs.
Review shows how tail models support regression, classification, dimension reduction and anomaly detection in rare-event settings.
· “Extrapolation in Statistical Learning with Extreme Value Theory”
A non-Euclidean inner product derived from counterfactuals makes high-level concepts linear directions and connects interpretation directly
· “The Linear Representation Hypothesis and the Geometry of Large Language Models”
A gradient-sign method generates them quickly and the same examples can be used to improve clean-data accuracy during training.
Gradient descent on principal component prediction alone yields parameters where each layer runs one power iteration.
· “Looped Transformers with Layer Normalization Provably Learn the Power Method”
Joint phase-space model with smooth second moment follows discrete steps for eight optimizers where prior approximations break down.
SIREN freezes the shortlist, separates selection from evaluation, and uses item-level bootstrap to recover accurate procedure performance.
· “Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking”
A smooth approximation of discrete choices lets policies plan full sequences of costly queries before predicting, lowering total error on A
· “Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients”
Isolated tests show dropout and confounders erase causal advantages over simple correlations in single-cell data
Predicting full Kantorovich potentials from sliced OT allows quick reuse across many measure pairs without depending on their structure.
With few denoising steps, scoring-rule-trained diffusion beats standard models on images and robot tasks.
A kernel filter pulls any iterated integral from a path in parallel time that no longer grows with depth.
Learned reverse-process variances also cut sampling steps by 10x with almost no quality loss and scale smoothly with compute.
Exact rank and a new lower bound show every convex calibrated surrogate pays quadratic dimension.
· “Exact Rank and Convex Calibration Dimension Lower Bounds for the Multi-Label F1 Loss”
Above that many points, almost surely no centered ellipsoid passes through them all.
· “The sharp SAT/UNSAT phase transition in random ellipsoid fitting”
No learner can beat this rate, and a single algorithm reaches it without knowing the oracle error or the confidence level.
New minimax bounds tie regret to the product (B-1)W of boundary-state bits and batch count.
· “Information Routing across Batch Boundaries: Memory--Batch Tradeoffs in Lipschitz Bandits”
Monotonicity under data processing and additivity on independent products force every such functional to an integral over four strata
Vertices are realizable label sequences of length m; edges mark label disagreements on shared points, fixing whether dimension meets or tops
CSA maintains per-round selective risk bounds under predictable updates without pooling across deployments.
· “Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs”
Three probe conditions suffice for identification and replace ambient d sqrt(T) regret with intrinsic r sqrt(T) plus detection cost.
· “Catching a Moving Subspace: Low-Rank Bandits Beyond Stationarity”
Tests on gridworlds, chess and terminals separate solved perceptual errors from persistent multi-step and abstention failures.
· “HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models”
Proves minimax Θ(s/√N) bound and faster Θ(s/N) rate when used only to warm-start the solver.
· “Provably Data-driven Lagrangian Relaxation for Mixed Integer Linear Programming”
Gradient flow no longer selects the vanishing ridge solution outside the kernel regime, so a function-space energy defines geodesic ridge as
· “Canonical Regularisation of Wide Feature-Learning Neural Networks”
when initial loss is small and log-sum-exp functions remain linearly independent modulo affine functions
Extending delay-thresholding to LMO momentum handles heterogeneous workers and matches best known bounds in smooth cases.
· “Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method”
Upper-tail accumulation scale unifies conflicting laws for rescaling inverse temperature with context length n.
· “A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention”
Under Gaussian transport the velocity field equals a Nadaraya-Watson smoother whose bandwidth shrinks from global to nearest-neighbor.
The method outperforms benchmarks in Delhi and Mumbai while staying stable during spikes and seasons and adding uncertainty estimates.
Benchmark on two cohorts shows CFN edges linear CLR in colon while compartment labels remove spurious dependency instability.
New proof separates an m-dependent warm-up from a stochastic term that does not grow with the number of quantiles.
Artist-to-artist distances from full acoustic distributions recover expert review co-mentions, rising to 0.865 when critics agree.
Fine-tuning Whisper Small on 1,373 elicited recordings gives Baniwa its first ASR baseline.
· “Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study”
A density-matrix representation of each driver also reproduces the fundamental diagram and hysteresis on I-24 highway data.
The proven convergence powers fast projections that separate clusters and outliers in high-dimensional data.
· “Efficient Estimation of High Information Projections using Nearest Neighbours”
It matches fixed bases on clean data and degrades 3–9x less when labels get noisy.
· “Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality”
Treating data domains as mixture components lets design pick the most informative proxy-training points.
An exact decision rule, tested on two 7B model families, values a probe by the boundary-crossing mass it reveals.
Bayesian priors that borrow shrinkage across neighbors yield tighter calibrated intervals, even with missing data.
· “Graph-dependent shrinkage priors for Bayesian trend filtering”
Proof and 12-month-window evidence show 12,000 random features add no real market timing beyond a simple mean.
· “(Mis)Understanding Benign Overfitting in Equity Return Prediction”
A leave-one-out sampler needs O(DTC/ε) parallel steps, bypassing the dimension-scaling barrier of standard τ-leaping.
· “Provably adaptive sampling with uniform and remasking discrete diffusion models”
ConvergeFlow trains with MSE alone and reaches a generative perplexity of 33.17 on OpenWebText at dataset entropy.
· “ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings”
A black-box teacher defines each neighborhood; randomized refits flag which coefficients are stable.
On glucose data it flags hypoglycemia about 18 minutes earlier than waiting through the full hour.
· “Primal--Dual Alternating Neural Learning for Timely Classification with Performance Guarantees”
A single dissipativity condition replaces ellipticity and yields an almost-sure pseudo-trajectory property for optimization.
Spearman 0.81 on shape errors, 0.69 on scale errors for simulated Ariel spectra.
Separating skips from engaged watching beats point regression and the prior mixture on ranking and completion.
· “Hierarchical Exponential-Gaussian Mixtures for Watch-Time Distribution Prediction”
Keeping the full spread of predicted probabilities, not the average, yields calibrated abstention under 0.6% ECE.
· “Credal Large Language Models for Semantic Commitment under Uncertainty”
At the Bayes limit, one DDIM inversion step is gradient descent on a strongly convex potential with modulus e^{-h_t}, giving exact…
· “One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion”
One conditioned training sweeps a dark-photon model over two parameters and finds m_phi in [8,13] MeV.
Initial-regularization SGD achieves noiseless excess risk up to O(m^{-3+epsilon}) for source parameter r=1, improving on the prior r=1/2…
Exact identities show feature-side alignment can hide cancellation or filtering instead of internal harmony.
· “A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks”
A diffusion model's probability-flow ODE maps pre-change data to Gaussian latents, and a closed-form MMD with Shiryaev-Roberts recursion…
· “Change Detection in Probability Flow ODE: Online Testing in Diffusion Latent Spaces”
A generative framework using convex neural networks learns worst-case distributions for Sinkhorn-based robust hypothesis testing, enabling…
· “Generative Neural Networks for Sinkhorn Distributionally Robust Hypothesis Testing”
Two online algorithms learn a latent coefficient field from a single trajectory at roughly n^{-1/2} sup-norm error.
· “Q-Learning with Stable Infinite-Dimensional Linear Function Approximation”
Optimal sampling probabilities no longer shift with predictor units, so inactive features can't hijack selection in imbalanced data.
· “Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models”
New guarantee: value estimates stay accurate when either trajectories or horizon length grows, even with thousands of state features.
Two independent calibrations deploy one prediction set with probability $1-\rho$, at a quantified cost in data and set size.
A single-scale score field determines center, homogeneity, and angular weights of a junction—no score derivatives needed.
· “Recovering Weighted Tangent Geometry from a Single-Scale Score Field”
For the A4/LEARN APOE4 contrast, transparent balancing beats uncertainty sampling; fitted scores win for age slopes.
· “Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN”
It learns both the grouping of variables and the causal graph between groups from data alone.
· “Joint Causal Structure and Cluster Discovery Using Variational Inference”
Learning over context-length windows hits AUC 0.98 on AI-text detection, beating full-context baselines on both tasks.
· “Token-Level Likelihood-Array Regression for Membership Inference and AI-Generated Text Detection”
Chaotic systems: short trajectories diverge, but long-run statistics match the true attractor.
· “Symbolic Neural ODEs: Learning interpretable models from time-series data”
Beyond the graph: cover, clusters, and nerve each add predictive signal; new distance and stability tools.
Aligning every treatment pattern to one barycenter cuts cost from quadratic to linear.
· “Barycentric Fused Gromov-Wasserstein Balancing for Causal Inference under Multiple Treatments”
The paper proves a single variance-based sampling score works across bandits, tree search, and Q-learning.
Modeling pixel and color correlations at log-linear cost cuts FID and NLL at 10–50 sampling steps.
· “Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling”
A pretrained inference model adapts to fresh priors at test time, no simulator reruns or Gaussian shortcuts.
GeoQ wraps pretrained surrogates in quantile error estimates plus a flag for unsupported queries.
· “GeoQ: Geometry-Aware Conditional Quantile Error Estimation for Scientific Surrogate Models”
A stochastic least-squares method finds latent low-rank structure that linear SVD misses in sparse connectomes.
It learns a continuous-time hazard from messy, asynchronous measurements, using only information available just before each moment.
Modeling complex arrays in place, not vectorized, cuts covariance error and imputes missing brain regions.
The right effective size is set by exceedance correlation, and it shifts with every threshold level.
· “The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering”
The dependence channel alone ranks Storm Atiyah 59th; the record-rule fallback explains why 2020 has no alerts.
· “TRACE-C: Rank-Calibrated Relational Anomaly Detection for Multi-Stream Operational Telemetry”
Exact formula: audits through m cannot certify best-of-N reliability; the gap shrinks only with the m²/N scale.
· “The geometry of AI validation: Exact certification limits for iid best-of-N search”
An audit finds differenced loaders hide signal, so per-node SARIMA residuals are a better target for graph neural networks
· “A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Baselines”
For uniform and masking diffusion, the states-per-sample error rate is unavoidable and a counting estimator reaches it.
The uncertainty band is provably the worst case the model class allows.
· “Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series”
Autoencoder objective provably balances numerical variability, categorical entropy, and cross-type dependence.
· “Conditional-Independence-Regularized Distributional Autoencoders for Mixed-Type Data”
The reliability diagram gains a real null distribution, so a p-value replaces a bare number.
· “EDGE: a closed-form directed test for the calibration of probabilistic binary classifiers”