Path patterns weight random forest trees for accuracy gains
Instance-adaptive signals from root-to-leaf paths yield +0.99 pp mean improvement on qualifying binary benchmarks without class trade-offs.
Statistics
sort pith recommended most recent
Instance-adaptive signals from root-to-leaf paths yield +0.99 pp mean improvement on qualifying binary benchmarks without class trade-offs.
Western Australia data show short-term uptake rose then converged as ineligible peers caught up, revealing pure timing shift.
RetiSEM organises variables into biological blocks and forbids edges, lowering structural error versus unconstrained models on synthetic tes
· “RetiSEM: Generalising Causal Models for Fragmented Biomedical Data”
Minimum free energy randomization controls imbalance and prevents unobserved factors from dominating error in effect estimates
· “Minimum free energy randomized design to improve covariate balance”
A scaling factor projects tail autocovariances onto parametric models without full prewhitening or recoloring penalties.
· “Tail postcoloring in long-run variance estimation of time series”
Review shows how tail models support regression, classification, dimension reduction and anomaly detection in rare-event settings.
· “Extrapolation in Statistical Learning with Extreme Value Theory”
A non-Euclidean inner product derived from counterfactuals makes high-level concepts linear directions and connects interpretation directly
· “The Linear Representation Hypothesis and the Geometry of Large Language Models”
A gradient-sign method generates them quickly and the same examples can be used to improve clean-data accuracy during training.
Gradient descent on principal component prediction alone yields parameters where each layer runs one power iteration.
· “Looped Transformers with Layer Normalization Provably Learn the Power Method”
Two-gradient scheme matches prior convergence guarantees for targets that are merely log-smooth rather than strongly convex.
Unsupervised online detector tracks truncated eigenvalue sequences in multivariate series without strong distributional assumptions.
· “CHASM: Online Changepoint Detection in Temporal and Cross-Variable Dependence”
Joint phase-space model with smooth second moment follows discrete steps for eight optimizers where prior approximations break down.
SIREN freezes the shortlist, separates selection from evaluation, and uses item-level bootstrap to recover accurate procedure performance.
· “Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking”
A smooth approximation of discrete choices lets policies plan full sequences of costly queries before predicting, lowering total error on A
· “Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients”
Isolated tests show dropout and confounders erase causal advantages over simple correlations in single-cell data
Robust scatter matrices deliver stronger privacy at comparable utility when influential outliers are present.
· “Data anonymization in the presence of outliers via invariant coordinate selection”
Six-step framework maps estimands and borrowing strategies for trials that lack randomized controls, supporting decisions in rare diseases.
· “Externally Controlled Trials: A Review of Design and Borrowing Through a Causal Lens”
Decomposing interviews into symptom-specific tasks yields lower error than original raters and 0.877 agreement with experts.
· “ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms”
New high-dimensional theory for dependent observations turns pointwise sieve estimators into simultaneous inference tools.
· “Simultaneous Inference for Nonlinear Time Series, a Sieve M-regression Approach”
New methods generate random numbers from this distribution at constant speed regardless of shape parameters and support Bayesian models.
· “The Pearson IV distribution: Random variate generation and applications”
A time-varying penalty derived from prior location information slots into the fast PELT algorithm and improves accuracy without losing exact
· “Pi-Change: A Prior-Informed Multiple Change Point Detection Algorithm”
A running product of threshold bets keeps type I error capped at every stop while growing exponentially under alternatives.
· “Betting on Bets: Anytime-Valid Tests for Stochastic Dominance”
Simulations with varying cluster sizes and correlations confirm reliable performance for win ratio, win odds, net benefit and DOOR tests.
· “Statistical inference with win statistics in cluster-randomized trials with composite outcomes”
Predicting full Kantorovich potentials from sliced OT allows quick reuse across many measure pairs without depending on their structure.
Extending hyperevent models to three-way hypergraphs allows testing what drives teams, citations, and topics to combine while controlling 3-
· “Modeling Tripartite Hyperevents in Scientific Collaboration Networks”
PToTR constrains the coefficient tensor to a low-rank form so that estimated rates remain positive for tensor responses and tensor inputs.
· “Poisson-response Tensor-on-Tensor Regression and Applications”
Teams that bridge distinct knowledge communities produce more breakthroughs and escape the usual penalty of larger size.
· “Structural Diversity Drives Disruptive Scientific Innovation”
With few denoising steps, scoring-rule-trained diffusion beats standard models on images and robot tasks.
A kernel filter pulls any iterated integral from a path in parallel time that no longer grows with depth.
A three-step estimator outdoes direct high-dimensional quantile methods at 99.9th-percentile tails.
Learned reverse-process variances also cut sampling steps by 10x with almost no quality loss and scale smoothly with compute.
The exact threshold is set by the coordinates' fourth moment alone: heavier tails make fitting fail sooner.
Proof shows expected trace error within a 1+ε factor of the optimal rank-r fit in O(r/ε + r√log r) steps.
· “A new analysis of the randomly pivoted Cholesky algorithm”
Every Bingham parameter matrix gets at least this acceptance probability, so the sampler runs in worst-case polynomial time.
· “A Complexity Bound for the Kent-Ganeiber-Mardia Sampler for the Bingham Distribution”
Even with deployment data in hand, calibration noise sets the floor: relative precision scales as N times the sixth power of s.
· “Physical-Support Confidence Sets for Highly Coherent Dictionaries”
Even with thousands of units, only a few covariates can be balanced; new designs target smaller function classes instead.
· “The Limits of Experimental Design: Covariate Balance Beyond Low Dimension”
Any alarm rule with average run length at least b is a threshold crossing of an e-detector; strong rules also bound every adaptive horizon.
Cox's original exact partial likelihood is asymptotically optimal when tied failure times occur at finitely many instants.
· “Semiparametric efficient estimation of the Cox regression coefficient when there can be ties”
Exact rank and a new lower bound show every convex calibrated surrogate pays quadratic dimension.
· “Exact Rank and Convex Calibration Dimension Lower Bounds for the Multi-Label F1 Loss”
Above that many points, almost surely no centered ellipsoid passes through them all.
· “The sharp SAT/UNSAT phase transition in random ellipsoid fitting”
Cube orientation plus suffix averaging settles agnostic PAC sample complexity up to universal constants.
No learner can beat this rate, and a single algorithm reaches it without knowing the oracle error or the confidence level.
New minimax bounds tie regret to the product (B-1)W of boundary-state bits and batch count.
· “Information Routing across Batch Boundaries: Memory--Batch Tradeoffs in Lipschitz Bandits”
One coherence condition turns per-intersection e-processes into stopped-FDR and SupFDR guarantees.
An explicit bivariate mixture breaks the conjectured six-mode cap; more peaks than components are real.
· “At least seven modes in a heteroscedastic three-component bivariate Gaussian mixture”
Monotonicity under data processing and additivity on independent products force every such functional to an integral over four strata
Proportional-regime analysis shows attack success peaks then falls while clean performance improves with training trigger strength.
· “When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks”
Vertices are realizable label sequences of length m; edges mark label disagreements on shared points, fixing whether dimension meets or tops
CSA maintains per-round selective risk bounds under predictable updates without pooling across deployments.
· “Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs”
Three probe conditions suffice for identification and replace ambient d sqrt(T) regret with intrinsic r sqrt(T) plus detection cost.
· “Catching a Moving Subspace: Low-Rank Bandits Beyond Stationarity”
Tests on gridworlds, chess and terminals separate solved perceptual errors from persistent multi-step and abstention failures.
· “HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models”
Proves minimax Θ(s/√N) bound and faster Θ(s/N) rate when used only to warm-start the solver.
· “Provably Data-driven Lagrangian Relaxation for Mixed Integer Linear Programming”
Gradient flow no longer selects the vanishing ridge solution outside the kernel regime, so a function-space energy defines geodesic ridge as
· “Canonical Regularisation of Wide Feature-Learning Neural Networks”
Different orderings and couplings reduce to the same canonical form whose non-escape coordinates are orthogonal and sum to the identity.
when initial loss is small and log-sum-exp functions remain linearly independent modulo affine functions
A van Trees bound shows each client's samples and privacy budget alone set the limit.
· “A Van Trees Lower Bound for Fully Interactive Differentially Private Federated Learning”
Extending delay-thresholding to LMO momentum handles heterogeneous workers and matches best known bounds in smooth cases.
· “Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method”
Upper-tail accumulation scale unifies conflicting laws for rescaling inverse temperature with context length n.
· “A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention”
Covariance decomposition isolates a universal Gaussian term plus explicit fourth-order adjustments for linear statistics of high-dimensional
Under Gaussian transport the velocity field equals a Nadaraya-Watson smoother whose bandwidth shrinks from global to nearest-neighbor.
The method outperforms benchmarks in Delhi and Mumbai while staying stable during spikes and seasons and adding uncertainty estimates.
Benchmark on two cohorts shows CFN edges linear CLR in colon while compartment labels remove spurious dependency instability.
Each control is regressed against Brownian labels before the value, giving stable joint value-and-control accuracy.
· “Deep-Control BSDE: Layerwise Brownian-Weighted Regression for High-Dimensional Semilinear PDEs”
Three-way dominance was held for years, but peak distance above the field did not eclipse the 1980s.
· “How exceptional was the Big Three era? Extremes and persistence in men's professional tennis”
A new framework separates evidence, power, and stability—none can rescue a failed one.
New proof separates an m-dependent warm-up from a stochastic term that does not grow with the number of quantiles.
At any finite sample size, the score 1/p(Y|X) yields the smallest average prediction set among all valid methods.
Optimal edge universality holds even when signal and noise covariance do not commute.
· “Local Laws and Edge Universality for Noncentral Sample Covariance Matrices”
Artist-to-artist distances from full acoustic distributions recover expert review co-mentions, rising to 0.865 when critics agree.
In a nicotine trial, model averaging trimmed the interval 13% beyond covariate adjustment.
A KL-penalized optimization framework for adaptive trial enrichment yields continuous enrollment mixtures and an asymptotic variance…
· “A Unified Adaptive Enrichment Design for Power Enhancement”
Centrosymmetric kernels make rate-optimal tight differencing compatible with kernel-based spectral density and long-run variance…
· “Tight differencing in spectral density estimation with centrosymmetric kernels”
Fine-tuning Whisper Small on 1,373 elicited recordings gives Baniwa its first ASR baseline.
· “Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study”
Support reduction makes the NPMLE a finite convex problem; matching impossibility appears at sqrt n noise.
· “Statistical Properties of Nonparametric MLE under Laplace Noise”
Adding explicit covariate controls to the network's final layer removes the bias that decorrelation methods miss.
· “Controlling for Omitted Variable Bias in Deep Neural Networks”
A rotation-based test beats classic omnibus tests on single-lag power and spots the 11-year solar cycle without Monte Carlo.
· “Random Invariance Testing on Quadratic Form Statistics with Application to Autocorrelation”
A density-matrix representation of each driver also reproduces the fundamental diagram and hysteresis on I-24 highway data.
A conditional-sparsity argument yields sharp weighted L^p estimates on every filtered probability space.
· “Weighted Estimation by Discrete-time Sparse Domination on Martingale Spaces”
The proven convergence powers fast projections that separate clusters and outliers in high-dimensional data.
· “Efficient Estimation of High Information Projections using Nearest Neighbours”
Consistent Bayesian inference for set-identified models, with no covariate binning or ad hoc moment selection.
· “Nonparametric Bayesian Inference for Partially Identified Discrete Response Models”