INDEX
Every paper Pith has read and judged, in one searchable index.
-
stat.ML arXiv submitted 2026-06-29Shell-core geometry produces grokking scaling laws
R\'ois\'in Luo +4 · “A Stochastic--Geometric Theory of Scaling Laws in Grokking”
2606.30388 -
math.ST arXiv submitted 2026-06-29Horseshoe posterior rules attain optimal detection boundary
Sayantan Banerjee +2 · “Multiple testing with the horseshoe”
2606.30375 -
math.ST arXiv submitted 2026-06-29Horseshoe prior controls FDR at optimal detection boundary
Sayantan Banerjee +2 · “Multiple testing with the horseshoe”
2606.30375 -
stat.ML arXiv submitted 2026-06-29Extrapolation beats single-nugget solves for ill-conditioned systems
Disha Hegde +2 · “Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning”
2606.30328 -
stat.ME arXiv submitted 2026-06-29Conditional test unifies HWE check and SNP association in GWAS
Stefan B\"ohringer +1 · “Evaluating HWE and Association in Genome Wide Association Studies: A Unified Procedure”
2606.30311 -
stat.ML arXiv submitted 2026-06-29CDFs replace sorting for sliced Wasserstein distance
Christophe Vauthier +2 · “Highly Data Parallelizable Estimation of the Sliced-Wasserstein Distance Using Cumulative Distribution Functions”
2606.30310 -
math.ST arXiv submitted 2026-06-29Differential algebra checks unique recovery of functions in DE models
Torkel E Loman +2 · “Structural functional identifiability and model discovery in differential equation models”
2606.30289 -
physics.geo-ph arXiv submitted 2026-06-29Map all of North America's crustal motion 1,000× faster
Lionel Voirol +5 · “Towards Open Science: Monitoring Crustal Deformations in North America”
2607.16264 -
cs.LG arXiv submitted 2026-06-29Level sets yield exact acceptance certificates for speculative decoding
Aaryam Sharma · “When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding”
2606.30265 -
cs.LG arXiv submitted 2026-06-29TabPFN v2 leads NHANES benchmark for HbA1c and CRP
Federico Felizzi · “Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification”
2606.30702 -
math.ST arXiv submitted 2026-06-29Optimal confidence sets induce Choquet-risk optimal contours
Max Raner · “Efficiency of Valid Inferential Models: Choquet-risk Optimal Possibility Measures, and Direct Comparisons”
2606.30229 -
stat.ML arXiv submitted 2026-06-29Optimal transport unifies flow matching and Schrödinger bridge
Titouan Vayer (COMPACT) · “Notes on generative modeling: flow matching, diffusion, optimal transport and Schr{\"o}dinger bridge”
2606.30053 -
math.ST arXiv submitted 2026-06-29Copulas obey |β|^3 ≤ 2ξ exactly for Chatterjee ξ and Blomqvist β
Jacob Israel Orenday Lares +1 · “The exact region between Chatterjee's $\xi$ and Blomqvist's $\beta$”
2606.30033 -
math.ST arXiv submitted 2026-06-29Explicit MSE bounds for adaptive rare MCMC under Wasserstein contraction
Julian Hofstadler +1 · “Error bounds for simultaneous Wasserstein contractive adaptive increasingly rare MCMC”
2606.30018 -
math.ST arXiv submitted 2026-06-29Common noise repeats speed nonparametric regression rates
Fabienne Comte (MAP5 - UMR 8145) +1 · “Adaptive nonparametric regression from repeated measurements under common noise”
2606.30000 -
math.ST arXiv submitted 2026-06-29Frank-Wolfe computes optimal e-values for non-convex voting tests
Adrienne Tuynman +1 · “Optimal Posterior E-values with Non-Convex Parameter Sets with Applications to Voting Systems”
2606.29998 -
stat.ME arXiv submitted 2026-06-29Ordinal time series model estimates category spacings from data
Anna Nalpantidi +2 · “Beyond Equidistant Assumptions: An Autoregressive Ordered Stereotype Model for Ordinal Time Series”
2606.29931 -
math.ST arXiv submitted 2026-06-29Four Lorenz curve forms fail validity conditions
Jos\'e Mar\'ia Sarabia +3 · “Revisiting "A universal model for the Lorenz curve with novel applications''”
2606.29923 -
cs.AI arXiv submitted 2026-06-29Personal decision theory optimal under population enforcement metric
Arvid Sj\"olander · “A causal modeling perspective on decision theory”
2606.29911 -
math.OC arXiv submitted 2026-06-29AdaGrad misses Hölder rate on composite problems
Matia Bojovic +2 · “AdaGrad does not adapt to H\"older-smoothness for composite objectives”
2606.29893 -
cs.LG arXiv submitted 2026-06-29Shapley values track decision value not forecast error
Konstantinos Ziliaskopoulos +2 · “Decision-Value Attribution in Predict-then-Optimize Systems”
2606.29878 -
physics.data-an arXiv submitted 2026-06-29Squeals flag poor model fits while users drag curves
Andrew Gelman +3 · “The Squealer: Sensification of model exploration and model misfit”
2606.29842 -
cs.CR arXiv submitted 2026-06-29Exact privacy math for 2020 Census now 1824 times faster
Buxin Su +2 · “A Sieve-Accelerated Quadrature Method for Exact Privacy Accounting in the 2020 U.S. Decennial Census”
2606.29835 -
stat.ME arXiv submitted 2026-06-29New downscaling method matches kriging accuracy at far lower compute cost
Daisuke Murakami +3 · “Scalable coarse-to-fine spatial downscaling”
2606.29798 -
cs.LG arXiv submitted 2026-06-29Autoencoder memorizes inliers before outliers in early training
Kunwoong Kim +1 · “What Drives the Inlier-Memorization Effect? A Theory of Outlier Detection via Early Training Dynamics”
2606.29791 -
stat.ME arXiv submitted 2026-06-29Historical data cuts bias and variance in model evaluations
Xinrui Ruan +6 · “HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data”
2606.29784 -
stat.AP arXiv submitted 2026-06-29Statistics students should probe LLMs as stochastic systems
Tian Zheng · “Probing the Stochastic Machine: Engaging with LLMs in Statistics Curricula Through Veridical Data Science”
2606.29754 -
stat.ME arXiv submitted 2026-06-29Symmetric noise enables valid tests after data-driven selection
Ameer Dharamshi +2 · “Testing hypotheses via orthogonalization”
2606.29732 -
cs.LG arXiv submitted 2026-06-29Distance matrix yields latent manifold dimension from eigenvalue multiplet
Igor Halperin · “I-BBS: Coordinate-Free Inference of Latent Sub-Manifolds Using Random Distance Matrix Theory”
2606.29675 -
stat.ML arXiv submitted 2026-06-29Adjusted Wasserstein distance improves MDS on heavy-tailed data
Flor Martinez-Sermeno +2 · “Adjusted Wasserstein distances for bridging empirical and true distributions with applications to MDS”
2606.29665 -
stat.ME arXiv submitted 2026-06-28Summary statistics enable privacy-preserving transfer for single-index models
Ye Tian · “Multi-Source Transfer Learning of Sparse Single-Index Models”
2606.29658 -
stat.ME arXiv submitted 2026-06-28Shared blocks enable consistent class graph recovery
Seunghyun Lee +1 · “Beyond Local Independence: High-Dimensional Latent Class Graphical Models with Shared Block Structure”
2606.29631 -
math.NA arXiv submitted 2026-06-28Malliavin weights cut fade-probability variance by up to 2516x
Francisco Delgado-Vences · “Stochastic Analysis of Fade Duration Using Wiener Chaos Expansion and Malliavin Calculus: Optimal Importance Sampling via Adaptive SGD”
2606.30692 -
stat.ML arXiv submitted 2026-06-28Bidirectional diffusion checks MHD prediction errors without ground truth
Alexander Scheinker · “Bidirectional Autoregressive Latent Diffusion for Forward and Inverse Magnetohydrodynamics”
2606.29620 -
cs.LG arXiv submitted 2026-06-28AI settles worst-case complexity of 1937 Kaczmarz algorithm
Micha{\l} Derezi\'nski +1 · “How AI settled the complexity of the oldest SGD algorithm”
2606.29593 -
cs.LG arXiv submitted 2026-06-28Optimizer memory makes shuffle order first-order fine-tuning noise
John Sweeney · “Optimizer Memory Makes Shuffle Order a First-Order Source of Fine-Tuning Noise”
2606.29554 -
cs.LG arXiv submitted 2026-06-28Reweight samples to reduce per-sample harm and lift accuracy
Apostolos Avranas · “Reducing Per-Sample Harm in Stochastic Optimization”
2607.16261 -
stat.ME arXiv submitted 2026-06-28Mixture models distinguish mild from gross anomalies in circular data
Antonio Punzo +4 · “Modelling and detecting mild and gross anomalies in circular data via double-contaminated models”
2606.29524 -
cs.LG arXiv submitted 2026-06-28Gradient tweak keeps primary descent intact while meeting secondary goals
Dara Varam +1 · “Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization”
2606.29521 -
cs.LG arXiv submitted 2026-06-28Expert priors enter best-subsets MIO as log-odds penalties
Nolan Alexander +1 · “A Mathematical Optimization Approach for Expert-Informed Bayesian Best Subset Selection”
2606.29516 -
stat.ME arXiv submitted 2026-06-28Bayesian model segments satellite images without local labels
Bao Khanh Nguyen +3 · “Scalable Bayesian Spatial Mixture Modelling for Remote Sensing Image Segmentation”
2606.29448 -
stat.ML arXiv submitted 2026-06-28Unsupervised maps cut regional coverage gaps in conformal prediction
Louis Berthier +4 · “Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery”
2606.29403 -
stat.AP arXiv submitted 2026-06-28Bayesian CDD alone stays accurate on gene directions across all DREAM5 networks
Xiaoying Wei +1 · “Bayesian Copula Directional Dependence is Cross-Network Robust for Gene-Regulatory Pair Direction: A Benchmark Study on DREAM5”
2606.29402 -
stat.AP arXiv submitted 2026-06-28Data on 37 nurses disproves modeling assumption in roster analysis
Richard D. Gill · “Critique of "Use of roster charts in the investigation and prosecution of nurses ..." by John O' Quigley”
2606.29394 -
stat.AP arXiv submitted 2026-06-28Language model embeddings beat hand-crafted features in car insurance pricing
Christopher Blier-Wong +1 · “Semantic insurance pricing with large language models”
2606.29371 -
cs.LG arXiv submitted 2026-06-28MMLU leaderboard ranks widen when subjects are the sampling unit
Bitya Neuhof +1 · “Quantifying Ranking Uncertainty in LLM Benchmarks”
2607.16259 -
cs.LG arXiv submitted 2026-06-28Symbolic trees generalize at rate L^d over sqrt(n)
\c{S}uayp Talha Kocabay +2 · “Sample Complexity of Scientific Discovery: PAC Learnability of Compositional Function Trees”
2606.29331 -
stat.ML arXiv submitted 2026-06-28Gradient boosting extended to vector outputs in tree leaves
David Cortes · “Gradient boosting with vector-valued leafs”
2606.29326 -
stat.ML arXiv submitted 2026-06-28Attention operator compresses distributions losslessly for Transformers
Peilin Liu +1 · “Generalization Analysis of Transformers in Distribution Regression”
2606.29256 -
cs.LG arXiv submitted 2026-06-28Prices that double in a week are still forecastable
Ranuga Weerasekara +8 · “When Prices Double in a Week: Forecasting of Agricultural Volatility in Import-Isolated Markets”
2606.29248