Exact risk ratios for data selection in linear regression
Proves 1+1/d at 2d-1, 5/3 for (3,4), 2 for (4,5); gives lower bound.
· “Exact Risk Ratios for Weighted Data Selection in Linear Regression”
Machine Learning
Papers on all aspects of machine learning research (supervised, unsupervised, reinforcement learning, bandit problems, and so on) including also robustness, explanation, fairness, and methodology. cs.LG is also an appropriate primary category for applications of machine learning methods.
sort pith recommended most recent
Proves 1+1/d at 2d-1, 5/3 for (3,4), 2 for (4,5); gives lower bound.
· “Exact Risk Ratios for Weighted Data Selection in Linear Regression”
Fine-tuning and sampling share one iteration with global descent and quadratic convergence.
· “Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling”
Bound splits into drift and noise budgets; correctly predicts learning order in linear regression.
· “Speed Limit for Information Acquisition in Stochastic Learning Dynamics”
An overhead h works exactly when the sum of 1/h(2^j)² converges; the √(log n/n) rate stays just out of reach.
· “The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives”
New lower bounds match known upper bounds, pinning the fastest polynomial convergence exponent
· “Silver Rate Is (Almost) Optimal for Gradient Descent Acceleration”
Measuring k copies together buys a √k speedup, and no adaptive protocol can beat the rate.
· “Tight Lower Bounds for State Tomography with Limited Entanglement”
Each generated token leaks at most bT of mutual information about the private secret; unanimous votes add no noise.
· “PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement”
Training a student model on one or two tokens per response matches or beats dense supervision in nine configurations.
· “Extremely Sparse Supervision Incentivizes Reasoning Ability”
A single scalar measured in seconds forecasts which LLMs keep instruction-following after VLM training
· “Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis”
Distance-based scores flag the most reproducible verdicts as most uncertain, making abstention worse than random.
· “Verdict Instability of OOD Scores under Reference Resampling”
Index-set-dependent weights detect smoothness invisible to the isotropic Barron scale; the result is uniformly sharp.
· “Sharp Mixed Spectral Barron Regularity of Coulombic Many-Electron Wave Functions”
Cutwidth and tree-cutwidth govern both the compression to MPS or TTN and the sample complexity of tomography, in realisable and agnostic…
For rank-r∗ targets, every global minimizer attains the optimal error at any search rank r ≥ r∗.
· “Sharp Restricted Isometry Thresholds for Global Minima of Rank-Restricted Matrix LASSO”
For sparse data the full nonlinear dynamics reduces to independent scalar flows with explicit convergence times.
Under ETH for PPAD, polynomial CCE computation is impossible; first radically uncoupled algorithm matches the hardness bound.
· “Independent Reinforcement Learning in Discounted Markov Games”
Bipartite message-passing GNNs outperform AMO and ICD, cutting inference time tenfold and resisting beam squint across 64 subcarriers.
· “Efficient Graph Neural Networks for Multicarrier Wideband Hybrid Beamforming Optimization”
Training-free preprocessing removes edges that inflate a generalization bound, improving GNN robustness across datasets.
· “Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks”
By adapting a classic propensity-score stratification to the private setting, the new DPBlocking algorithm outperforms earlier methods…
· “Differentially Private Average Treatment Effect Estimation by Propensity Score Blocking”
A novel transformer with bidirectional time awareness forecasts events, outcomes, and counterfactual effects from lifelong multimodal…
By constructing probe inputs and masking low-consensus updates, zero-shot transfer improves on multiple benchmarks.
Upper and lower bounds match up to log factors, removing the earlier sqrt(n) factor and showing input dimension does not control sample…
· “Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks”
Solving the dual recovers the likelihood from the posterior, with Amari's coordinate swap as a special case.
Exact law: one scalar sets effective step size; accuracy peaks at the instability boundary from MNIST to GPT-2.
· “When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay”
Ten agent configs near expert selectivity but score half as well on steering — causal validation is the missing skill.
· “SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?”
A controlled synthetic study across 12 tasks shows that easy-to-hard ordering, pacing, and smoothness each matter in different contexts.
· “Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics”
Using a cheap offline model to estimate difficulty before training, ThinkPrior cuts silent groups by 55% with no accuracy loss.
· “ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR”
Multi-task learning across 21 grape cultivars consistently beats single-task models and the current scientific standard.
· “Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling”
A browser-free backend runs the same game file humans play, beating classic benchmark environments on speed.
The method converts activation steering directions into rank-one weight updates, allowing addition, subtraction, and composition of…
Controlled study shows users prefer planning-based traces but simple chain-of-thought yields better error detection and trust calibration.
· “Do Reasoning Representations Help Humans Evaluate LLM Outputs?”
Novel representation tracks how probability shifts among answers, revealing differences that accuracy and entropy miss.
· “Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning”
Method avoids degeneracy that plagues joint fits, recovers self-gravitating disk shape and first analytical distribution function.
· “Closed-Form of the Local Galactic Potential and Stellar Distribution Function from Gaia DR3”
Continuous diffusion models learn constraints, but standard DDPM sampling locks in early errors.
· “Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks”
A hybrid structural-semantic KG metric reveals domain complexity as a key bottleneck in LLM QA.
· “Evaluation of Contextual Understanding in Large Language Models”
A three-channel scattering layer guarantees T+R+A=1 to machine precision but shows no accuracy gain over a six-word rule.
Requiring plausible blood-pressure waveforms removes ICU artifacts that mimic ventricular tachycardia.
· “Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU”
Proving that frozen attention layers simulate closed-form diffusion and particle-based generation from prompts alone.
· “Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling”
A 9B omni interaction agent with Cerebellum-Brain architecture sets new interaction records while handling realime agentic tasks.
A distibuted system produces interpretable structural statistics that beat expert crafting and GNNs on Alipay transaction netwks with…
Where your connectives sit in Post's lattice decides what you can fit, compress, and learn.
· “Fitting and Learning Basis-Restricted Propositional Formulas”
Level-set encoding and divergence regularization stabilize long-horizon predictions across laminar, transitional, and turbulent regimes.
· “ONE CYLinder: A Benchmark for Graph-Based Surrogate Modeling of Unsteady Bluff-Body Flows”
A weight-preserving BatchNorm update unmasks a measurement artifact that skews unlearning evaluation
· “The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation”
Fixing the spin count removes the low-temperature barrier and cuts sparse Bayesian sampling to k^1.5 measurements.
Transition-action pretraining lets the model respond to user edits without manual action labels.
· “Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems”
Run-length-compressed training strings make the required training data single-exponential in size.
Under a single-common-factor model, anchor–judge error correlation is identified in closed form with a diagnostic battery for violations.
Keeps gradients reliable at 0.1 s steps: 8,192 worlds on one GPU, control tasks MJX cannot solve.
· “Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics”
OPRD amplifies the aligned component of a student's own reward gradient along a teacher's learned shift, beating the teacher with fewer…
· “Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation”
An axis-factorized transformer gains 3–5 balanced-accuracy points on six EEG tasks, and the same split carries to audio.
· “Adaptive Anisotropic Attention for Axis-Structured Signals”
A certificate-gated adjudication rule bounds innocent-naming probability to 10⁻³ per investigation, verified on federated GNSS and…
Pointwise +1 identity holds exactly for each Haar rotation; classical relation only in large-d limit
Theorem 3.1 gives O((ln N)^r/√N) rates for bounded, Lipschitz, and quadratic loss, plus parameter estimation.
Dropping only logs where control and treatment disagree keeps the model closer to an unbiased reference.
· “BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests”
One architecture handles node, link, and graph classification on unseen datasets with few labeled examples.
· “Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning”
A four-stage transformer model built on a semantic-first codec runs on a single RTX 5090 and rivals top commercial voices in audiobook…
· “TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context”
Two power laws add the activation ratio as a third dimension, letting hyperparameters transfer across sparsity levels.
A plug-in framework uses pseudo-extrapolation along cross-class edges to reject unknowns on heterophilic graphs.
· “HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation”
A method combining oscillatory neurons and predictive self-supervision achieves strong adversarial robustness without ever generating…
Estimator matches optimal prediction error of continuous setting when grid is fine enough.
· “Optimal estimation for Functional Linear Regression with Noisy Discretized Data”
A novel extension of soft actor-critic handles mixed actions and graph states, achieving real-robot closed-loop assembly.
· “Learning to build covering structures with continuous adjustments”
MOEMB achieves 70.4% on MMEB-V2 with 3.1B parameters, surpassing TTE models using over 12B parameters and less compute.
· “MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models”
Vision-language models identify chart roles and links but cannot resolve which component is in front, a new benchmark shows.
· “Charts Are Beyond Pixels: Probing for Layer-Wise Chart Understanding and Editing”
DATPO combines adaptive rollout and sentence-level diversity to improve test-time scaling in mathematical reasoning.
· “Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR”
A product of successor value and pseudocount jointly scores novelty and reachability, outperforming tuned additive methods with fewer…
A margin-only preference loss keeps helpfulness and compliance at SFT level while matching DPO's harmlessness.
· “Suan: Rectifying Direct Preference Safety Alignment in Large Language Models”
A conditional VAE learns patient-specific deformation modes and generates realistic synthetic repeat CTs for adaptive proton therapy.
· “SynthRCT: Scalable Conditional Deformation Synthesis for Synthetic Repeat CT Generation”
Four encoding strategies tested; best model uses a distinct context graph with full prefix histories, outperforming prior graph-based…
· “Leveraging contextual events on structure-aware next activity prediction”
Auto-hetSNGP achieves 96.8% accuracy on dynamic aperture while converging 3× faster than standard DDU and producing uncertainty maps…
· “Flexible Spectral-Normalized Neural Gaussian Process for Dynamic Aperture Prediction”
Standardized pulses of micro-training reveal a checkpoint's hidden response state, cutting prediction error by up to 78% over benchmark…
Replacing one salience per feature with a per-outcome matrix prevents both bound collapse and sign cancellation.
· “Why shared attention vectors fail: a case for outcome-indexed tuning”
Multi-level-set physics-driven solver yields clear boundaries and uniform material regions without manual tuning.
· “Multi-Level-Set-Based Physics-Driven Neural Network to Solve 3-D Inverse Scattering Problems”
FedGenSC uses local discriminators and a prototype bank to improve text reconstruction across all SNRs.
· “FedGenSC: Federated Generative Semantic Communication with Channel-Aware Adaptation”
When LLMs evaluate explanations, they often just solve the task instead, and anonymization merely rewards label leakage.
· “Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations”
Contrastive pre-training on CMR data transfers structural heart knowledge to ECG, boosting performance even for a disease never seen…
AlphaRJM wins 10 of 12 metric comparisons by retaining evaluation history across formula episodes.
· “AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery”
For all moment regimes, no query interaction needed; full sample-interval tradeoff characterized.
· “Non-Adaptive 1-Bit Mean Estimation: Minimax Rates and the Sample-Interval Tradeoff”
Certified topological scans of 111 networks show pairs carry nearly all class-overlap structure
A single-backward-pass method reconstructs per-variable gradients and aligns them, improving forecasts on seven datasets.
A signed diagnostic reveals when temporal attention is too sharp or too diffuse; a query-temperature regulator restores coherence at…
· “Temporal State Transport in Video Generation: Diagnosing and Correcting Spectral Imbalance”
Lattice reduction attacks can reconstruct private model updates from secure neighborhood aggregates in decentralized federated learning…