Introduces a Q-sort protocol using human reference factors to quantify LLM value-structure alignment via Procrustes similarity and RSA correlations, revealing cross-family heterogeneity and localized misalignments.
hub Mixed citations
An Introduction to the Bootstrap
Mixed citation behavior. Most common role is background (60%).
hub tools
citation-role summary
citation-polarity summary
fields
cs.CL 5 cs.LG 3 cs.AI 2 cs.IR 2 math.ST 2 cs.CV 1 econ.EM 1 hep-ph 1 physics.flu-dyn 1 physics.soc-ph 1years
2026 24representative citing papers
Direct fixed-weight solver for free-support Wasserstein medians relocates atoms using OT barycentric projections and inverse-distance weights, achieving monotone descent on smoothed objectives with fewer subproblems than nested Weiszfeld baselines.
HTEB introduces dynamic, multi-axis evaluation of text embedding robustness using LLM transformations, finding decoupled profiles across models and that scaling does not close all robustness gaps.
SensorFault-Bench is a new CPS-grounded benchmark showing that clean-MSE rankings of forecasting models often disagree with their robustness under standardized sensor-fault scenarios across four real datasets.
Heat-kernel smoothing over weighted points on a compact manifold yields a scale-dependent geometric effective sample size that discounts nearby and duplicate particles.
Fine-tuning Tesseract on synthetic Maltese line images plus lexicon-gated arbitration of five recognizer streams reduces development-set CER from 0.0234 to 0.01317 (and 0.00700 with label normalization).
Proves finite-shot mean-squared-error laws for virtual distillation and symmetry verification that define certified operating windows and a selection trichotomy for their comparison.
The IM interval is the shortest valid prior-free procedure for the Behrens-Fisher problem, established via cylindrical predictive random sets, minimaxity, admissibility, and a projection argument.
An adversarial methodology generates a multilingual cross-platform dataset of paired human-AI social messages, and models trained on it outperform prior detectors on real-world out-of-distribution data.
AEGIS uses activation probes for early-warning detection of high-risk steps in weak policies and selectively escalates to stronger policies, recovering 10.1% of lost trajectories on LIBERO-Spatial while activating the strong policy on only 38% of steps.
Instruction-tuned LLMs exhibit an ownership bias, assigning up to 26% higher confidence to their own responses than identical user-provided answers; reframing the answer as user input during elicitation reduces overconfidence by up to 26%.
The authors propose target-space recovery profiles to diagnose which reproducible dimensions of fMRI brain responses are captured by model predictions, showing that accuracy alone can mask alignment mismatches in visual cortex.
EnergyAgentBench, a 70-task live-data agentic benchmark, finds Claude Sonnet 4.6 best overall and Haiku 4.5 best on long-horizon siting among nine LLM agents.
CARE, a context-aware LLM judge, outperforms standard methods when evaluating multi-hop retrieval quality in RAG systems.
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
Global QCD analysis extracts genuine twist-three PDFs from g2, d2 and SIDIS asymmetries, confirming their universality and factorization validity.
A review reframing density estimation as 'density evolution' across scales, linking kernel smoothing to heat flow, mixtures to compression, and topology to level sets, while stating three structural results on modes, Gaussian semigroups, and log-concavity.
The OSS Challenge provides benchmarks showing spatiotemporal video models excel at open suturing skill classification and OSATS scoring but struggle with keypoint tracking under occlusion.
Crowdsourced judgments reliably flag authentic videos but frequently miss manipulations and struggle to identify whether changes are audio-only, video-only, or both.
AVVA is a new framework adapting verbal analysis for classroom discourse with triangulation across ten steps and a four-criterion validation scheme for temporal stability, applied to 23 hours of recordings.
The deep SPAR model shows concurrent floods and droughts becoming more likely in the Upper Danube by 2100 under high emissions, with changes in the dependence between catchments contributing substantially to the increase.
Bayesian-ARGOS is a hybrid frequentist-Bayesian method that discovers equations from limited noisy observations more efficiently than SINDy or bootstrap-ARGOS while adding uncertainty quantification.
Empirical study of frontier AI on Project Euler finds power-law machine effort scaling with human difficulty (b<1 for 20/25 models) and moderate support for exponential success probability decay, with SOTA 50% horizons at 2.5-4.3 human hours.
Data-driven equation discovery applied to liquid film flows identifies identifiability issues from multi-collinearity in monomial bases and early-time transients with large residuals.
citing papers explorer
-
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts
Introduces a Q-sort protocol using human reference factors to quantify LLM value-structure alignment via Procrustes similarity and RSA correlations, revealing cross-family heterogeneity and localized misalignments.
-
Fast Computation of Free-Support Wasserstein Medians
Direct fixed-weight solver for free-support Wasserstein medians relocates atoms using OT barycentric projections and inverse-distance weights, achieving monotone descent on smoothed objectives with fewer subproblems than nested Weiszfeld baselines.
-
The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness
HTEB introduces dynamic, multi-axis evaluation of text embedding robustness using LLM transformations, finding decoupled profiles across models and that scaling does not close all robustness gaps.
-
Benchmarking Sensor-Fault Robustness in Forecasting
SensorFault-Bench is a new CPS-grounded benchmark showing that clean-MSE rankings of forecasting models often disagree with their robustness under standardized sensor-fault scenarios across four real datasets.
-
Heat-Kernel Entropy Profiles and Geometric Effective Sample Size for Weighted Measures on Manifolds
Heat-kernel smoothing over weighted points on a compact manifold yields a scale-dependent geometric effective sample size that discounts nearby and duplicate particles.
-
LV-ROVER-MLT: Low-Resource Maltese OCR by Synthetic Fine-Tuning and Multi-Stream Arbitration
Fine-tuning Tesseract on synthetic Maltese line images plus lexicon-gated arbitration of five recognizer streams reduces development-set CER from 0.0234 to 0.01317 (and 0.00700 with label normalization).
-
Certified Finite-Shot Operating Windows for Virtual Distillation and Symmetry Verification
Proves finite-shot mean-squared-error laws for virtual distillation and symmetry verification that define certified operating windows and a selection trichotomy for their comparison.
-
Revisiting the Behrens-Fisher Problem: Validity-First Optimality
The IM interval is the shortest valid prior-free procedure for the Behrens-Fisher problem, established via cylindrical predictive random sets, minimaxity, admissibility, and a projection argument.
-
Adversarial Creation and Detection of AI-Generated Social Bot Content
An adversarial methodology generates a multilingual cross-platform dataset of paired human-AI social messages, and models trained on it outperform prior detectors on real-world out-of-distribution data.
-
AEGIS: A Backup Reflex for Physical AI
AEGIS uses activation probes for early-warning detection of high-risk steps in weak policies and selectively escalates to stronger policies, recovering 10.1% of lost trajectories on LIBERO-Spatial while activating the strong policy on only 38% of steps.
-
Large Language Models Are Overconfident in Their Own Responses
Instruction-tuned LLMs exhibit an ownership bias, assigning up to 26% higher confidence to their own responses than identical user-provided answers; reframing the answer as user input during elicitation reduces overconfidence by up to 26%.
-
Beyond Prediction Accuracy: Target-Space Recovery Profiles for Evaluating Model-Brain Alignment
The authors propose target-space recovery profiles to diagnose which reproducible dimensions of fMRI brain responses are captured by model predictions, showing that accuracy alone can mask alignment mismatches in visual cortex.
-
EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data
EnergyAgentBench, a 70-task live-data agentic benchmark, finds Claude Sonnet 4.6 best overall and Haiku 4.5 best on long-horizon siting among nine LLM agents.
-
Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies
CARE, a context-aware LLM judge, outperforms standard methods when evaluating multi-hop retrieval quality in RAG systems.
-
The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
-
Phenomenology of genuine twist-three distributions from a global QCD analysis
Global QCD analysis extracts genuine twist-three PDFs from g2, d2 and SIDIS asymmetries, confirming their universality and factorization validity.
-
Density Evolution: A Multiscale View of Density Estimation
A review reframing density estimation as 'density evolution' across scales, linking kernel smoothing to heat flow, mixtures to compression, and topology to level sets, while stating three structural results on modes, Gaussian semigroups, and log-concavity.
-
OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025
The OSS Challenge provides benchmarks showing spatiotemporal video models excel at open suturing skill classification and OSATS scoring but struggle with keypoint tracking under occlusion.
-
Beyond Seeing Is Believing: On Crowdsourced Detection of Audiovisual Deepfakes
Crowdsourced judgments reliably flag authentic videos but frequently miss manipulations and struggle to identify whether changes are audio-only, video-only, or both.
-
Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues
AVVA is a new framework adapting verbal analysis for classroom discourse with triangulation across ten steps and a four-criterion validation scheme for temporal stability, applied to 23 hours of recordings.
-
Exploring climate change effects on concurrent floods and concurrent droughts via statistical deep learning
The deep SPAR model shows concurrent floods and droughts becoming more likely in the Upper Danube by 2100 under high emissions, with changes in the dependence between catchments contributing substantially to the increase.
-
Fast and principled equation discovery from chaos to climate
Bayesian-ARGOS is a hybrid frequentist-Bayesian method that discovers equations from limited noisy observations more efficiently than SINDy or bootstrap-ARGOS while adding uncertainty quantification.
-
Human vs Machine Mathematical Difficulty on Project Euler: An Experimental Analysis
Empirical study of frontier AI on Project Euler finds power-law machine effort scaling with human difficulty (b<1 for 20/25 models) and moderate support for exponential success probability decay, with SOTA 50% horizons at 2.5-4.3 human hours.
-
Data-Driven Equation Discovery for Nonlinear Liquid Film Flows
Data-driven equation discovery applied to liquid film flows identifies identifiability issues from multi-collinearity in monomial bases and early-time transients with large residuals.