Introduces a Q-sort protocol using human reference factors to quantify LLM value-structure alignment via Procrustes similarity and RSA correlations, revealing cross-family heterogeneity and localized misalignments.
hub Mixed citations
An Introduction to the Bootstrap
Mixed citation behavior. Most common role is background (60%).
hub tools
citation-role summary
citation-polarity summary
fields
cs.CL 5 cs.LG 3 cs.AI 2 cs.IR 2 math.ST 2 cs.CV 1 econ.EM 1 hep-ph 1 physics.flu-dyn 1 physics.soc-ph 1years
2026 24representative citing papers
Direct fixed-weight solver for free-support Wasserstein medians relocates atoms using OT barycentric projections and inverse-distance weights, achieving monotone descent on smoothed objectives with fewer subproblems than nested Weiszfeld baselines.
HTEB introduces dynamic, multi-axis evaluation of text embedding robustness using LLM transformations, finding decoupled profiles across models and that scaling does not close all robustness gaps.
SensorFault-Bench is a new CPS-grounded benchmark showing that clean-MSE rankings of forecasting models often disagree with their robustness under standardized sensor-fault scenarios across four real datasets.
Heat-kernel smoothing over weighted points on a compact manifold yields a scale-dependent geometric effective sample size that discounts nearby and duplicate particles.
Fine-tuning Tesseract on synthetic Maltese line images plus lexicon-gated arbitration of five recognizer streams reduces development-set CER from 0.0234 to 0.01317 (and 0.00700 with label normalization).
Proves finite-shot mean-squared-error laws for virtual distillation and symmetry verification that define certified operating windows and a selection trichotomy for their comparison.
The IM interval is the shortest valid prior-free procedure for the Behrens-Fisher problem, established via cylindrical predictive random sets, minimaxity, admissibility, and a projection argument.
An adversarial methodology generates a multilingual cross-platform dataset of paired human-AI social messages, and models trained on it outperform prior detectors on real-world out-of-distribution data.
AEGIS uses activation probes for early-warning detection of high-risk steps in weak policies and selectively escalates to stronger policies, recovering 10.1% of lost trajectories on LIBERO-Spatial while activating the strong policy on only 38% of steps.
Instruction-tuned LLMs exhibit an ownership bias, assigning up to 26% higher confidence to their own responses than identical user-provided answers; reframing the answer as user input during elicitation reduces overconfidence by up to 26%.
The authors propose target-space recovery profiles to diagnose which reproducible dimensions of fMRI brain responses are captured by model predictions, showing that accuracy alone can mask alignment mismatches in visual cortex.
EnergyAgentBench, a 70-task live-data agentic benchmark, finds Claude Sonnet 4.6 best overall and Haiku 4.5 best on long-horizon siting among nine LLM agents.
CARE, a context-aware LLM judge, outperforms standard methods when evaluating multi-hop retrieval quality in RAG systems.
Task-aligned supervised geometric stability predicts linear steerability with high accuracy while unsupervised stability detects representational drift earlier and with lower false alarms than CKA or Procrustes.
Global QCD analysis extracts genuine twist-three PDFs from g2, d2 and SIDIS asymmetries, confirming their universality and factorization validity.
A review reframing density estimation as 'density evolution' across scales, linking kernel smoothing to heat flow, mixtures to compression, and topology to level sets, while stating three structural results on modes, Gaussian semigroups, and log-concavity.
The OSS Challenge provides benchmarks showing spatiotemporal video models excel at open suturing skill classification and OSATS scoring but struggle with keypoint tracking under occlusion.
Crowdsourced judgments reliably flag authentic videos but frequently miss manipulations and struggle to identify whether changes are audio-only, video-only, or both.
AVVA is a new framework adapting verbal analysis for classroom discourse with triangulation across ten steps and a four-criterion validation scheme for temporal stability, applied to 23 hours of recordings.
The deep SPAR model shows concurrent floods and droughts becoming more likely in the Upper Danube by 2100 under high emissions, with changes in the dependence between catchments contributing substantially to the increase.
Bayesian-ARGOS is a hybrid frequentist-Bayesian method that discovers equations from limited noisy observations more efficiently than SINDy or bootstrap-ARGOS while adding uncertainty quantification.
Empirical study of frontier AI on Project Euler finds power-law machine effort scaling with human difficulty (b<1 for 20/25 models) and moderate support for exponential success probability decay, with SOTA 50% horizons at 2.5-4.3 human hours.
Data-driven equation discovery applied to liquid film flows identifies identifiability issues from multi-collinearity in monomial bases and early-time transients with large residuals.
citing papers explorer
-
Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies
CARE, a context-aware LLM judge, outperforms standard methods when evaluating multi-hop retrieval quality in RAG systems.
-
Fast and principled equation discovery from chaos to climate
Bayesian-ARGOS is a hybrid frequentist-Bayesian method that discovers equations from limited noisy observations more efficiently than SINDy or bootstrap-ARGOS while adding uncertainty quantification.