Large language models display the identifiable victim effect at roughly twice the human baseline, strongly amplified by instruction tuning and chain-of-thought prompting but inverted by reasoning-specialized models.
hub Tool reference
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing , volume =
Tool reference. 83% of classified Pith citations use this work as a method, library, or software dependency, not as a substantive claim.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
LLMs show a probe-distance sign flip between finite and infinitival embedded clauses that UD cannot explain, claimed as evidence for Minimalist phase structure.
A hyperparameter-free angular kernel scan framework detects marginal distributional shifts in HDLSS regimes with theoretical guarantees under cross-coordinate mixing.
Proposes a scale-calibrated median-of-means estimator for robust aggregation of distributed PCA estimates on the product of Euclidean space and Grassmann manifold.
Human face perception aligns with neural networks trained on inverse-generative and naturalistic discriminative tasks, as these best predict human dissimilarity judgments on controversial and random face pairs.
ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.
The profile maximum likelihood estimator for the location in anisotropic hyperbolic wrapped normal models is strongly consistent, asymptotically normal, and attains the Hájek-Le Cam minimax lower bound under squared geodesic loss.
MuoFuzz improves greybox fuzzing by learning mutator sequence interactions to select effective orders, outperforming AFL++ and MOPT on coverage and unique bugs in FuzzBench and MAGMA.
Scoring the same cervical-spine segmentations against silver rather than expert labels overestimates Dice by ~8 points and turns a non-significant age fairness gap into a significant one via variance collapse.
Galaxy structure depends on environment at fixed stellar mass at >5σ, secondary to mass, with cluster-only effects at z<0.5 and multi-scale effects for massive systems at z≥0.5.
Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.
Fine-tuning Tesseract on synthetic Maltese line images plus lexicon-gated arbitration of five recognizer streams reduces development-set CER from 0.0234 to 0.01317 (and 0.00700 with label normalization).
Proposes sBDCA with preconditioning for the LTS estimator, claiming up to 3.25 times faster runtime and up to 90% lower objective values than Fast-LTS on synthetic and real data.
Introduces a matched four-condition protocol and ONCU metric to diagnose evidence utilization in long-context and RAG models across synthetic and multi-hop QA tasks.
No extraterrestrial technosignatures detected in an all-sky ultra-narrowband imaging survey at 50-86 MHz with OVRO-LWA, achieving ~100 Jy sensitivity and EIRP limits of 10^14 W at 10 pc.
Derives exact operating characteristic corrections and a numerical search over sample sizes to obtain optimal two-stage Bayes factor designs for two-arm binary-endpoint phase II trials that minimize expected sample size under the null.
A Bayesian framework decomposes mLLM variance, showing language features explain 79-92% of language identity variance and that model identity vs. benchmark-model interactions dominate differently for understanding versus reasoning tasks.
NLRΔ applied to 50 primary measures from five cognitive task families on public test-retest data yields median -0.138 nats with zero cells passing the headline rule across a 24-specification multiverse.
Sparse autoencoders applied to GPT-2 and Llama models recover semantic features accounting for 94% of peak brain encoding performance and map onto distinct cortical semantic regions across three languages.
Analysis of SATD in Dockerfiles shows 27% of admissions and 40% of repayments are coupled to non-Dockerfile artifacts, with coupled events repaid faster overall and external dependencies as a key trigger.
Proposes adaptive multiple importance sampling for robust Bayesian model evidence estimation under parameter non-identifiability, shown to outperform deterministic methods on ecological case studies while being cheaper than MCMC.
SPICE-RACS DR2 delivers the largest single Faraday rotation measure catalog from a radio survey, with 250,000-340,000 RMs across most of the sky at median uncertainty of 2 rad m^{-2}.
SVAR-FM uses simulator clamping to produce interventional distributions and flow matching to identify time series causal structures, with an error bound that predicts sign reversal of causal effects below a simulator accuracy threshold.
Several fairness impossibility results share an RKHS geometry where linear mean constraints are overdetermined by unequal base rates, yielding the Pokémon theorem on residual MMD violations and feature-learning collapse.
citing papers explorer
-
Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models
Large language models display the identifiable victim effect at roughly twice the human baseline, strongly amplified by instruction tuning and chain-of-thought prompting but inverted by reasoning-specialized models.
-
Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English
LLMs show a probe-distance sign flip between finite and infinitival embedded clauses that UD cannot explain, claimed as evidence for Minimalist phase structure.
-
High-Dimensional Change-Point Detection via Angular Kernel Statistics
A hyperparameter-free angular kernel scan framework detects marginal distributional shifts in HDLSS regimes with theoretical guarantees under cross-coordinate mixing.
-
Scale-Calibrated Median-of-Means for Robust Distributed Principal Component Analysis
Proposes a scale-calibrated median-of-means estimator for robust aggregation of distributed PCA estimates on the product of Euclidean space and Grassmann manifold.
-
Human face perception reflects inverse-generative and naturalistic discriminative objectives
Human face perception aligns with neural networks trained on inverse-generative and naturalistic discriminative tasks, as these best predict human dissimilarity judgments on controversial and random face pairs.
-
ProactBench: Beyond What The User Asked For
ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.
-
Profile Likelihood Inference for Anisotropic Hyperbolic Wrapped Normal Models on Hyperbolic Space
The profile maximum likelihood estimator for the location in anisotropic hyperbolic wrapped normal models is strongly consistent, asymptotically normal, and attains the Hájek-Le Cam minimax lower bound under squared geodesic loss.
-
On Interaction Effects in Greybox Fuzzing
MuoFuzz improves greybox fuzzing by learning mutator sequence interactions to select effective orders, outperforming AFL++ and MOPT on coverage and unique bugs in FuzzBench and MAGMA.
-
False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation
Scoring the same cervical-spine segmentations against silver rather than expert labels overestimates Dice by ~8 points and turns a non-significant age fairness gap into a significant one via variance collapse.
-
Environmental Sculpting of Galaxy Structure at Fixed Stellar Mass: A Multi-Scale Analysis Across Cosmic Time using 3 Million HSC Galaxies
Galaxy structure depends on environment at fixed stellar mass at >5σ, secondary to mass, with cluster-only effects at z<0.5 and multi-scale effects for massive systems at z≥0.5.
-
Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection
Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.
-
LV-ROVER-MLT: Low-Resource Maltese OCR by Synthetic Fine-Tuning and Multi-Stream Arbitration
Fine-tuning Tesseract on synthetic Maltese line images plus lexicon-gated arbitration of five recognizer streams reduces development-set CER from 0.0234 to 0.01317 (and 0.00700 with label normalization).
-
Faster than Fast-LTS: Robust Regression and Outlier Detection with DC Programming
Proposes sBDCA with preconditioning for the LTS estimator, claiming up to 3.25 times faster runtime and up to 90% lower objective values than Fast-LTS on synthetic and real data.
-
Diagnosing Evidence Utilization in Long-Context and Retrieval-Augmented Language Models under Matched Evidence Conditions
Introduces a matched four-condition protocol and ONCU metric to diagnose evidence utilization in long-context and RAG models across synthetic and multi-hop QA tasks.
-
All-Sky Ultra-Narrowband Spectral Imaging with the OVRO-LWA: Technosignature Constraints and Axion-Like Particle Prospects
No extraterrestrial technosignatures detected in an all-sky ultra-narrowband imaging survey at 50-86 MHz with OVRO-LWA, achieving ~100 Jy sensitivity and EIRP limits of 10^14 W at 10 pc.
-
Optimal sequential two-stage Bayes Factor Design for two-arm clinical Phase II Trials with binary Endpoints
Derives exact operating characteristic corrections and a numerical search over sample sizes to obtain optimal two-stage Bayes factor designs for two-arm binary-endpoint phase II trials that minimize expected sample size under the null.
-
DEPART: DEcomposing PARiTy across Multilingual LLMs
A Bayesian framework decomposes mLLM variance, showing language features explain 79-92% of language identity variance and that model identity vs. benchmark-model interactions dominate differently for understanding versus reasoning tasks.
-
Information-Theoretic Reliability is Robust to Analytic Choice: A 24-Specification Multiverse on Public Cognitive Test-Retest Data
NLRΔ applied to 50 primary measures from five cognitive task families on public test-retest data yields median -0.138 nats with zero cells passing the headline rule across a 24-specification multiverse.
-
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
Sparse autoencoders applied to GPT-2 and Llama models recover semantic features accounting for 94% of peak brain encoding performance and map onto distinct cortical semantic regions across three languages.
-
Beyond the Tip of the Iceberg: Understanding SATD in Dockerfiles through the Lens of Co-evolution
Analysis of SATD in Dockerfiles shows 27% of admissions and 40% of repayments are coupled to non-Dockerfile artifacts, with coupled events repaid faster overall and external dependencies as a key trigger.
-
Reliable model selection in the presence of parameter non-identifiability
Proposes adaptive multiple importance sampling for robust Bayesian model evidence estimation under parameter non-identifiability, shown to outperform deterministic methods on ecological case studies while being cheaper than MCMC.
-
The Rapid ASKAP Continuum Survey VII: Spectra and Polarisation In Cutouts of Extragalactic Sources (SPICE-RACS) Second Data Release -- Unveiling the Magnetised Sky
SPICE-RACS DR2 delivers the largest single Faraday rotation measure catalog from a radio survey, with 250,000-340,000 RMs across most of the sky at median uncertainty of 2 rad m^{-2}.
-
Intervention-Based Time Series Causal Discovery via Simulator-Generated Interventional Distributions
SVAR-FM uses simulator clamping to produce interventional distributions and flow matching to identify time series causal structures, with an error bound that predicts sign reversal of causal effects below a simulator accuracy threshold.
-
The Pok\'emon Theorem and other Fairness Impossibility Results
Several fairness impossibility results share an RKHS geometry where linear mean constraints are overdetermined by unequal base rates, yielding the Pokémon theorem on residual MMD violations and feature-learning collapse.
-
What Software Engineering Looks Like to AI Agents? -- An Empirical Study of AI-Only Technical Discourse on MoltBook
Empirical analysis of 4707 MoltBook posts shows AI-only technical discourse focuses on security, trust, and abstract topics while lacking concrete runtime and project details found in human GitHub discussions.
-
Scale selection for geometric medians on product manifolds
Joint location-scale minimization for geometric medians on product manifolds degenerates to marginal medians, and three new scale-selection methods restore identifiability with asymptotic guarantees.
-
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
Machine translation preserves embedding similarity structure for ten languages but distorts it for four in the Manifesto Corpus, via a new non-inferiority testing framework.
-
M-CaStLe: Uncovering Local Causal Structures in Multivariate Space-Time Gridded Data
M-CaStLe generalizes local stencil-based causal discovery to the multivariate case and decomposes resulting graphs into reaction and spatial components for interpretation in space-time gridded data.
-
The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious
42% of significant turn-level associations in LLM conversation analysis are spurious due to unaccounted autocorrelation, with a validated two-stage correction framework improving replication.
-
Fairboard: a quantitative framework for equity assessment of healthcare models
Patient identity and clinical features predict brain tumor segmentation accuracy more strongly than model choice, with localized spatial biases consistent across models and no formal fairness guarantees in any.
-
Rethinking player evaluation in sports: Goals above expectation and beyond
A double machine learning framework that residualizes standard outcome-above-expectation metrics to support valid frequentist inference and player-specific effect estimation in sports analytics.
-
Topological Signatures of Diffusive Release in Porous Media
Persistent homology signatures of the solid phase in synthetic porous media correlate with diffusion-release regimes even after stratifying by target porosity, and classify early-fast, late-release, and long-tail behavior with 0.64–0.76 test accuracy.
-
Do Waders, Swimmers, and Divers Exist? A GPS-Based Pilot Study of Site-Dependent Visitor Movement in Theme Parks
GPS tracking across theme parks shows visitor movement forms a continuum rather than discrete types, diverges from self-reports, and reverses feature relationships from site to site, requiring local calibration.
-
A Dual Edge Spatial Jacobian Image Graph for Interpretable Diabetic Retinopathy Grading
A dual-edge graph fuses vessel-lesion geometry and embedding-biomarker sensitivity from four aligned streams to produce interpretable DR grades on APTOS images with 0.8076 accuracy.
-
Continuous Hidden Markov Models for Equity Returns: Heavy-Tail Emission Families and Regime-Conditional Value-at-Risk
Heavy-tailed continuous HMMs recover volatility clustering and produce regime-conditional VaR that passes joint conditional coverage tests on US equity data.
-
Deep Slice Interpolation for Reducing Through-Plane Anisotropy and Noise in Head CT
Deep learning system synthesizes intermediate head CT slices to halve through-plane anisotropy while providing implicit denoising, outperforming baselines on structural metrics.
-
Density Evolution: A Multiscale View of Density Estimation
A review reframing density estimation as 'density evolution' across scales, linking kernel smoothing to heat flow, mixtures to compression, and topology to level sets, while stating three structural results on modes, Gaussian semigroups, and log-concavity.
-
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
LLMs reach moderate accuracy on a new psychiatric interview benchmark but systematically discount explicit symptoms when preserved functioning or protective factors are present.
-
Guiding Multi-Objective Genetic Programming with Description Length Improves Symbolic Regression Solutions
Post-selection with DL or FBF after multi-objective GP search improves test-set performance over AIC/BIC baselines on noisy synthetic and real regression tasks, while using DL directly as fitness often causes premature convergence to overly simple models.
-
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
A kernel-based regularized learning framework for FDR control that unifies arbitrary structures and supplies provably valid decision rules with likelihood-based tuning.
-
A Method for Characterizing Disease Progression from Acute Kidney Injury to Chronic Kidney Disease
Clustering of EHR-derived patient vectors identifies 15 post-AKI states with varying CKD transition probabilities, enabling identification of state-specific risk factors.
-
Learning Nonlinear Dynamics: Improving the Estimation Efficiency and Reliability of Gaussian Process State-Space Models
Modifies Gibbs sampler for GP state-space models, introduces CFA measurement structure, and validates software via simulation-based calibration to enable reliable learning of nonlinear latent dynamics.
-
Community detection in small-sample ordinal regimes: A benchmarking framework for Delphi data
A simulation benchmark shows community detection on correlation graphs can extract thematic structure from high-dimensional ordinal Delphi data where traditional factor models fail due to rank deficiency.
-
Behavioral and Performance Indicators of Depression and Anxiety in Electronic Learning Systems
Observational study links specific Moodle usage behaviors to higher depression and anxiety scores in 97 computer engineering students.
-
AI and physics-based weather forecasting: A comparative study
Raw IFS forecasts outperform raw AIFS for wind speed at all horizons, but post-processing with EMOS or QR reduces the gap, leaving IFS ahead mainly at short leads.
-
Prototyping and Evaluating a Real-time Neuro-Adaptive Virtual Reality Flight Training System
No performance difference was found between neuro-adaptive and fixed-difficulty VR flight training, yet pilots preferred the adaptive version after briefing.
-
Mono- and Polyauxic Growth Kinetics: A Semi-Mechanistic Framework for Complex Biological Dynamics
Reformulates Boltzmann and Gompertz equations into semi-mechanistic forms for mono- and polyauixic microbial growth as weighted sums of constrained sigmoidal phases, paired with a two-stage optimization workflow and information-criterion model selection, evaluated on anaerobic digestion datasets.
-
Statistical significance in choice modelling: computation, usage and reporting
A commentary paper that reviews common misuses of statistical significance tests in choice modelling and advocates for greater attention to behavioural and policy relevance alongside proper uncertainty reporting.
- Visualizing Local Maxima of the Ohio overdose epidemic with Vineyards
- Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform