Structural uncertainty from self-preference-induced rankings of LLM reasoning paths complements answer dispersion for identifying unreliable instances on logical tasks while collapsing on factual retrieval.
Rank analysis of incomplete block designs: I. the method of paired comparisons,
7 Pith papers cite this work, alongside 1,372 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7roles
method 1polarities
use method 1representative citing papers
The only aggregation rule satisfying same-scale normalization, recursive consistency, and marginal Elo-strength consistency converts ratings to strengths, takes their weighted arithmetic mean, and converts back.
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
LMs systematically inflate expressed certainty during rewriting, affecting up to 75% of outputs with a 1.5-2x bias toward increasing rather than decreasing certainty, and the effect compounds over iterations.
LearnedCache shows a quantized one-layer perceptron, trained on eBPF kernel traces and deployed through cache_ext, can beat FIFO page-cache eviction on some Filebench workloads.
CivBench trains models on turn-level states in Civilization V to predict victory probabilities, providing a progress-based evaluation of LLM strategic capabilities across 307 games with 7 models.
AURA is an adaptive uncertainty-aware refinement method for auditing LLM-as-a-judge pairwise decisions that learns human-consistency signals through selective human verification on uncertain cases.
citing papers explorer
-
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty
Structural uncertainty from self-preference-induced rankings of LLM reasoning paths complements answer dispersion for identifying unreliable instances on logical tasks while collapsing on factual retrieval.
-
Aggregating Elo Ratings: An Axiomatization
The only aggregation rule satisfying same-scale normalization, recursive consistency, and marginal Elo-strength consistency converts ratings to strengths, takes their weighted arithmetic mean, and converts back.
-
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
-
From `May' to `Is': Certainty Distortion in Language Model Rewriting
LMs systematically inflate expressed certainty during rewriting, affecting up to 75% of outputs with a 1.5-2x bias toward increasing rather than decreasing certainty, and the effect compounds over iterations.
-
LearnedCache: eBPF-Integrated Perceptron-Based Eviction Policies for the Linux Page Cache
LearnedCache shows a quantized one-layer perceptron, trained on eBPF kernel traces and deployed through cache_ext, can beat FIFO page-cache eviction on some Filebench workloads.
-
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
CivBench trains models on turn-level states in Civilization V to predict victory probabilities, providing a progress-based evaluation of LLM strategic capabilities across 307 games with 7 models.
-
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
AURA is an adaptive uncertainty-aware refinement method for auditing LLM-as-a-judge pairwise decisions that learns human-consistency signals through selective human verification on uncertain cases.