Prover-verifier deliberation yields a high-confidence subset of LLM answers with ~30pp higher precision than the complement on GPQA Diamond by using defender-challenger dialogues.
hub
IEEE Transactions on Information Theory , author=
15 Pith papers cite this work, alongside 872 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
Presents first online L2D algorithm for multiclass classification with bandit feedback and varying experts, achieving O((n+n_e)T^{2/3}) regret generally and O((n+n_e)√T) under low noise.
Fuzzy ARTMAP models are highly vulnerable to a new white-box attack aligned with their category competition, but progressive selective training yields stronger replay-free robustness than offline adversarial training under adaptive evaluation.
PLACE delivers a closed-form certified classification method for point clouds and graphs based on persistent homology with explicit excess-risk bounds, selection rules, and training-time certificates.
Multimodal contrastive learning using multilinear products is fragile to single bad modalities, and a gated version improves top-1 retrieval accuracy on synthetic and real trimodal data.
A decoupled softmax-plus-per-expert-sigmoid surrogate for multi-expert learning-to-defer avoids augmented-action pathologies and yields an excess-risk bound independent of expert-pool growth under fixed per-expert weight.
Heat-kernel smoothing over weighted points on a compact manifold yields a scale-dependent geometric effective sample size that discounts nearby and duplicate particles.
Post-hoc learning to defer is cast as density-ratio learning between model and expert ideal distributions, producing DR CPE losses that recover Chow's rule for KL-based ideals and support adjustable deferral via thresholding.
JTS trains reasoning models via supervised warm-up and missing-premise RL to make an explicit answerability commitment that triggers early termination on unanswerable inputs, raising Abstention@Detection near saturation.
On binary verdicts, Pearson, Spearman, Kendall's tau-b, phi, and the Matthews correlation are a single statistic, so most multi-metric agreement reports repeat one number under different names.
Inference-time distillation combines dynamic in-context learning from teacher demonstrations with self-consistency cascades to cut LLM agent costs 2.5-3.5x while recovering most accuracy, without training or manual prompts.
A residual-adequacy architecture unifies interpretation, learning, and empathy as one constraint through accountable abstention driven by representational mismatch.
Object-oriented RGB pixel distribution analysis from satellite images classifies paved versus unpaved roads in Greater Maputo.
citing papers explorer
-
Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction
Prover-verifier deliberation yields a high-confidence subset of LLM answers with ~30pp higher precision than the complement on GPQA Diamond by using defender-challenger dialogues.
-
Online Learning-to-Defer with Varying Experts
Presents first online L2D algorithm for multiclass classification with bandit feedback and varying experts, achieving O((n+n_e)T^{2/3}) regret generally and O((n+n_e)√T) under low noise.
-
Streaming Adversarial Robustness in Fuzzy ARTMAP: Mechanism-Aligned Evaluation, Progressive Training, and Interpretable Diagnostics
Fuzzy ARTMAP models are highly vulnerable to a new white-box attack aligned with their category competition, but progressive selective training yields stronger replay-free robustness than offline adversarial training under adaptive evaluation.
-
A Closed-Form Persistence-Landmark Pipeline for Certified Point-Cloud and Graph Classification
PLACE delivers a closed-form certified classification method for point clouds and graphs based on persistent homology with explicit excess-risk bounds, selection rules, and training-time certificates.
-
Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning
Multimodal contrastive learning using multilinear products is fragile to single bad modalities, and a gated version improves top-1 retrieval accuracy on synthetic and real trimodal data.
-
Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer
A decoupled softmax-plus-per-expert-sigmoid surrogate for multi-expert learning-to-defer avoids augmented-action pathologies and yields an excess-risk bound independent of expert-pool growth under fixed per-expert weight.
-
Heat-Kernel Entropy Profiles and Geometric Effective Sample Size for Weighted Measures on Manifolds
Heat-kernel smoothing over weighted points on a compact manifold yields a scale-dependent geometric effective sample size that discounts nearby and duplicate particles.
-
Density-Ratio Losses for Post-Hoc Learning to Defer
Post-hoc learning to defer is cast as density-ratio learning between model and expert ideal distributions, producing DR CPE losses that recover Chow's rule for KL-based ideals and support adjustable deferral via thresholding.
-
Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information
JTS trains reasoning models via supervised warm-up and missing-premise RL to make an explicit answerability commitment that triggers early termination on unanswerable inputs, raising Abstention@Detection near saturation.
-
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
On binary verdicts, Pearson, Spearman, Kendall's tau-b, phi, and the Matthews correlation are a single statistic, so most multi-metric agreement reports repeat one number under different names.
-
Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
Inference-time distillation combines dynamic in-context learning from teacher demonstrations with self-consistency cascades to cut LLM agent costs 2.5-3.5x while recovering most accuracy, without training or manual prompts.
-
Interpretation, Learning, and Empathy as One Constraint: A Residual-Adequacy Architecture with Accountable Abstention
A residual-adequacy architecture unifies interpretation, learning, and empathy as one constraint through accountable abstention driven by representational mismatch.
-
Monitoring road infrastructures from satellite images in Greater Maputo
Object-oriented RGB pixel distribution analysis from satellite images classifies paved versus unpaved roads in Greater Maputo.
- MCMit: Hardware-Software Co-Design for Mid-Circuit Measurement Error Mitigation
- Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees