REVIEW 26 cited by
Conformal Risk Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Conformal Risk Control
read the original abstract
We extend conformal prediction to control the expected value of any monotone loss function. The algorithm generalizes split conformal prediction together with its coverage guarantee. Like conformal prediction, the conformal risk control procedure is tight up to an $\mathcal{O}(1/n)$ factor. We also introduce extensions of the idea to distribution shift, quantile risk control, multiple and adversarial risk control, and expectations of U-statistics. Worked examples from computer vision and natural language processing demonstrate the usage of our algorithm to bound the false negative rate, graph distance, and token-level F1-score.
Forward citations
Cited by 26 Pith papers
-
Conformal Risk-Averse Decision Making with Action Conditional Guarantee
Action-conditional conformal prediction sets provide per-action safety guarantees for risk-averse policies that optimize conditional value-at-risk through pinball-loss minimization.
-
BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies
BOKBO is the first conformal abstention method for K-sample VLA policies that supplies finite-sample distribution-free guarantees on executed violation rates, with global and Mondrian per-task variants.
-
Post-Selection Distributional Model Evaluation
PS-DME is a new framework that controls post-selection false coverage rate for distributional KPI estimates via e-values and is provably more sample-efficient than data splitting under explicit conditions.
-
Decomposition-Based Modular Conformal Prediction for Two-Stage Modeling
A decomposition-based modular conformal prediction method for two-stage models with FWER-controlled stage-wise scaling and adaptive extension for non-stationary data.
-
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
On hard multi-table text-to-SQL, verification-based LLM judges beat self-consistency and log-probability for predicting execution correctness, and fine-tuned verifiers fail to transfer across schemas.
-
Reliable Conformal Prediction for Ordinal Classification Using the Ranked Probability Score
RPS-based conformal prediction for ordinal classification yields median-centered contiguous sets with a favorable width-miscoverage tradeoff compared to prior methods.
-
The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions
Comparative evaluation of seven confidence constructions across 25 LLM-dataset pairs reveals that verbalized scores provide good ranking but coarse granularity for thresholding, while multi-query aggregation helps wea...
-
Density Ridge Selective Prediction for LLM and VLM Hallucination Detection under Calibration Label Scarcity
Density ridge scoring on 6D kinematic features from hidden states yields 5-20 point AUROC gains over Semantic Entropy and log-probability baselines for hallucination detection under a 200-query calibration protocol ac...
-
A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control
A joint finite-sample certificate for adaptive selective conformal risk control that treats selected risk as a ratio and couples empirical-Bernstein, Clopper-Pearson, and closeness bounds.
-
Proof-Carrying Agent Actions: Model-Agnostic Runtime Governance for Heterogeneous Agent Systems
PCAA introduces a runtime-neutral governance model for heterogeneous agent systems based on action certificates with five checkpoints, externality awareness, and enforceability classes.
-
CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control
CALIBURN integrates Bayesian change-point detection, isotonic calibration, cost-sensitive thresholding, conformal risk control, and burn-rate alerting into a single streaming substrate, showing that calibration and CR...
-
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
ARC-STAR reduces velocity rollout error by at least 36x over raw Poseidon across all tested regime cells via auditable global and local correction stages on five flow benchmarks.
-
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
ARC-STAR is an auditable, budget-aware post-hoc correction method that reduces velocity rollout error by at least 36x over raw Poseidon across five flow benchmarks.
-
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
ARC-STAR is a frozen, auditable post-hoc correction method that reduces velocity rollout error by at least 36x over raw Poseidon across five flow benchmarks using global and local stages with budget-aware triage.
-
SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction
SAVER proposes a conformal groundability gate plus submodular image selector that activates vision only when needed for multimodal named entity recognition and relation extraction, improving F1 while lowering compute.
-
Uncertainty Quantification for LLM-based Code Generation
RisCoSet applies multiple hypothesis testing to construct risk-controlling partial-program prediction sets for LLM code generation, achieving up to 24.5% less code removal than prior methods at equivalent risk levels.
-
Robust Conditional Conformal Prediction via Branched Normalizing Flow
Branched Normalizing Flow improves conditional coverage robustness of conformal prediction under distribution shift by normalizing test inputs to the calibration distribution and mapping prediction sets back.
-
Geometry-Calibrated Conformal Abstention for Language Models
Geometry-calibrated conformal abstention lets language models abstain from uncertain queries with finite-sample guarantees on both participation rate and conditional correctness of answers.
-
On a Probability Inequality for Order Statistics with Applications to Bootstrap, Conformal Prediction, and more
An approximate inequality for the probability involving order statistics under near-i.i.d. conditions is established and applied to justify resampling-based statistical procedures.
-
Hybrid Decision Making via Conformal VLM-generated Guidance
ConfGuide uses conformal risk control to generate targeted guidance sets in a learning-to-guide hybrid decision framework and demonstrates it on multi-label medical diagnosis.
-
Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking
SBBT separates Brier-score calibration gains from AUROC ranking gains in prefix-conditioned success estimation for LLM math reasoning, with structure-aware signals yielding up to +0.110 AUROC over baselines.
-
Explainable Wastewater Digital Twins: Adaptive Context-Conditioned Structured Simulators with Self-Falsifying Decision Support
CCSS-IX is a context-conditioned structured simulator for wastewater digital twins that uses adaptive expert mixing and self-falsifying conformal decision rules to reduce unsafe actions while maintaining low predictio...
-
A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
Pith review generated a malformed one-line summary.
-
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
Foundation model representations from images and transcriptomics carry complementary signals for cancer classification; multimodal fusion improves results mainly when no modality dominates, and conformal prediction re...
-
Uncertainty Quantification on Graph Learning: A Survey
A survey that categorizes uncertainty quantification approaches for graphical models into representation and handling dimensions to identify challenges and opportunities.
-
Online Safety Monitoring for LLMs
Simple thresholding on an external verifier signal, calibrated by risk control, performs competitively with sequential hypothesis testing monitors on math reasoning and red-teaming datasets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.