A DETR-style probe distills multi-sample claim uncertainty into single-pass span detection and continuous Mixture-of-Beta scores, outperforming baselines on a new 293K-span benchmark.
Truthfulqa: Measuring how models mimic human falsehoods
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
Empirical evaluation of three LLMs finds prevalent overconfidence in insecure code generation, with security calibration outperforming functional calibration but both degrading in repository-level settings.
Case study of CMBAgent on 18 astrophysical tasks finds strong performance on well-specified problems but frequent silent failures yielding physically inconsistent outputs.
Multi-agent AI agents answer questions alone then exchange reasoning to revise decisions, tested via experiments for net reliability gains versus error propagation.
citing papers explorer
-
SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation
A DETR-style probe distills multi-sample claim uncertainty into single-pass span detection and continuous Mixture-of-Beta scores, outperforming baselines on a new 293K-span benchmark.
-
An Empirical Study of Security Calibration in Large Language Models for Code
Empirical evaluation of three LLMs finds prevalent overconfidence in insecure code generation, with security calibration outperforming functional calibration but both degrading in repository-level settings.
-
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows
Case study of CMBAgent on 18 astrophysical tasks finds strong performance on well-specified problems but frequent silent failures yielding physically inconsistent outputs.
-
Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring
Multi-agent AI agents answer questions alone then exchange reasoning to revise decisions, tested via experiments for net reliability gains versus error propagation.