LLMs suppress causal caution in practical advisory contexts (rates drop from 91.7-100% to 6.7-18.3%) but recover it with a self-correction prompt (to 71.4-100%).
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6verdicts
UNVERDICTED 6representative citing papers
Mechanical enforcement of governance rules in LLM-based financial decision systems reduces non-compliant deferrals by 73% and raises task accuracy from MCC 0.43 to 0.88, revealing that governance and task performance are distinct axes.
SABER combines self-prior with multi-trace PK and CK reasoning representations to estimate reliability beliefs and drive trust-or-abstain decisions in knowledge-conflict RAG, improving accuracy over baselines.
MedVIGIL provides a 300-case evaluation suite with 2556 probes that measures silent failures in medical VLMs under broken evidence, showing the best model at 69.2 on the composite score versus a human radiologist at 83.3.
A multi-strategy interrogation method with auxiliary expert assessment reduces expected calibration error by 40% on average across three medical VQA datasets for MLLMs.
EARS fine-tunes sub-agents on LLM-as-Judge labeled data to output explanatory abstentions for failure modes, raising response pass rate from 68.5% to 78.9% in a production e-commerce MAS.
citing papers explorer
-
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
LLMs suppress causal caution in practical advisory contexts (rates drop from 91.7-100% to 6.7-18.3%) but recover it with a self-correction prompt (to 71.4-100%).
-
Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems
Mechanical enforcement of governance rules in LLM-based financial decision systems reduces non-compliant deferrals by 73% and raises task accuracy from MCC 0.43 to 0.88, revealing that governance and task performance are distinct axes.
-
Trust or Abstain? A Self-Aware RAG Approach
SABER combines self-prior with multi-trace PK and CK reasoning representations to estimate reliability beliefs and drive trust-or-abstain decisions in knowledge-conflict RAG, improving accuracy over baselines.
-
MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence
MedVIGIL provides a 300-case evaluation suite with 2556 probes that measures silent failures in medical VLMs under broken evidence, showing the best model at 69.2 on the composite score versus a human radiologist at 83.3.
-
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA
A multi-strategy interrogation method with auxiliary expert assessment reduces expected calibration error by 40% on average across three medical VQA datasets for MLLMs.
-
EARS: Explanatory Abstention for Reliable Sub-Agent Modeling in Large-scale Multi-Agent Systems
EARS fine-tunes sub-agents on LLM-as-Judge labeled data to output explanatory abstentions for failure modes, raising response pass rate from 68.5% to 78.9% in a production e-commerce MAS.