Coding agents struggle to infer least-privilege file permissions by omitting needed accesses while granting unused or sensitive ones, but Sufficiency-Tightness Decomposition improves sensitive-task success by up to 15.8% and reduces attacks.
Knowrl: Teaching language models to know what they know
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2representative citing papers
Linear probes trained on pre-solution hidden states, supervised by post-solution correctness probe outputs, recover 32–66% of the calibration gap between pre- and post-solution confidence across five open-source LLMs.
citing papers explorer
-
Do Coding Agents Understand Least-Privilege Authorization?
Coding agents struggle to infer least-privilege file permissions by omitting needed accesses while granting unused or sensitive ones, but Sufficiency-Tightness Decomposition improves sensitive-task success by up to 15.8% and reduces attacks.
-
Future Confidence Distillation in Large Language Models
Linear probes trained on pre-solution hidden states, supervised by post-solution correctness probe outputs, recover 32–66% of the calibration gap between pre- and post-solution confidence across five open-source LLMs.