LoID pulls informative priors directly from LLM token predictions instead of generated text, recovering up to 59% of the oracle performance gap on 10 OOD tabular datasets.
Harnessing large language models as post-hoc correctors
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
BLUEmed combines hybrid RAG and multi-agent debate to detect clinical terminology substitution errors, reporting 69.13% accuracy and 74.45% ROC-AUC under few-shot prompting.
citing papers explorer
-
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
LoID pulls informative priors directly from LLM token predictions instead of generated text, recovering up to 59% of the oracle performance gap on 10 OOD tabular datasets.
-
BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection
BLUEmed combines hybrid RAG and multi-agent debate to detect clinical terminology substitution errors, reporting 69.13% accuracy and 74.45% ROC-AUC under few-shot prompting.