REVIEW 3 major objections 5 minor 3 cited by
Uncertainty Quantification for Retrieval-Augmented Reasoning
T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This paper introduces R2C, a consistency-based uncertainty score that improves AUROC by over 5% on average across five retrieval-augmented reasoning systems and three QA datasets.
desk verdict Useful empirical UQ for RAR with real AUROC gains; 'theoretically grounded' is overreach and the perturbation bias needs systematic analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the perturbation–consistency loop: R2C formalizes RAR as an MDP, then temporarily swaps the action set for one of three stochastic perturbation actions at a randomly chosen state. Query paraphrasing rewrites the search query; critical rethinking forces the model to reject previously retrieved information and issue a new query; answer validation summarizes the retrieved evidence and checks the final answer for groundedness and completeness. The loop matters because each perturbation changes the retriever's input, which changes the documents, which changes the generator's next input — so uncertainty flows through both components. The score itself is the simple majority-vo
What would settle it
Run R2C on a dataset of correct most-likely answers and measure the flip rate of the critical-rethinking action alone. If a large fraction of correct answers are flipped to incorrect ones, the consistency score systematically over-estimates uncertainty on correct answers and AUROC would drop on that subset — exactly the signature of Appendix C's failed case generalized.
Extended reading notes
Core claim
R2C models a retrieval-augmented reasoning system as a Markov decision process. The most-likely answer is obtained by rolling out the system once. To estimate uncertainty, R2C samples B alternative rolls: at a random step it injects one of three actions — query paraphrasing, critical rethinking, or answer validation — and then continues the original system's loop. The uncertainty score is one minus the fraction of the B rolls that return the same answer as the most-likely roll. The paper reports that, with only three rolls, this score reaches an AUROC of roughly 77–82% depending on the dataset, outperforming all baselines, and that it also produces better abstention decisions (about 5% gain
Load-bearing premise
The method assumes that the three perturbation actions — paraphrasing, critical rethinking, and answer validation — produce a representative sample of plausible reasoning outcomes, so that disagreement with the most-likely answer tracks error; the paper's own reported failed case shows critical rethinking can derail a correct path and inflate uncertainty.
Editorial extensions
If this is right
- Future UQ methods for RAR systems should operate at the level of the reasoning path rather than token probabilities; path-level consistency is the signal that matches correctness.
- R2C scores can be used directly as an abstention rule: withholding answers above a threshold improves both abstention accuracy and F1 by about 5%.
- The scores also provide a model-selection signal: choosing the answer with lowest uncertainty among candidate systems improves exact match by ~7% over single models and ~3% over existing selection methods.
- R2C reaches baseline-level AUROC with about three generations and 700 tokens, roughly 2.5 times fewer than ten-generation baselines, which makes deployment more feasible.
- The observed increase in query and document diversity suggests that measuring uncertainty requires deliberately covering a diverse set of retrieval and reasoning outcomes, not just sampling tokens.
Reading between the lines
- Because R2C's score is built from a small number of perturbed rolls, one could train an amortized predictor that estimates the score from a single path, cutting inference cost further.
- The MDP perturbation recipe is domain-agnostic: the same consistency-over-perturbations idea could quantify uncertainty in other agentic loops, such as tool use or multi-agent debate, by defining actions that perturb each environment's state.
- The failed case in Appendix C suggests that action selection — which perturbation to apply at which step — is itself tunable; an adaptive scheduler could push AUROC higher than uniform random action/state choice.
- R2C's success at model selection hints that uncertainty scores like this could serve as a reward signal for training RAR systems to be explicitly uncertainty-aware, rather than using them only post hoc.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes R2C, a consistency-based uncertainty quantification method for retrieval-augmented reasoning (RAR) systems. R2C models RAR as an MDP and perturbs the most-likely reasoning path with three actions—query paraphrasing, critical rethinking, and answer validation—then computes the uncertainty score as one minus the majority-vote agreement of B=10 sampled responses (Eqs. 1–2). Experiments on five RAR systems and three QA datasets show a mean AUROC gain of over 5% against strong UQ baselines. Extrinsic evaluations on abstention and model selection report ~5% gains in abstention metrics and ~3–7% gains in exact match over selection methods. The paper also claims a ~2.5x token-efficiency advantage.
Significance. If the reported results hold, R2C offers a practical, black-box-compatible UQ signal for a setting (multi-step retrieval-augmented reasoning) where existing methods are indeed limited. The paper's strengths include a broad evaluation across five RAR systems and three datasets, significance testing on the main AUROC comparisons, and honest reporting of a failed case in Appendix C and of action-set sensitivity in Sec. 6.5. However, the core consistency assumption—that the three perturbation actions are neutral probes of reasoning-path variability—is not systematically validated, and the claimed efficiency advantage is based on a configuration different from the main protocol. The MDP framing is currently a notational device rather than a theoretical grounding.
major comments (3)
- [Sec. 4.2, Eq. (2), Sec. 6.5, Appendix C] The central score in Eq. (2) assumes that the three perturbation actions produce a representative sample of plausible reasoning outcomes. The Critical Rethinking prompt (Figure 8) is explicitly adversarial—it instructs the model to 'strongly reject the entire retrieved information as unhelpful, irrelevant, or misleading' and to produce a 'fundamentally new search query.' Appendix C shows the consequence: a correct answer ('Adrian Lyne') receives uncertainty 0.6 because QP and CR flipped six of ten sampled generations to an incorrect answer. The paper provides only this anecdote, not per-action flip rates for correct versus incorrect generations. Sec. 6.5 further shows that no single action set dominates (e.g., QP alone beats 'All' on HotpotQA for Search-R1; AV alone beats 'All' for SelfAsk), so the headline result is contingent on a fixed heuristic choice. To support the generality of th
- [Sec. 5.1 and Sec. 6.4] The Introduction claims that R2C requires 'only about 3 generations on average, 2.5 times fewer token generations than the 10 used by the baselines.' However, the main evaluation (Sec. 5.1) fixes B=10 for R2C and all baselines. The '3 generations' figure comes from the efficiency analysis in Sec. 6.4, where R2C at B=3 approximately matches baselines at B=10 in AUROC—it is not the configuration used for the headline 5% improvement. As written, the efficiency claim is ambiguous and potentially misleading. It should be clearly framed as a break-even analysis, or the main result should be reported under the configuration used for the claimed efficiency.
- [Algorithm 1 and Sec. 4.2] Algorithm 1 is not a faithful executable specification of R2C. For the Answer Validation action, lines 5–9 set s_t = s_N and directly apply a*, but Sec. 4.2 describes an additional path-summarization step using a separate model M that produces \hat{s}_N before the AV prompt is applied (Figures 9–10). The algorithm omits this summarization call. This is a reproducibility issue; it also affects the efficiency accounting, since the summarization is an extra LLM call not shown in the token-efficiency analysis.
minor comments (5)
- [Contribution (1)] The phrase 'novel theocratically grounded UQ method' contains a typo ('theocratically' should be 'theoretically'). More substantively, the MDP formulation in Sec. 3 is a notation-level re-description of RAR; no axioms, assumptions, or formal guarantees are derived. I suggest softening the 'theoretically grounded' claim unless formal properties are provided.
- [Sec. 6.5, Figure 6] Figure 6 and the surrounding text appear to show results for only two of the five RAR models (SelfAsk and Search-R1) and two datasets (PopQA and HotpotQA). Please state this explicitly in the caption and text, or extend the analysis to all model–dataset combinations, since the discussion generalizes to the method as a whole.
- [Table 1 header] The column header 'RAG' should be 'RAR' for consistency with the paper's scope. Also, 'Uncer. M.' is undefined; use 'Uncertainty Method' or spell out in the caption.
- [Reference [25]] The citation for P(true) is incomplete ('Saurav Kadavath et al. 2022'). Please provide the full reference (authors, venue, or arXiv ID) as done for other entries.
- [Sec. 6.2] There is a grammatical error: 'while captures the effectiveness...' is missing a subject. Also, the sentence 'This indicates that R2C scores are more reliable in detecting than correct abstentions than correct non-abstentions' is garbled; please rewrite.
Circularity Check
No significant circularity: R2C's uncertainty score is a direct majority-vote definition evaluated on external benchmarks; self-citations are background and not load-bearing.
full rationale
R2C's uncertainty score is defined directly as U(x,r)=1 - (1/B) * sum_b I(r_b = r), i.e., one minus the fraction of perturbed generations that agree with the most-likely response (Eqs. 1-2 and Algorithm 1). This is a constructive definition, not a fitted or reverse-engineered quantity: no parameter is fitted to correctness labels, and the method is evaluated against external ground-truth exact-match correctness on PopQA, HotpotQA, and Musique (Table 1), with abstention and model-selection as extrinsic tasks (Tables 2-3). The MDP formalization and the three perturbation actions (QP, CR, AV) are heuristic components whose validity is tested empirically rather than derived from the target results; the Appendix C failed case and Section 6.5 action-set ablations are acknowledged limitations and ablations, not circular reductions. The paper's self-citations—Search-R1 [24], the axiomatic UQ analysis [48], and other prior author work—are background references or benchmark components, not load-bearing justifications; no uniqueness theorem, fitted input masquerading as a prediction, or ansatz smuggled in via self-citation is present. Therefore, the derivation chain is self-contained and no circular step is identifiable.
Assumptions & free parameters
free parameters (4)
- Number of generations B =
10 for main results; 3 for efficiency claim
- Perturbation action set A* =
{query paraphrasing, critical rethinking, answer validation}
- Abstention threshold tau_abs =
0.9
- Sampling temperature =
T=0.7 for most-likely generation, T=1.0 for perturbed generations
assumptions (3)
- domain assumption Consistency of answers under perturbations correlates with the probability of correctness
- domain assumption Query paraphrasing, critical rethinking, and answer validation generate diverse yet relevant reasoning paths
- domain assumption The same LLM can reliably judge semantic equivalence between two answers
Cite this review
Pith. "Pith review of Uncertainty Quantification for Retrieval-Augmented Reasoning." pith.science (2026). https://pith.science/paper/NRHLX6GM
@misc{pith2026251011483,
author = {Pith},
title = {Pith review of: Uncertainty Quantification for Retrieval-Augmented Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/NRHLX6GM}},
note = {Machine review of arXiv:2510.11483}
}
read the original abstract
Retrieval-augmented reasoning (RAR) is a recent evolution of retrieval-augmented generation (RAG) that employs multiple reasoning steps for retrieval and generation. While effective for some complex queries, RAR remains vulnerable to errors and misleading outputs. Uncertainty quantification (UQ) offers methods to estimate the confidence of systems' outputs. These methods, however, often handle simple queries with no retrieval or single-step retrieval, without properly handling RAR setup. Accurate estimation of UQ for RAR requires accounting for all sources of uncertainty, including those arising from retrieval and generation. In this paper, we account for all these sources and introduce Retrieval-Augmented Reasoning Consistency (R2C)--a novel UQ method for RAR. The core idea of R2C is to perturb the multi-step reasoning process by applying various actions to reasoning steps. These perturbations alter the retriever's input, which shifts its output and consequently modifies the generator's input at the next step. Through this iterative feedback loop, the retriever and generator continuously reshape one another's inputs, enabling us to capture uncertainty arising from both components. Experiments on five popular RAR systems across diverse QA datasets show that R2C improves AUROC by over 5% on average compared to the state-of-the-art UQ baselines. Extrinsic evaluations using R2C as an external signal further confirm its effectiveness for two downstream tasks: in Abstention, it achieves ~5% gains in both F1Abstain and AccAbstain; in Model Selection, it improves the exact match by ~7% over single models and ~3% over selection methods.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 3 Pith papers
-
When Confidence Takes the Wrong Path: Diagnosing Retrieval-State Lock-In in RAG
Retrieval-state lock-in causes zero-dispersion errors in 42% of KG-RAG and 59% of dense-retrieval failures; a three-object check rule reaches 91.9% pooled precision at 7.7% coverage.
-
NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems
A rule-guided self-generated fine-tuning method reduces verbal confidence miscalibration (ECE) in RAG question-answering by roughly 0.1 absolute across four open-weight models.
-
Evaluating RAG Reliability under Clean, Misleading, and Mixed Retrieval
Proposes an evaluation framework using parametric override and confidence metrics to assess RAG robustness to clean, poisoned, and mixed retrieval evidence on factoid questions.
Reference graph
Works this paper leans on
-
[1]
Pierre Achkar, Tim Gollub, and Martin Potthast. 2025. Ask, Retrieve, Summarize: A Modular Pipeline for Scientific Literature Summarization.CoRR(2025)
2025
-
[2]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi
-
[3]
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, and Salman Avestimehr. 2024. MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics ACL. 7752–7767
2024
-
[4]
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Sungmin Kang, Tuo Zhang, Baturalp Buyukates, Salman Avestimehr, and Sai Praneeth Karimireddy. 2025. Reconsider- ing LLM Uncertainty Estimation Methods in the Wild. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 29531–29556
2025
-
[5]
Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen
Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen. 2025. ReSearch: Learning to Reason with Search for LLMs via Reinforce- ment Learning.CoRR(2025)
2025
-
[6]
Yifei Chen, Guanting Dong, Yutao Zhu, and Zhicheng Dou. 2025. Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration.arXiv preprint arXiv:2508.13828(2025)
arXiv 2025
-
[7]
Zhijun Chen, Jingzheng Li, Pengpeng Chen, Zhuoran Li, Kai Sun, Yuankai Luo, Qianren Mao, Dingqi Yang, Hailong Sun, and Philip S. Yu. 2025. Harnessing Mul- tiple Large Language Models: A Survey on LLM Ensemble.CoRRabs/2502.18036 (2025). https://doi.org/10.48550/ARXIV.2502.18036 arXiv:2502.18036
-
[8]
Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian Von Der Weth, and Brian Y. Lim. 2021. Directed diversity: Leveraging language embedding dis- tances for collective creativity in crowd ideation. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–35
2021
Show all 72 references
-
[9]
Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024. The Power of Noise: Redefining Retrieval for RAG Systems. InProceedings of the 47th International ACM SIGIR Conference on...
2024
-
[10]
Elizabeth R DeLong, David M DeLong, and Daniel L Clarke-Pearson. 1988. Com- paring the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach.Biometrics(1988), 837–845
1988
-
[11]
Jinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2024. Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models. InProceedings of the 62nd Annual Meeting of t...
2024
-
[12]
Jinhao Duan, James Diffenderfer, Sandeep Madireddy, Tianlong Chen, Bhavya Kailkhura, and Kaidi Xu. 2025. UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making.CoRRabs/2506.17419 (2025)
2025 arXiv
-
[13]
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024. Detecting hallucinations in large language models using semantic entropy.Nat.630, 8017 (2024), 625–630
2024
-
[14]
Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Vidhisha Balachandran, and Yulia Tsvetkov. 2024. Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistic...
2024
-
[15]
Chen, Trevor Chow, Ishan S
Neel Guha, Mayee F. Chen, Trevor Chow, Ishan S. Khare, and Christopher Ré
-
[16]
James Harrison, John Willes, and Jasper Snoek. 2024. Variational Bayesian Last Layers. InThe Twelfth International Conference on Learning Representations, ICLR. OpenReview.net
2024
-
[17]
InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS
Smoothie: Label Free Language Model Routing. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS
2024
-
[18]
Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, and Yang Zhang. 2024. Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2...
2024
-
[19]
Bolei He, Nuo Chen, Xinran He, Lingyong Yan, Zhenkai Wei, Jinchang Luo, and Zhen-Hua Ling. 2024. Retrieving, Rethinking and Revising: The Chain-of- Verification Can Improve Retrieval Augmented Generation. InFindings of the Association for Computational Linguistics: EMNLP. Asso...
2024
-
[20]
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023. LLM-Blender: Ensem- bling Large Language Models with Pairwise Ranking and Generative Fusion. In Proceedings of the 61st Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), ACL. Associati...
2023
-
[21]
Asib Rahman, K
Shayekh Bin Islam, Md. Asib Rahman, K. S. M. Tozammel Hossain, Enamul Hoque, Shafiq Joty, and Md. Rizwan Parvez. 2024. Open-RAG: Enhanced Retrieval Aug- mented Reasoning with Open-Source Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP....
2024
-
[22]
Pengcheng Jiang, Xueqiang Xu, Jiacheng Lin, Jinfeng Xiao, Zifeng Wang, Jimeng Sun, and Jiawei Han. 2025. s3: You Don’t Need That Much Data to Train a Search Agent via RL.arXiv preprint arXiv:2505.14146(2025)
2025
-
[23]
Jinhao Jiang, Jiayi Chen, Junyi Li, Ruiyang Ren, Shijie Wang, Xin Zhao, Yang Song, and Tao Zhang. 2025. RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement. InProceedings of the 2025 Conference of the Nations of the Americas Chapter ...
2025
-
[24]
Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.CoRR(2025)
2025
-
[25]
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active Retrieval Aug- mented Generation. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7969–7992
2023
-
[26]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP....
2020
-
[27]
Saurav Kadavath et al. 2022. Language Models (Mostly) Know What They Know. abs/2207.05221 (2022)
2022 arXiv
-
[28]
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic Uncertainty: Lin- guistic Invariances for Uncertainty Estimation in Natural Language Generation. InThe Eleventh International Conference on Learning Representations ICLR
2023
-
[29]
Anton Korikov, Pan Du, Scott Sanner, and Navid Rekabsaz. 2025. Batched Self-Consistency Improves LLM Relevance Assessment and Ranking.CoRR abs/2505.12570 (2025)
2025
-
[30]
Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou. 2025. Search-o1: Agentic Search-Enhanced Large Reasoning Models.CoRR(2025)
2025
-
[31]
Baixuan Li, Yunlong Fan, Tianyi Ma, Miao Gao, Chuanqi Shi, and Zhiqiang Gao. 2025. RASPberry: Retrieval-Augmented Monte Carlo Tree Self-Play with Reasoning Consistency for Multi-Hop Question Answering. InFindings of the Association for Computational Linguistics: ACL 2025. 11258–11276
2025
-
[32]
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query Rewriting in Retrieval-Augmented Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 5303–5315
2023
-
[33]
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2024. Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models.Trans. Mach. Learn. Res.2024 (2024)
2024
-
[34]
Andrey Malinin and Mark J. F. Gales. 2021. Uncertainty Estimation in Au- toregressive Structured Prediction. In9th International Conference on Learning Representations, ICLR
2021
-
[35]
Nishanth Madhusudhan, Sathwik Tejaswi Madhusudhan, Vikas Yadav, and Ma- soud Hashemi. 2025. Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models. InProceedings of the 31st In- ternational Conference on Computational Linguistics, COLING. ...
2025
-
[36]
Viktor Moskvoretskii, Maria Marina, Mikhail Salnikov, Nikolay Ivanov, Sergey Pletenev, Daria Galimzianova, Nikita Krayko, Vasily Konovalov, Irina Nikishina, and Alexander Panchenko. 2025. Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home. InProceedings ...
2025
-
[37]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. InProceedings of the Conference acronym ’XX, June 03–05, 2018, Woodsto...
2023
-
[38]
Laura Perez-Beltrachini and Mirella Lapata. 2025. Uncertainty Quantification in Retrieval Augmented Question Answering.arXiv preprint arXiv:2502.18108 (2025)
2025 arXiv
-
[39]
Malik Sajjad Ahmed Nadeem, Jean-Daniel Zucker, and Blaise Hanczar. 2010. Accuracy-Rejection Curves (ARCs) for Comparing Classification Methods with a Reject Option. InProceedings of the third International Workshop on Machine Learning in Systems Biology, MLSB (JMLR Proceedings...
2010
-
[40]
Zhenting Qi, Mingyuan MA, Jiahang Xu, Li Lyna Zhang, Fan Yang, and Mao Yang. 2025. Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solver. In The Thirteenth International Conference on Learning Representations
2025
-
[41]
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. Measuring and Narrowing the Compositionality Gap in Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2023. 5687–5711
2023
-
[42]
Robertson and Steve Walker
Stephen E. Robertson and Steve Walker. 1994. Some Simple Effective Approxima- tions to the 2-Poisson Model for Probabilistic Weighted Retrieval. InProceedings of the 17th Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval. 232–241
1994
-
[43]
I Can’t Believe It’s Not Better: Failure Modes in the Age of Foundation Models
Jie Ren, Yao Zhao, Tu Vu, Peter J. Liu, and Balaji Lakshminarayanan. 2023. Self- Evaluation Improves Selective Generation in Large Language Models. InPro- ceedings on "I Can’t Believe It’s Not Better: Failure Modes in the Age of Foundation Models" at NeurIPS 2023 Workshops (Pr...
2023
-
[44]
Alireza Salemi and Hamed Zamani. 2024. Evaluating Retrieval Quality in Retrieval-Augmented Generation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR. ACM, 2395–2400
2024
- [45]
-
[46]
Heydar Soudani. 2025. Enhancing Knowledge Injection in Large Language Models for Efficient and Trustworthy Responses. InProceedings of the 48th International ACM Conference on Research and Development in Information Retrieval, SIGIR. 4211
2025
-
[47]
Yaorui Shi, Sihang Li, Chang Wu, Zhiyuan Liu, Junfeng Fang, Hengxing Cai, An Zhang, and Xiang Wang. 2025. Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs.arXiv preprint arXiv:2505.11277(2025)
2025
-
[48]
Heydar Soudani, Evangelos Kanoulas, and Faegheh Hasibi. 2025. Why Uncer- tainty Estimation Methods Fall Short in RAG: An Axiomatic Analysis. InFindings of the Association for Computational Linguistics: ACL 2025. 16596–16616
2025
-
[49]
Heydar Soudani, Evangelos Kanoulas, and Faegheh Hasibi. 2024. Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge. InProceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific...
2024
-
[50]
Heydar Soudani, Roxana Petcu, Evangelos Kanoulas, and Faegheh Hasibi. 2024. A Survey on Recent Advances in Conversational Data Generation.CoRR abs/2405.13003 (2024)
2024 arXiv
-
[51]
Heydar Soudani, Roxana Petcu, Evangelos Kanoulas, and Faegheh Hasibi. 2024. Data Augmentation for Conversational AI. InCompanion Proceedings of the ACM on Web Conference 2024, WWW. 1234–1237
2024
-
[52]
Zhongxiang Sun, Qipeng Wang, Weijie Yu, Xiaoxue Zang, Kai Zheng, Jun Xu, Xiao Zhang, Yang Song, and Han Li. 2025. ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding. InProceedings of the 48th International ACM SIGIR Conference on Research and Development...
2025
-
[53]
Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
2024
-
[54]
Hieu Tran, Zonghai Yao, Zhichao Yang, Junda Wang, Yifan Zhang, Shuo Han, Feiyun Ouyang, and Hong Yu. 2025. RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volu...
2025
-
[55]
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. 2023. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine- Tuned with Human Feedback. InProcee...
2023
-
[56]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal
-
[58]
Le, Ed H
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InThe Eleventh International Conference on Learning Representations, ICLR
2023
-
[59]
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. The Art of Abstention: Selective Prediction and Error Regularization for Natural Language Processing. InProceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International ...
2021
-
[60]
Duygu Nur Yaldiz, Yavuz Faruk Bakman, Baturalp Buyukates, Chenyang Tao, Anil Ramakrishna, Dimitrios Dimitriadis, Jieyu Zhao, and Salman Avestimehr
-
[61]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Lan- guage Models are not Fair Evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (...
2024
-
[62]
An Yang et al. 2024. Qwen2.5 Technical Report.CoRRabs/2412.15115 (2024)
2024 arXiv
-
[63]
Cohen, Rus- lan Salakhutdinov, and Christopher D
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Rus- lan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. InProceedings of the 2018 Conference on Empirical Methods in Natural Lang...
2018
-
[64]
Narasimhan, and Yuan Cao
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InThe Eleventh International Conference on Learning Representations, ICLR
2023
-
[65]
Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu. 2024. Benchmarking LLMs via Uncertainty Quantification. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing ...
2024
-
[66]
Duygu Nur Yaldiz, Yavuz Faruk Bakman, Sungmin Kang, Alperen Öziş, Hayret- tin Eren Yildiz, Mitash Ashish Shah, Zhiqi Huang, Anoop Kumar, Alfy Samuel, Daben Liu, et al. 2025. TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs.arXiv preprint arXiv:2...
2025 arXiv
-
[67]
Qiwei Zhao, Dong Li, Yanchi Liu, Wei Cheng, Yiyou Sun, Mika Oishi, Takao Osaki, Katsushi Matsuda, Huaxiu Yao, Chen Zhao, Haifeng Chen, and Xujiang Zhao. 2025. Uncertainty Propagation on LLM Agent. InProceedings of the 63rd Annual Meeting of the Association for Computational Li...
2025
-
[71]
Tianhui Zhang, Bei Peng, and Danushka Bollegala. 2025. Evaluating the Evalua- tion of Diversity in Commonsense Generation. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 24258–24275
2025
-
[73]
Thebarton Oval
Is the response grounded in the provided information? 2) Does the response correctly and fully answer the user query? You must provide a single, coherent reasoning step that examines both criteria and suggests how the response could be improved. After your reasoning, you must ...
2018
-
[2022]
MuSiQue: Multihop Questions via Single-hop Question Composition.Trans. Assoc. Comput. Linguistics10 (2022), 539–554
2022
-
[2023]
InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge- Intensive Multi-Step Questions. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10014–10037
-
[2024]
InThe Twelfth International Conference on Learning Representations, ICLR
Self-RAG: Learning to Retrieve, Generate, and Critique through Self- Reflection. InThe Twelfth International Conference on Learning Representations, ICLR. OpenReview.net
-
[2025]
InFindings of the Association for Computational Linguistics: NAACL 2025
Do Not Design, Learn: A Trainable Scoring Function for Uncertainty Estimation in Generative LLMs. InFindings of the Association for Computational Linguistics: NAACL 2025
2025
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.