Pith. sign in

REVIEW 3 major objections 57 references

Treating retrieved sources as probabilistic evidence, not fixed context, lets a RAG system expose disagreement and abstain instead of hallucinating.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 11:18 UTC pith:WITSMLPU

load-bearing objection Clean systems paper: conflict-preserving DS fusion + routing for multi-source RAG, with real gains on CRAG-ambiguous and honest ablations—evaluator validation is the soft spot, not the math. the 3 major comments →

arxiv 2607.10491 v1 pith:WITSMLPU submitted 2026-07-11 cs.LG

EvidentialRAG: Quantifying and Mitigating Information Conflict in Multi-Source Retrieval-Augmented Generation via Evidential Deep Learning

classification cs.LG
keywords retrieval-augmented generationevidential deep learningknowledge conflictDempster-Shafer theoryhallucinationcalibrationlarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most retrieval-augmented generation systems paste retrieved passages into a prompt as if those passages agree. This paper argues that open-world retrieval often returns conflicting, stale, or incomplete evidence, and that a useful system must measure that disagreement before it answers. EvidentialRAG turns each retrieved chunk into Dirichlet evidence over candidate claims, fuses those evidences with a rule that moves unresolved conflict into epistemic uncertainty rather than normalizing it into false confidence, then routes the generator to answer directly, explain the conflict, or abstain. On conflict-heavy benchmarks it lowers hallucination and raises conflict resolution and calibration while staying competitive on ordinary multi-hop and standard question answering. The claim is that evidential modeling is a practical way to make foundation-model retrieval systems more trustworthy when sources disagree.

Core claim

The paper claims multi-source RAG should be treated as probabilistic evidence aggregation rather than deterministic context concatenation. Mapping chunks to Dirichlet evidence, fusing them with a conflict-preserving Dempster-Shafer operator that transfers unresolved disagreement into frame mass (global uncertainty), and routing generation by that score reduces hallucination on ambiguous evidence from 45.3% for Corrective RAG to a human-calibrated 34.8% on the CRAG ambiguous subset, raises conflict resolution from 35.2% to 51.2%, and improves expected calibration error to 0.122 without collapsing ordinary answer quality.

What carries the argument

Conflict-preserving Dempster-Shafer fusion of Dirichlet evidence: each chunk yields singleton belief masses and frame mass (ignorance); pairwise conflict mass is partially transferred into global uncertainty U_global via a tunable transfer parameter rather than being normalized away; U_global then selects direct answer, conflict-aware answer, or abstention.

Load-bearing premise

The method assumes a smaller prompt-based evaluator can turn each retrieved chunk into reliable evidence scores that truly reflect support, irrelevance, and contradiction.

What would settle it

On a held-out conflict set, if human-labeled instances show that fused global uncertainty fails to separate genuine contradiction from mere irrelevance better than a matched corrective baseline—so that routing no longer cuts hallucination while keeping refusal F1 high—the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • RAG pipelines can surface source disagreement to users instead of silently synthesizing incompatible claims.
  • Answer confidence can be better calibrated by using fused singleton mass rather than generator-only confidence.
  • Abstention and conflict-aware modes can be controlled by two uncertainty thresholds without retraining the main generator.
  • Ordinary multi-hop QA need not be sacrificed if routing thresholds leave coherent evidence in the direct-answer path.
  • Production systems can swap evaluators, retrievers, or generators while keeping the same fusion and routing layer.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same fused uncertainty score could trigger targeted follow-up retrieval or human handoff only when conflict is high, rather than always paying full evaluator cost.
  • Domains with structured identifiers (law, medicine, finance) could replace prompt-based claim alignment with ontology keys and may obtain cleaner conflict signals.
  • Distilling the evaluator into a smaller classifier could make the reliability gains cheap enough for interactive enterprise search.
  • Logging U_global and the chosen route would give an auditable record of when the system chose to hedge, useful for regulated deployments.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper proposes EvidentialRAG, a multi-source RAG pipeline that converts retrieved chunks into Dirichlet evidence via a prompt-based Llama-3-8B evaluator, fuses those masses with a conflict-preserving Dempster–Shafer operator (parameter λ) that transfers unresolved conflict into frame mass rather than normalizing it away, and routes the generator to direct answer, conflict-aware answer, or abstention using global uncertainty U_global. On matched baselines (Naive RAG, Self-RAG, Corrective RAG) with the same retriever/generator, the method is competitive on standard CRAG/ConflictQA/MuSiQue QA while improving conflict behavior: on the CRAG ambiguous subset, human-calibrated hallucination falls from 45.3% to 34.8%, conflict resolution rises from 35.2% to 51.2%, and ECE reaches 0.122. Supporting analyses include λ and threshold sweeps, ablations, efficiency accounting, a 200-response human audit of the LLM judge, and a 100-chunk evaluator validity check.

Significance. If the gains are attributable to the evidential layer rather than evaluator artifacts or routing conservatism, the work offers a practical, modular mechanism for making inter-source conflict a first-class signal in RAG—relevant to expert systems, enterprise search, and high-stakes information access. Strengths include a clear EDL–DST bridge that treats ignorance as frame mass rather than an unknown singleton, matched-system evaluation with three seeds and paired bootstrap tests, ablations and sensitivity tables, a MuSiQue non-conflict sanity check against over-abstention, and an explicit human audit of judge bias. The contribution is systems-level rather than a new theoretical result, but the conflict-preserving fusion and uncertainty routing are concrete and deployable.

major comments (3)
  1. §3.2, §4.3, Table 13: The central claim that fused U_global is a reliable conflict-vs-ignorance routing signal rests on a prompt-based Llama-3-8B evaluator whose validation is limited to 100 held-out chunks (91.4% claim fidelity, Pearson r=0.84). There is no direct measurement of whether the evaluator systematically under-detects mutual exclusivity (treating contradictions as weak support or vacuity) or over-assigns singleton mass on ambiguous/stale CRAG passages. Because iterative fusion (Eqs. 7–13) and fixed thresholds (τ1=0.35, τ2=0.65) amplify correlated evaluator errors, the human-calibrated 34.8%/51.2% gains in Table 4 cannot yet be firmly attributed to the evidential layer. A conflict-labeled evaluator audit or error analysis on CRAG-ambiguous chunks is needed.
  2. §4.4–4.5, Tables 4 and 6: Full-benchmark hallucination and CRR are scored by an LLM judge; the 200-response human audit only supplies fixed offsets (33.1→34.8 hallucination; 52.5→51.2 CRR). The audit is stratified but small relative to the benchmark, and contradiction agreement is the weakest dimension (κ=0.79). The headline conflict results should either be fully human-scored on a larger stratified sample or accompanied by confidence intervals that propagate judge–human disagreement, so that the reported gains are not sensitive to a single calibration offset.
  3. §4.2, Tables 3–4, 7–8: Baselines are run under a shared harness that the authors note may differ from original Self-RAG/Corrective RAG hyperparameters. The paper correctly frames this as a matched-system comparison, but the abstract and §5 still present large relative gains against Corrective RAG as if they were method-level improvements. Either re-tune the published baselines under their recommended settings or more carefully bound claims to the matched configuration so readers do not over-interpret the absolute deltas.

Circularity Check

0 steps flagged

Empirical systems paper on external benchmarks; no derivation reduces reported metrics to fitted inputs or self-citation by construction.

full rationale

EvidentialRAG is an engineering/systems contribution: map chunks to Dirichlet evidence via a prompt evaluator, fuse with a conflict-preserving Dempster–Shafer operator (Eqs. 7–13), and route generation by U_global (Eq. 15). Headline numbers (CRAG-ambiguous hallucination 34.8%, CRR 51.2%, ECE 0.122) are measured against external benchmarks (CRAG, ConflictQA, MuSiQue) and matched baselines under a shared harness, not derived as first-principles predictions. λ=0.6 and (τ1,τ2)=(0.35,0.65) are free hyperparameters chosen on held-out validation before final testing (Sec. 4.3); sensitivity tables (Tables 9–10) show other settings change outcomes, so the reported metrics are not forced by the fit. Confidence for ECE is defined as max fused singleton mass (Eq. 27)—a definition of the signal being calibrated, not a circular prediction of accuracy. Human-audit offsets (Table 6) recalibrate the LLM judge; they do not bake the method’s gains into the labels by construction. No load-bearing self-citation, uniqueness theorem from the authors, or ansatz smuggled via prior author work appears in the derivation chain. Evaluator fidelity (Table 13) is a validity/correctness risk, not circularity. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 3 invented entities

The central empirical claim rests on standard DS/EDL mappings plus several hand-chosen operating parameters and the assumption that a prompt-based evaluator yields usable evidence masses. No new physical entities; the invented pieces are methodological constructs (fusion operator, routing policy, evidence interface).

free parameters (4)
  • conflict transfer λ = 0.6
    Controls how much pairwise conflict mass K is moved into frame ignorance; default 0.6 chosen on validation (Section 4.3, Table 9).
  • routing thresholds τ1, τ2 = 0.35, 0.65
    Partition U_global into direct answer / conflict-aware / abstention; default (0.35, 0.65) selected on validation only (Section 4.3, Table 10).
  • top-k retrieved chunks = 5
    Fixed retrieval depth for all systems; affects both evidence mass and cost (Section 4.3).
  • ECE bin count B = 15
    Calibration binning hyperparameter for reported ECE (Section 4.3).
axioms (5)
  • standard math Dirichlet evidence maps to singleton belief b_ij = e_ij/S_i and vacuity u_i = M/S_i summing to a valid mass function (Eqs. 2–6).
    Standard EDL / subjective-logic mapping (Sensoy et al., Jøsang); used as the bridge to DS masses.
  • domain assumption Unresolved conflict should be transferred into epistemic frame mass rather than normalized away (λ-transfer / Yager-style special case).
    Design choice justified by high-conflict DS literature (Yager, Smets); not forced by data.
  • ad hoc to paper A prompt-based Llama-3-8B evaluator can extract claims and non-negative evidence scores reliable enough for fusion without task-specific fine-tuning.
    Core operational premise of Section 3.2; only partially validated on 100 annotated chunks (Table 13).
  • domain assumption Claim normalization can merge paraphrases while keeping incompatible numerical/entity/proposition claims distinct without creating artificial conflict.
    Section 3.3 two-stage alignment; conservative merges assumed sufficient for open-domain answers.
  • ad hoc to paper Fused max singleton mass is a well-calibrated confidence for ECE without further (1−U_global) discounting.
    Explicit design choice in Section 4.4 to avoid double-discounting; underpins calibration claims.
invented entities (3)
  • Conflict-preserving fusion operator ⊗_λ no independent evidence
    purpose: Aggregate multi-chunk DS masses while transferring fraction λ of conflict into frame ignorance instead of classical Dempster normalization.
    Defined in Eqs. 7–13; special cases recover Yager (λ=1) and no-transfer (λ=0). Methodological construct, not a physical entity.
  • Uncertainty-guided generation router (direct / conflict-aware / abstention) no independent evidence
    purpose: Map U_global to three generator behaviors via thresholds τ1, τ2.
    Eq. 15 and Figure 2; the policy is the paper’s control interface from evidence to output style.
  • Prompt-based evidential claim extractor producing Dirichlet evidence vectors no independent evidence
    purpose: Convert unstructured retrieved chunks into candidate claims and non-negative evidence without fine-tuning the main generator.
    Section 3.2 modular evaluator; fidelity measured only in-paper on a small held-out set.

pith-pipeline@v1.1.0-grok45 · 27853 in / 3624 out tokens · 43362 ms · 2026-07-14T11:18:53.519082+00:00 · methodology

0 comments
read the original abstract

Retrieval-augmented generation grounds large language models in external evidence, but most pipelines still treat retrieved passages as deterministic and mutually consistent context. In open information environments, retrieved sources may disagree because of temporal drift, source error, ambiguity, or genuine uncertainty. This paper introduces ERAG, an uncertainty-aware RAG framework that converts retrieved chunks into probabilistic evidence before generation. A lightweight evaluator extracts candidate claims and maps chunk-level support to Dirichlet evidence. A conflict-preserving Dempster-Shafer fusion rule then transfers unresolved disagreement into epistemic uncertainty rather than normalizing it away. The generator is routed to direct answering, conflict-aware answering, or abstention according to the fused uncertainty score. Experiments on CRAG, ConflictQA, and MuSiQue show that ERAG remains competitive with the strongest matched baseline on standard question answering while improving behavior under conflict. On the CRAG ambiguous subset, hallucination decreases from 45.3% for Corrective RAG to a human-calibrated estimate of 34.8%, conflict resolution increases from 35.2% to 51.2%, and expected calibration error improves to 0.122. These results suggest that evidential modeling is a practical mechanism for trustworthy information processing in foundation-model-based retrieval systems.

Figures

Figures reproduced from arXiv: 2607.10491 by M. F. Mridha, Ruksat Khan Shayoni, S M Asif Hossain.

Figure 1
Figure 1. Figure 1: Conflict-preserving fusion of retrieved evidence. Equivalent claims reinforce the [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Uncertainty-guided generation routes. The thresholds separate ordinary evidence [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of EvidentialRAG. The sequence starts with query-specific retrieval, converts chunks into evidential masses, fuses source evidence while preserving conflict, and routes the generator according to global uncertainty. 13 [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Inference procedure for EvidentialRAG. The procedure distinguishes per-chunk evidence decisions from batched evaluator execution. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: CRAG ambiguous conflict outcomes. The left panel shows hallucination reduction, [PITH_FULL_IMAGE:figures/full_fig_p024_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

57 extracted references · 7 linked inside Pith

  1. [1]

    Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks , booktitle =

    Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and K. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks , booktitle =. 2020 , url =

  2. [2]

    Proceedings of the 37th International Conference on Machine Learning , series =

    Guu, Kelvin and Lee, Kenton and Tung, Zora and Pasupat, Panupong and Chang, Mingwei , title =. Proceedings of the 37th International Conference on Machine Learning , series =. 2020 , publisher =

  3. [3]

    Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics , pages =

    Izacard, Gautier and Grave, Edouard , title =. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics , pages =. 2021 , doi =

  4. [4]

    Borgeaud, Sebastian and Mensch, Arthur and Hoffmann, Jordan and Cai, Trevor and Rutherford, Eliza and Millican, Katie and van den Driessche, George B. and Lespiau, Jean-Baptiste and Damoc, Bogdan and Clark, Aidan and de Las Casas, Diego and Guy, Aurelia and Menick, Jacob and Ring, Roman and Hennigan, Tom and Huang, Saffron and Maggiore, Loren and Jones, C...

  5. [5]

    Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages =

    Karpukhin, Vladimir and Oguz, Barlas and Min, Sewon and Lewis, Patrick and Wu, Ledell and Edunov, Sergey and Chen, Danqi and Yih, Wen-tau , title =. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages =. 2020 , doi =

  6. [6]

    Foundations and Trends in Information Retrieval , volume =

    Robertson, Stephen and Zaragoza, Hugo , title =. Foundations and Trends in Information Retrieval , volume =. 2009 , doi =

  7. [7]

    Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , pages =

    Reimers, Nils and Gurevych, Iryna , title =. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , pages =. 2019 , doi =

  8. [8]

    Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =

    Khattab, Omar and Zaharia, Matei , title =. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =. 2020 , doi =

  9. [9]

    Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics , pages =

    Santhanam, Keshav and Khattab, Omar and Saad-Falcon, Jon and Potts, Christopher and Zaharia, Matei , title =. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics , pages =. 2022 , doi =

  10. [10]

    arXiv preprint arXiv:1901.04085 , year =

    Nogueira, Rodrigo and Cho, Kyunghyun , title =. arXiv preprint arXiv:1901.04085 , year =. doi:10.48550/arXiv.1901.04085 , url =

  11. [11]

    arXiv preprint arXiv:2312.10997 , year =

    Gao, Yunfan and Xiong, Yun and Gao, Xinyu and Jia, Kangxiang and Pan, Jinliu and Bi, Yuxi and Dai, Yi and Sun, Jiawei and Wang, Meng and Wang, Haofen , title =. arXiv preprint arXiv:2312.10997 , year =. doi:10.48550/arXiv.2312.10997 , url =

  12. [12]

    The Twelfth International Conference on Learning Representations , year =

    Asai, Akari and Wu, Zeqiu and Wang, Yizhong and Sil, Avirup and Hajishirzi, Hannaneh , title =. The Twelfth International Conference on Learning Representations , year =

  13. [13]

    arXiv preprint arXiv:2401.15884 , year =

    Yan, Shi-Qi and Gu, Jia-Chen and Zhu, Yun and Ling, Zhen-Hua , title =. arXiv preprint arXiv:2401.15884 , year =. doi:10.48550/arXiv.2401.15884 , url =

  14. [14]

    arXiv preprint arXiv:2406.04744 , year =

    Yang, Xiao and Sun, Kai and Xin, Hao and Sun, Yushi and Bhalla, Nikita and Chen, Xiangsen and Choudhary, Sajal and Gui, Rongze Daniel and Jiang, Ziran Will and Jiang, Ziyu and Kong, Lingkun and Moran, Brian and Wang, Jiaqi and Xu, Yifan Ethan and Yan, An and Yang, Chenyu and Yuan, Eting and Zha, Hanwen and Tang, Nan and Chen, Lei and Scheffer, Nicolas and...

  15. [15]

    arXiv preprint arXiv:2604.11209 , year =

    Zhao, Tianzhe and Chen, Jiaoyan and Zhang, Shuxiu and Zhu, Haiping and Lin, Qika and Liu, Jun , title =. arXiv preprint arXiv:2604.11209 , year =. doi:10.48550/arXiv.2604.11209 , url =

  16. [16]

    Advances in Neural Information Processing Systems , year =

    Su, Zhaochen and Zhang, Jun and Qu, Xiaoye and Zhu, Tong and Li, Yanshu and Sun, Jiashuo and Li, Juntao and Zhang, Min and Cheng, Yu , title =. Advances in Neural Information Processing Systems , year =

  17. [17]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages =

    Xu, Rongwu and Qi, Zehan and Guo, Zhijiang and Wang, Cunxiang and Wang, Hongru and Zhang, Yue and Xu, Wei , title =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages =. 2024 , doi =

  18. [18]

    Transactions of the Association for Computational Linguistics , volume =

    Trivedi, Harsh and Balasubramanian, Niranjan and Khot, Tushar and Sabharwal, Ashish , title =. Transactions of the Association for Computational Linguistics , volume =. 2022 , doi =

  19. [19]

    and Salakhutdinov, Ruslan and Manning, Christopher D

    Yang, Zhilin and Qi, Peng and Zhang, Saizheng and Bengio, Yoshua and Cohen, William W. and Salakhutdinov, Ruslan and Manning, Christopher D. , title =. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages =. 2018 , doi =

  20. [20]

    and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav , title =

    Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav , title...

  21. [21]

    Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =

    Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy , title =. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =. 2016 , doi =

  22. [22]

    Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics , pages =

    Thorne, James and Vlachos, Andreas and Christodoulopoulos, Christos and Mittal, Arpit , title =. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics , pages =. 2018 , doi =

  23. [23]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    Niu, Cheng and Wu, Yuanhao and Zhu, Juno and Xu, Siliang and Shum, KaShun and Zhong, Randy and Song, Juntong and Zhang, Tong , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2024 , publisher =. doi:10.18653/v1/2024.acl-long.585 , url =

  24. [24]

    Manakul, Potsawee and Liusie, Adian and Gales, Mark J. F. , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =. 2023 , doi =

  25. [25]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =

    Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang Wei and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =. 2023 , doi =

  26. [26]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages =

    Lin, Stephanie and Hilton, Jacob and Evans, Owain , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages =. 2022 , doi =

  27. [27]

    ACM Computing Surveys , volume =

    Ji, Ziwei and Lee, Nayeon and Frieske, Rita and Yu, Tiezheng and Su, Dan and Xu, Yan and Ishii, Etsuko and Bang, Yejin and Chen, Delong and Dai, Wenliang and Chan, Ho Shu and Madotto, Andrea and Fung, Pascale , title =. ACM Computing Surveys , volume =. 2023 , doi =

  28. [28]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =

    Maynez, Joshua and Narayan, Shashi and Bohnet, Bernd and McDonald, Ryan , title =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =. 2020 , doi =

  29. [29]

    , title =

    Gneiting, Tilmann and Raftery, Adrian E. , title =. Journal of the American Statistical Association , volume =. 2007 , doi =

  30. [30]

    , title =

    Guo, Chuan and Pleiss, Geoff and Sun, Yu and Weinberger, Kilian Q. , title =. Proceedings of the 34th International Conference on Machine Learning , pages =. 2017 , url =

  31. [31]

    and Nowozin, Sebastian and Dillon, Joshua V

    Ovadia, Yaniv and Fertig, Emily and Ren, Jie and Nado, Zachary and Sculley, D. and Nowozin, Sebastian and Dillon, Joshua V. and Lakshminarayanan, Balaji and Snoek, Jasper , title =. Advances in Neural Information Processing Systems , volume =. 2019 , url =

  32. [32]

    Advances in Neural Information Processing Systems , volume =

    Lakshminarayanan, Balaji and Pritzel, Alexander and Blundell, Charles , title =. Advances in Neural Information Processing Systems , volume =. 2017 , url =

  33. [33]

    Advances in Neural Information Processing Systems , volume =

    Kendall, Alex and Gal, Yarin , title =. Advances in Neural Information Processing Systems , volume =. 2017 , url =

  34. [34]

    Proceedings of the 33rd International Conference on Machine Learning , pages =

    Gal, Yarin and Ghahramani, Zoubin , title =. Proceedings of the 33rd International Conference on Machine Learning , pages =. 2016 , url =

  35. [35]

    Advances in Neural Information Processing Systems , volume =

    Sensoy, Murat and Kaplan, Lance and Kandemir, Melih , title =. Advances in Neural Information Processing Systems , volume =. 2018 , url =

  36. [36]

    Subjective Logic: A Formalism for Reasoning Under Uncertainty , publisher =

    J. Subjective Logic: A Formalism for Reasoning Under Uncertainty , publisher =. 2016 , doi =

  37. [37]

    , title =

    Dempster, Arthur P. , title =. The Annals of Mathematical Statistics , volume =. 1967 , doi =

  38. [38]

    1976 , url =

    Shafer, Glenn , title =. 1976 , url =

  39. [39]

    , title =

    Yager, Ronald R. , title =. Information Sciences , volume =. 1987 , doi =

  40. [40]

    IEEE Transactions on Pattern Analysis and Machine Intelligence , volume =

    Smets, Philippe , title =. IEEE Transactions on Pattern Analysis and Machine Intelligence , volume =. 1990 , doi =

  41. [41]

    and Kaiser, Lukasz and Polosukhin, Illia , title =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia , title =. Advances in Neural Information Processing Systems , volume =. 2017 , url =

  42. [42]

    Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pages =

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , title =. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pages =. 2019 , doi =

  43. [43]

    Brown, Tom B. and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom and Child, Rewon and Ramesh, Aditya and Ziegler, Daniel M. and Wu, Jeffrey and W...

  44. [44]

    2023 , doi =

    GPT-4 Technical Report , journal =. 2023 , doi =

  45. [45]

    LLaMA: Open and Efficient Foundation Language Models , journal =

    Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth. LLaMA: Open and Efficient Foundation Language Models , journal =. 2023 , doi =

  46. [46]

    arXiv preprint arXiv:2407.21783 , year =

    Grattafiori, Aaron and others , title =. arXiv preprint arXiv:2407.21783 , year =. doi:10.48550/arXiv.2407.21783 , url =

  47. [47]

    Jiang, Albert Q. and Sablayrolles, Alexandre and Mensch, Arthur and Bamford, Chris and Chaplot, Devendra Singh and de Las Casas, Diego and Bressand, Florian and Lengyel, Gianna and Lample, Guillaume and Saulnier, Lucile and Lavaud, L. Mistral 7B , journal =. 2023 , doi =

  48. [48]

    and Zhang, Hao and Stoica, Ion , title =

    Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph E. and Zhang, Hao and Stoica, Ion , title =. Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , pages =. 2023 , doi =

  49. [49]

    , title =

    Wolf, Thomas and Debut, Lysandre and Sanh, Victor and Chaumond, Julien and Delangue, Clement and Moi, Anthony and Cistac, Pierric and Rault, Tim and Louf, Remi and Funtowicz, Morgan and Davison, Joe and Shleifer, Sam and von Platen, Patrick and Ma, Clara and Jernite, Yacine and Plu, Julien and Xu, Canwen and Le Scao, Teven and Gugger, Sylvain and Drame, M...

  50. [50]

    and Artzi, Yoav , title =

    Zhang, Tianyi and Kishore, Varsha and Wu, Felix and Weinberger, Kilian Q. and Artzi, Yoav , title =. International Conference on Learning Representations , year =

  51. [51]

    Text Summarization Branches Out , pages =

    Lin, Chin-Yew , title =. Text Summarization Branches Out , pages =. 2004 , url =

  52. [52]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =

    Gao, Tianyu and Yen, Howard and Yu, Jiatong and Chen, Danqi , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =. 2023 , publisher =. doi:10.18653/v1/2023.emnlp-main.398 , url =

  53. [53]

    The Twelfth International Conference on Learning Representations , year =

    Xie, Jian and Zhang, Kai and Chen, Jiangjie and Lou, Renze and Su, Yu , title =. The Twelfth International Conference on Learning Representations , year =

  54. [54]

    arXiv preprint arXiv:2407.11005 , year =

    Friel, Robert and Belyi, Masha and Sanyal, Atindriyo , title =. arXiv preprint arXiv:2407.11005 , year =. doi:10.48550/arXiv.2407.11005 , url =

  55. [55]

    Educational and Psychological Measurement , volume =

    Cohen, Jacob , title =. Educational and Psychological Measurement , volume =. 1960 , doi =

  56. [56]

    The Annals of Statistics , volume =

    Efron, Bradley , title =. The Annals of Statistics , volume =. 1979 , doi =

  57. [57]

    Findings of the Association for Computational Linguistics: ACL 2024 , pages =

    Chen, Jianlyu and Xiao, Shitao and Zhang, Peitian and Luo, Kun and Lian, Defu and Liu, Zheng , title =. Findings of the Association for Computational Linguistics: ACL 2024 , pages =. 2024 , publisher =. doi:10.18653/v1/2024.findings-acl.137 , url =