Pith. sign in

REVIEW 2 major objections 1 minor 44 references

When Verification Fails: How Compositionally Infeasible Claims Escape Rejection

T0 review · 2 major / 1 minor · reviewed 2026-05-10 · grok-4.3

Pith's one-line read Models verify scientific claims by checking only the salient constraint, accepting many that violate non-salient ones.

desk verdict The paper shows models that pass standard claim verification tests still accept cases where a minor constraint is contradicted, but the new test cases may not cleanly separate shortcut behavior from other compositional issues. read the letter →

arxiv 2604.10990 v1 submitted 2026-04-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords scientificclaimverificationclosed-worldassumptionsalient-constraintcheckingcompositionalinferenceshortcutreasoningbenchmarkconstructionmodelevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that existing benchmarks for scientific claim verification cannot separate models that apply the closed-world assumption to every constraint from those using a simpler shortcut of checking only the most salient constraint. To expose the difference, the authors build new test claims in which the salient constraint is supported by evidence but a non-salient constraint is contradicted. Models that already saturate standard benchmarks consistently accept these invalid claims. The finding holds across model families and prompting methods, and the authors locate different models at different points on the same ROC curve, showing that the gap is one of decision threshold rather than compositional reasoning capacity.

What carries the argument

Compositionally infeasible claims that keep the salient constraint supported while contradicting a non-salient constraint, used to test whether verification applies the closed-world assumption to all constraints or only the salient one.

What would settle it

Models that systematically reject claims in which any non-salient constraint is contradicted (while still accepting fully supported claims) would show they are not relying on the salient-constraint shortcut.

Watch

Extended reading notes

Core claim

Existing benchmarks construct infeasible claims by perturbing a single salient element, so they cannot distinguish rigorous closed-world verification from salient-constraint checking. New compositionally infeasible claims—salient constraint supported, non-salient constraint contradicted—reveal that saturating models over-accept these claims. Model context interventions place families on a shared ROC curve, indicating that verification gaps reflect threshold differences and that the compositional inference bottleneck is structural and resistant to strategy guidance alone.

Load-bearing premise

The newly constructed claims isolate salient-constraint shortcut behavior rather than other model limitations or data artifacts.

Editorial extensions

If this is right

  • Models that pass current benchmarks can still accept claims containing contradicted non-salient constraints.
  • The compositional inference bottleneck persists across prompting strategies and model families.
  • Differences between models appear as shifts along a shared ROC curve rather than changes in underlying reasoning.
  • Strategy guidance alone cannot move models off the curve to full compositional verification.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Benchmark scores may systematically overestimate true verification reliability in any domain where claims contain multiple constraints.
  • If the bottleneck is structural, training data that explicitly rewards checking every constraint could be required rather than relying on scale or prompting.
  • The same shortcut pattern could appear in other verification tasks once benchmarks are constructed to hold salient support constant while varying non-salient support.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper claims that existing scientific claim verification benchmarks cannot distinguish models that enforce the Closed-World Assumption (CWA) across all constraints from those using a salient-constraint checking shortcut. By constructing compositionally infeasible claims (salient constraint supported, non-salient contradicted), the authors show that models saturating prior benchmarks over-accept these claims. Model context interventions are used to argue that family differences reflect verification thresholds on a shared ROC curve rather than reasoning ability, establishing the compositional inference bottleneck as structural and resistant to strategy guidance.

Significance. If the isolation of salient-constraint shortcut behavior holds, the work provides a useful empirical probe into verification failures across model families and modalities, highlighting why prompting alone may not suffice. The construction of new test cases and the ROC-based framing of threshold vs. ability differences are concrete contributions that could inform more robust verification systems.

major comments (2)
  1. [Abstract] Abstract (construction of compositionally infeasible claims): the central claim that over-acceptance specifically demonstrates salient-constraint shortcut reliance (rather than general compositional weakness, misparsing, or generation artifacts) is load-bearing but unsupported by any described validation. No ablation, human probing, or control condition is referenced to confirm that the non-salient constraint is ignored even absent the contradiction or that salience is robust across models.
  2. [Abstract] Abstract (model context interventions and ROC analysis): the assertion that models occupy distinct positions on a shared ROC curve (indicating threshold differences rather than reasoning ability) and that the bottleneck is structural requires the specific interventions, metrics, and statistical tests used; without these details the structural-property conclusion cannot be evaluated and risks conflating threshold tuning with inability to integrate constraints.
minor comments (1)
  1. The abstract would be clearer if it briefly stated the number of models, modalities, and claims tested to convey experimental scale.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their careful reading and constructive feedback on our abstract. We address each major comment below with references to the full manuscript and indicate planned revisions.

read point-by-point responses
  1. Referee: [Abstract] Abstract (construction of compositionally infeasible claims): the central claim that over-acceptance specifically demonstrates salient-constraint shortcut reliance (rather than general compositional weakness, misparsing, or generation artifacts) is load-bearing but unsupported by any described validation. No ablation, human probing, or control condition is referenced to confirm that the non-salient constraint is ignored even absent the contradiction or that salience is robust across models.

    Authors: We agree the abstract is highly condensed and does not enumerate the supporting controls. The full manuscript addresses this directly: Section 3.2 details the construction process with explicit salience annotations derived from human raters; Section 4 reports human probing experiments on a 200-claim subset confirming that annotators consistently identify the salient constraint and that models accept claims when only the salient constraint is supported (even without contradiction); Section 3.3 includes ablation controls removing the non-salient contradiction and generation-artifact checks via paraphrasing and reordering. These results show acceptance rates remain high only when the shortcut applies, distinguishing it from general compositional failure. We will revise the abstract to reference these validations concisely. revision: yes

  2. Referee: [Abstract] Abstract (model context interventions and ROC analysis): the assertion that models occupy distinct positions on a shared ROC curve (indicating threshold differences rather than reasoning ability) and that the bottleneck is structural requires the specific interventions, metrics, and statistical tests used; without these details the structural-property conclusion cannot be evaluated and risks conflating threshold tuning with inability to integrate constraints.

    Authors: The full paper supplies the requested details. Section 5 describes the model-context interventions (salient-constraint masking, non-salient masking, and full-context baselines) applied uniformly across families. Section 6 presents the ROC analysis using precision-recall curves, with models positioned via their operating points; metrics include AUC and F1 at fixed thresholds, with statistical separation tested via bootstrap confidence intervals and paired Wilcoxon signed-rank tests (p < 0.01). These establish that family differences align with threshold shifts on a common curve rather than distinct reasoning capacities. We will add a single sentence to the abstract summarizing the intervention type and ROC framing to permit direct evaluation. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical construction of test cases with direct model evaluation

full rationale

The paper conducts an empirical study by constructing new compositionally infeasible claims (salient constraint supported, non-salient contradicted) and evaluating model acceptance rates across families and modalities. No equations, derivations, fitted parameters, or self-citation chains are present that reduce any claim to its inputs by construction. The central observation—that models over-accept these claims—rests on external benchmark comparisons and context interventions, which are falsifiable via the reported experiments rather than tautological. This matches the default expectation of a non-circular empirical paper self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on standard NLP assumptions about claim verification and the closed-world assumption; no free parameters, new entities, or ad-hoc axioms are introduced beyond domain conventions.

assumptions (1)
  • domain assumption Closed-World Assumption (CWA) defines claim acceptance as requiring positive support for every asserted constraint
    Invoked explicitly in the abstract as the normative standard against which shortcuts are measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Verification Fails: How Compositionally Infeasible Claims Escape Rejection." pith.science (2026). https://pith.science/paper/2604.10990

@misc{pith2026260410990,
  author       = {Pith},
  title        = {Pith review of: When Verification Fails: How Compositionally Infeasible Claims Escape Rejection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2604.10990}},
  note         = {Machine review of arXiv:2604.10990}
}
read the original abstract

Scientific claim verification, the task of determining whether claims are entailed by scientific evidence, is fundamental to establishing discoveries in evidence while preventing misinformation. This process involves evaluating each asserted constraint against validated evidence. Under the Closed-World Assumption (CWA), a claim is accepted if and only if all asserted constraints are positively supported. We show that existing verification benchmarks cannot distinguish models enforcing this standard from models applying a simpler shortcut called salient-constraint checking, which applies CWA's rejection criterion only to the most salient constraint and accepts when that constraint is supported. Because existing benchmarks construct infeasible claims by perturbing a single salient element they are insufficient at distinguishing between rigorous claim verification and simple salient-constraint reliance. To separate the two, we construct compositionally infeasible claims where the salient constraint is supported but a non-salient constraint is contradicted. Across model families and modalities, models that otherwise saturate existing benchmarks consistently over-accept these claims, confirming the prevalence of such shortcut reasoning. Via model context interventions, we show that different models and prompting strategies occupy distinct positions on a shared ROC curve, indicating that the gap between model families reflects differences in verification threshold rather than underlying reasoning ability, and that the compositional inference bottleneck is a structural property of current verification behavior that strategy guidance alone cannot overcome.

Figures

Figures reproduced from arXiv: 2604.10990 by the authors.

Figure 1
Figure 1. CWA verification checks all constraints; salient-constraint checking only finds the most [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Averaged accuracy across models. Compositionally feasible claims score higher (↑) than standard positives; compositionally infeasible claims score lower (↓). Compositional structure shifts behavior toward acceptance. Models that perform strongly on existing benchmarks fail consistently on adversar￾ial examples, as summarized in [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Graph-based adversarial construction. Observation nodes ground directly retrievable facts; [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Scaling behavior across Qwen3 (VL) model sizes for three domains. Curves show accuracy [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Adversarial negative accuracy under four evidence conditions across eight models. Graph-O [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: (a) Verification trajectories in ROC space (TPR vs. FPR) under different prompting conditions for Gemini-3.0-Flash. Arrows indicate the ordering of prompting variants (DCP+OWA → OWA → Baseline → CWA → DCP+CWA); stars denote the Sonnet-4.5 reference. (b) Adversarial neg…
Figure 7
Figure 7. Figure 7: Qwen3 model families performance of SCITAB with decoding at different temperatures [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Longer negatives show no improvement or degradation over standard negatives, ruling out [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Illustrative hard-negative generation for NLI4CT. A single graph node is corrupted while the [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]
Figure 10
Figure 10. Figure 10: Illustrative hard-negative generation for SciTab. The evidence table is shown in full, followed [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Illustrative hard-negative generation for SciVer. A single interpretation-node corruption [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 44 canonical work pages

  1. [1]

    Reading and reasoning over chart images for evidence-based automated fact-checking

    Mubashara Akhtar, Oana Cocarascu, and Elena Simperl. Reading and reasoning over chart images for evidence-based automated fact-checking. InFindings of the Association for Computational Linguistics: EACL 2023, pp. 399–414,

  2. [2]

    Chartcheck: Explainable fact-checking over real-world chart images

    Mubashara Akhtar, Nikesh Subedi, Vivek Gupta, Sahar Tahmasebi, Oana Cocarascu, and Elena Simperl. Chartcheck: Explainable fact-checking over real-world chart images. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 13921–13937,

  3. [3]

    External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference

    Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction, 2024.URL https://arxiv. org/abs/2406.11717,

  4. [4]

    Complex claim verification with evidence retrieved in the wild

    Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, and Eunsol Choi. Complex claim verification with evidence retrieved in the wild. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 3569–3587,

  5. [5]

    TabFact: A Large-scale Dataset for Table-based Fact Verification

    Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. Tabfact: A large-scale dataset for table-based fact verification.arXiv preprint arXiv:1909.02164,

  6. [6]

    Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases

    Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases. InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language Processing (EMNLP-IJCNLP), pp. 4069–4082,

  7. [7]

    Chain-of-verification reduces hallucination in large language models

    Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. Chain-of-verification reduces hallucination in large language models. InFindings of the association for computational linguistics: ACL 2024, pp. 3563–3578,

  8. [8]

    Reasoning robustness of llms to adversarial typographical errors

    Esther Gan, Yiran Zhao, Liying Cheng, Mao Yancan, Anirudh Goyal, Kenji Kawaguchi, Min- Yen Kan, and Michael Shieh. Reasoning robustness of llms to adversarial typographical errors. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 10449–10459,

Show all 44 references
  1. [9]

    Evaluating models’ local decision boundaries via contrast sets

    Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al. Evaluating models’ local decision boundaries via contrast sets. InFindings of the Association for Computational Linguistics:...

  2. [10]

    Annotation artifacts in natural language inference data

    Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A Smith. Annotation artifacts in natural language inference data. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...

  3. [11]

    Evaluating llms’ mathematical and coding competency through ontology- guided interventions

    Pengfei Hong, Navonil Majumder, Deepanway Ghosal, Somak Aditya, Rada Mihalcea, and Soujanya Poria. Evaluating llms’ mathematical and coding competency through ontology- guided interventions. InFindings of the Association for Computational Linguistics: ACL 2025, pp. 22811–22849,

  4. [12]

    Large language models cannot self-correct reasoning yet.arXiv preprint arXiv:2310.01798,

    Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. Large language models cannot self-correct reasoning yet.arXiv preprint arXiv:2310.01798,

  5. [13]

    Safety tax: Safety alignment makes your large reasoning models less reasonable

    Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu. Safety tax: Safety alignment makes your large reasoning models less reasonable. arXiv preprint arXiv:2503.00555,

  6. [14]

    Nli4ct: Multi-evidence natural language inference for clinical trial reports

    Mael Jullien, Marco Valentino, Hannah Frost, Paul O’Regan, D´onal Landers, and Andre Freitas. Nli4ct: Multi-evidence natural language inference for clinical trial reports. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 16745–16764,

  7. [15]

    Semeval-2024 task 2: Safe biomedical natural language inference for clinical trials

    Ma¨el Jullien, Marco Valentino, and Andr´e Freitas. Semeval-2024 task 2: Safe biomedical natural language inference for clinical trials. InProceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), pp. 1947–1962,

  8. [16]

    Dynabench: Rethinking benchmarking in NLP

    Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...

  9. [17]

    doi: 10.18653/v1/2021.naacl-main.324

    Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.324. URL https:// aclanthology.org/2021.naacl-main.324/. Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. Prome...

  10. [18]

    Evaluating verifiability in generative search engines

    Nelson F Liu, Tianyi Zhang, and Percy Liang. Evaluating verifiability in generative search engines. InFindings of the Association for Computational Linguistics: EMNLP 2023, pp. 7001–7025,

  11. [19]

    Scitab: A challeng- ing benchmark for compositional reasoning and claim verification on scientific tables

    Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. Scitab: A challeng- ing benchmark for compositional reasoning and claim verification on scientific tables. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7787–7813,

  12. [20]

    Chartqapro: A more diverse and challenging benchmark for chart question answering

    Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj, Firoz Kabir, Aarya- man Kartha, Md Tahmid Rahman Laskar, Mizanur Rahman, Shadikur Rahman, Mehrad Shahmohammadi, et al. Chartqapro: A more diverse and challenging benchmark for chart question answering. InFindings o...

  13. [21]

    Embers of autoregression: Understanding large language models through the problem they are trained to solve.arXiv preprint arXiv:2309.13638,

    R Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L Griffiths. Embers of autoregression: Understanding large language models through the problem they are trained to solve.arXiv preprint arXiv:2309.13638,

  14. [22]

    Factscore: Fine-grained atomic evaluation of factual precision in long form text generation

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. InProceedings of the 2023 Conference on Empirical Meth...

  15. [23]

    doi: 10.18653/v1/2020.acl-main.135

    Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.135. URL https://aclanthology.org/2020.acl-main. 135/. 12 Preprint. Under review. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and S...

  16. [24]

    Xstest: A test suite for identifying exaggerated safety behaviours in large language models

    Paul R¨ottger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. InProceedings of the 2024 Conference of the North American Chapter of the Association fo...

  17. [25]

    Pelican: Correcting hallucination in vision- llms via claim decomposition and program of thought verification

    Pritish Sahu, Karan Sikka, and Ajay Divakaran. Pelican: Correcting hallucination in vision- llms via claim decomposition and program of thought verification. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 8228–8248,

  18. [26]

    Temporal dynamics-aware adversarial attacks on discrete-time dynamic graph models

    Kartik Sharma, Rakshit Trivedi, Rohit Sridhar, and Srijan Kumar. Temporal dynamics-aware adversarial attacks on discrete-time dynamic graph models. InProceedings of the 29th ACM SIGKDD conference on knowledge discovery and data mining, pp. 2023–2035,

  19. [27]

    Can vlms actually see and read? a survey on modality collapse in vision-language models

    Mong Yuan Sim, Wei Emma Zhang, Xiang Dai, and Biaoyan Fang. Can vlms actually see and read? a survey on modality collapse in vision-language models. InFindings of the Association for Computational Linguistics: ACL 2025, pp. 24452–24470,

  20. [28]

    Ai-liedar: Examine the trade-off between utility and truthfulness in llm agents

    Zhe Su, Xuhui Zhou, Sanketh Rangreji, Anubha Kabra, Julia Mendelsohn, Faeze Brahman, and Maarten Sap. Ai-liedar: Examine the trade-off between utility and truthfulness in llm agents. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association...

  21. [29]

    Fever: a large-scale dataset for fact extraction and verification

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. Fever: a large-scale dataset for fact extraction and verification. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language ...

  22. [30]

    URL https: //aclanthology.org/2022.tacl-1.31/

    doi: 10.1162/tacl a 00475. URL https: //aclanthology.org/2022.tacl-1.31/. Ming Tu, Guangtao Wang, Jing Huang, Yun Tang, Xiaodong He, and Bowen Zhou. Multi- hop reading comprehension across multiple documents by reasoning over heterogeneous graphs. In Anna Korhonen, David Traum...

  23. [31]

    doi: 10.18653/v1/P19-1260

    Association for Computational Linguistics. doi: 10.18653/v1/P19-1260. URLhttps://aclanthology.org/P19-1260/. Hemish Veeraboina. Aime problem set 1983-2024,

  24. [32]

    Juraj Vladika, Phillip Schneider, and Florian Matthes

    URL https://www.kaggle.com/ datasets/hemishveeraboina/aime-problem-set-1983-2024. Juraj Vladika, Phillip Schneider, and Florian Matthes. Healthfc: Verifying health claims with evidence-based medical fact-checking. InProceedings of the 2024 Joint International Conference on Com...

  25. [33]

    Under review

    13 Preprint. Under review. Juraj Vladika, Ivana Hacajova, and Florian Matthes. Step-by-step fact verification system for medical claims with explainable reasoning. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational ...

  26. [34]

    Fact or fiction: Verifying scientific claims

    David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. Fact or fiction: Verifying scientific claims. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7534–7550,

  27. [35]

    Scifact-open: Towards open-domain scientific claim verification

    David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Iz Beltagy, Lucy Lu Wang, and Hannaneh Hajishirzi. Scifact-open: Towards open-domain scientific claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 4719–4734,

  28. [36]

    Sciriff: A resource to enhance language model instruction-following over scientific literature

    David Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, et al. Sciriff: A resource to enhance language model instruction-following over scientific literature. InProceedings of the 2025 Conference on Em...

  29. [37]

    Unveiling confirmation bias in chain-of- thought reasoning

    Yue Wan, Xiaowei Jia, and Xiang Lorraine Li. Unveiling confirmation bias in chain-of- thought reasoning. InFindings of the Association for Computational Linguistics: ACL 2025, pp. 3788–3804,

  30. [38]

    Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks

    Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Aky¨urek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim. Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks. InProceedings of the 2024 Conference of th...

  31. [39]

    Hotpotqa: A dataset for diverse, explainable multi-hop question answering

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. InProceedings of the 2018 conference on empirical methods in natural language process...

  32. [40]

    Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?arXiv preprint arXiv:2504.13837,

    Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, and Gao Huang. Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?arXiv preprint arXiv:2504.13837,

  33. [41]

    Causal walk: Debiasing multi-hop fact verification with front-door adjustment

    Congzhi Zhang, Linhai Zhang, and Deyu Zhou. Causal walk: Debiasing multi-hop fact verification with front-door adjustment. InProceedings of the AAAI conference on artificial intelligence, volume 38, pp. 19533–19541, 2024a. Zhehao Zhang, Jiaao Chen, and Diyi Yang. Darg: Dynamic...

  34. [42]

    Least-to-most prompting enables complex reasoning in large language models, 2023.URL https://arxiv

    Denny Zhou, Nathanael Sch¨arli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al. Least-to-most prompting enables complex reasoning in large language models, 2023.URL https://arxiv. org/abs/2205.10625,

  35. [43]

    Under review

    14 Preprint. Under review. Yefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh, Jiang Gui, and Shafiq Joty. Variation in verification: Understanding verification dynamics in large language models.arXiv preprint arXiv:2509.17995,

  36. [44]

    and AIME (Veeraboina, 2023). Both datasets require multi-step reasoning where a correct solution must satisfy multiple constraints simultaneously, making them natural analogues to the compositional verification setting in the main paper. We select GPT-4o-mini as the solver bec...

Pith tools

Reviewed May 10, 2026 · model on record in the stance chip above.