Pith. sign in

REVIEW 4 major objections 6 minor 128 references

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that scientific fact-checking must move from abstract-only corpora to full-paper evidence retrieval, and shows that adding verification feedback to a semantic reranker improves evidence retrieval.

desk verdict A competent perspective paper with one real preliminary result—verification-feedback reranking improves evidence Recall@k—but its central claim that retrieval gates verification accuracy is asserted, not tested. read the letter →

arxiv 2506.20844 v2 pith:IFLAD24X submitted 2025-06-25 cs.IR cs.CL

classification cs.IRcs.CL
keywords scientificfact-checkingevidenceretrievalverificationfeedbackfull-papertime-awarecitationtrackingmultimodalretrieval-augmented
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that scientific fact-checking has been held back by simplified, abstract-only benchmarks that avoid the real difficulties of full academic papers, and that the next phase should be specialised evidence retrieval over complete documents. Its preliminary experiment shows that combining a semantic relevance score with verification probabilities from a downstream verifier improves document evidence retrieval on SciFact-Open and Check-COVID across nearly all cutoff thresholds. This supports the paper's central premise that retrieval quality is a prerequisite for accurate claim verification, and that retrieval should be optimised for evidential utility rather than semantic similarity alone. The paper also lays out eight research directions covering time-aware retrieval, citation tracking, structured long-context parsing, multimodal evidence, credibility, and terminology modelling.

What carries the argument

The central mechanism is the +Verification reranker, whose final score for a claim $c$ and document $d$ is $s^{r+v}_{c,d} = 1/2 \cdot (s^r_{c,d} + p^r_{c,d} + p^s_{c,d})$, combining the monoT5-3B semantic relevance score with MultiVerS probabilities for support and refute. This formulation gives higher retrieval scores to documents that contribute to support or refute labels, demonstrating that verification feedback can act as a fine-grained evidence-utility signal. A second mechanism is the analysis of relevance score distributions for well-studied versus less-studied claims, showing that statistics such as first-document score, mean score, total score, initial-to-final ratio, and exponential decay differ between the two groups, which the paper proposes as features for flexible retrieval cutoffs.

What would settle it

Run the full pipeline end-to-end on SciFact-Open and Check-COVID using the same verifier with evidence retrieved by +Verification versus monoT5-3B alone; if final verification accuracy does not improve with the higher Recall@k, the claim that verification feedback yields useful evidence retrieval is falsified.

Watch

Extended reading notes

Core claim

The paper establishes that current scientific fact-checking systems operate on simplified versions of the problem built on small-scale abstract datasets, and that moving to full-paper evidence retrieval requires retrieval models that account for evidential value, document structure, time, credibility, and multimodal content. Its experimental contribution is a reranker called +Verification, which sums a monoT5-3B semantic relevance score with MultiVerS support and refute probabilities for each candidate document. This verification-aware reranker outperforms semantic-only reranking, with Recall@5 rising from 57.17% to 62.61% on SciFact-Open and from 82.06% to 84.04% on Check-COVID, and with larger gains at lower cutoffs where retrieving the most relevant evidence matters most. A case study shows that verification feedback rescues documents that semantic reranking ranks far down, such as one gold evidence document moved from rank 302 to rank 27.

Load-bearing premise

The load-bearing premise is that evidence retrieval, not verification capacity or annotation quality, is the binding constraint on scientific fact-checking accuracy, and that verification probabilities learned from 809 SciFact abstracts transfer as reliable evidence-utility signals to a 500,000-document corpus.

Editorial extensions

If this is right

  • Retrieval systems for fact-checking should be trained to optimise evidential utility, not just semantic similarity, by incorporating verification feedback into reranking.
  • Full-paper evidence retrieval would enable time-aware filtering and citation tracking to reduce the influence of outdated scientific evidence.
  • Relevance score distribution features can support flexible cutoffs that reduce verifier workload for claims with scarce evidence while preserving recall for well-studied claims.
  • A specialised IR benchmark for scientific fact-checking should include verification accuracy, decision latency, and noise robustness as evaluation metrics, not only recall at fixed cutoffs.
  • Structured document parsing, section-aware retrieval, multimodal evidence alignment, credibility indicators, and ontology-based concept modelling are all needed to move from abstract-level to full-paper scientific fact-checking.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If verification-feedback retrieval holds up in end-to-end evaluation, the same utility-aware retrieval idea could transfer to other knowledge-intensive tasks such as retrieval-augmented question answering, where evidence usefulness matters more than surface relevance.
  • The disclosed change in MultiVerS negative sampling from 20 to 5 should be ablated to confirm that the retrieval gains come from verification feedback itself rather than from altered training dynamics.
  • The paper's own logic implies that future benchmarks should measure whether improved retrieval actually raises final verification accuracy, since Recall@k gains alone do not guarantee better verdicts.
  • Time-aware retrieval could be tested directly on claim sets with known scientific reversals, such as early COVID-19 treatment claims, to see whether recency weighting changes the final veracity labels.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This perspective paper argues that scientific fact-checking is currently bottlenecked by evidence retrieval over simplified, abstract-level corpora, and that the field should move toward full-paper retrieval over large-scale, structured, multimodal, and time-evolving literature. It synthesizes prior work into eight research directions (RD.1-RD.8) covering evidence-aware retrieval, evidence imbalance, temporal and citation-aware retrieval, structured long-context parsing, multimodal content, credibility, and terminology. The paper includes a preliminary experiment in Section 2 in which verification probabilities from MultiVerS (trained on the SciFact train split) are combined with monoT5-3B relevance scores to rerank candidate documents; the +Verification model is evaluated via Recall@k on SciFact-Open and Check-COVID and is reported to outperform monoT5-3B at most cut-offs. The authors use these results to motivate RD.1 and the broader claim that evidence-aware retrieval is a prerequisite for accurate scientific claim verification.

Significance. The paper is a useful and well-cited synthesis that identifies a genuine gap: SciFact/SciFact-Open are abstract-only, and the challenges of full-paper evidence (section structure, long-range dependencies, tables and figures, citations, temporal drift, credibility) are underexplored in scientific fact-checking. Its main strength is that it converts the gap into concrete research directions with references to adjacent IR/NLP work. The +Verification experiment is a clean and honest preliminary study: it uses a held-out evaluation (verifier trained on SciFact train; retrieval evaluated on SciFact-Open and Check-COVID), the result is not forced by construction, and the code for the base verifier is linked. The main weakness is that the experimental evidence supports only gold-evidence Recall@k, not the paper's stronger framing that retrieval is the binding constraint on verification accuracy; the roadmap therefore rests partly on an assumed link. With the framing tightened and the experimental details clarified, the paper would make a solid contribution to the ICTIR perspective track.

major comments (4)
  1. [Section 1 and Section 2 (Results)] The paper's central premise, that effective evidence retrieval is a prerequisite for accurate claim verification and that retrieval quality gates end-to-end performance, is not tested by the reported experiment. Section 2 evaluates only gold-evidence Recall@k; no claim-level verification accuracy or F1 is reported for either monoT5-3B or +Verification. Improving Recall@k does not establish that a downstream verifier will make better veracity decisions: the verifier may be insensitive to the rank changes produced, or the documents promoted by +Verification could confuse it. Since this premise motivates RD.1 and the overall roadmap, please either soften the framing to 'retrieval is a necessary component supported by prior noise-sensitivity findings' or add an end-to-end experiment in which top-k outputs from both systems are passed to a fixed verifier and claim-level accuracy is reported.
  2. [Section 2, Eq. (3)] Equation (3) is inconsistent with the prose. The text says the final score combines the semantic score with the verification probabilities p^r and p^n, but Eq. (3) adds p^r and p^s, where p^s denotes the refutation probability according to Eq. (2). The intended combination rule is therefore unclear. Please correct the notation, define all symbols once, and state explicitly whether the feedback term is the support probability, the refute probability, or a combination, and how the normalisation factor 1/2 is derived.
  3. [Section 2, Approaches] The negative-sampling parameter of MultiVerS is changed from 20 to 5 without an ablation or validation-based justification. Because the verification feedback is the novel component of +Verification, this unablated change confounds the comparison: the Recall@k gain could be due to the retrained verifier rather than to the feedback mechanism. Please report results for the original setting (or a small sweep) and for the chosen setting, or justify the change with a validation experiment. Relatedly, the transfer of a verifier trained on 809 SciFact abstracts to a 500K-document corpus is assumed rather than measured; a sentence acknowledging this limitation and, ideally, a per-dataset analysis of verifier confidence would strengthen the claim.
  4. [Section 4.2, Figure 2] The GPT-4o-based analysis of cited versus self-contained evidence on 22 papers is presented as an empirical observation motivating RD.4, but no validation of the LLM classifications is reported. The sentence 'we observed a few inaccurate outcomes' indicates known errors, yet no accuracy rate, agreement measure, or manual correction procedure is provided. Please either report a human-validated accuracy figure or explicitly label the analysis as anecdotal and preliminary.
minor comments (6)
  1. [Section 2, Table 2] The ordinal suffixes '141th' and '163th' should be '141st' and '163rd'.
  2. [Section 5, paragraph beginning 'Fact-checking requires explicit retrieval of explicit.'] The phrase contains a duplicated or incomplete word; it should likely read 'explicit retrieval of explicit evidence' or 'explicit retrieval of evidence'.
  3. [Section 4.1] The sentence 'A time-aware system achieved a 15% improvement in macro F1 score' should identify the baseline and dataset for this figure rather than only citing [8].
  4. [Section 3, Table 4] Please report the number of claims and the variance or standard deviation for the reported statistics; without them, the contrast between 'well-studied' and 'less-studied' claims is difficult to assess.
  5. [Section 4.2] The sentence 'Citation graph analysis [17,96] can help map citation paths using directed graphs and applying search algorithms' has a grammatical mismatch; change 'applying' to 'apply'.
  6. [Section 3, paragraph on fixed cut-offs] The phrase 'one of the most inefficient approaches is setting the cut-off equal to the maximum number of evidence per claim' is awkward; consider 'a particularly inefficient approach is setting the cut-off to the maximum number of gold evidence documents per claim'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the +Verification experiment is a held-out evaluation, and the roadmap's premise, while not end-to-end validated, is not equivalent to its inputs.

full rationale

The paper's only derived result is the Section 2 comparison: "The integration of verification feedback consistently enhanced document evidence retrieval, as +Verification outperformed monoT5-3B across nearly all cut-off thresholds in both SciFact-Open and Check-COVID (Table 1)." This is not circular by construction: MultiVerS is trained on the SciFact train split ("MultiVerS is trained on the SciFact train set to create a verifier model that provides verification feedback"), and the evaluation uses SciFact-Open (claims from the SciFact test set over a 500K corpus) and Check-COVID, so the verifier feedback could have failed to transfer and the recall gains are a genuine held-out outcome. The 0.5 combination weight and the negative-sampling change from 20 to 5 are disclosed heuristics, not parameters fitted to the test sets. Equation 3 combines semantic relevance with verifier support/refute probabilities; although the surrounding prose mentions p_n while the equation uses p_s (an apparent typo), this does not make the score equal to the gold evidence labels. The central premise "effective retrieval is a prerequisite for accurate claim verification" is asserted with supporting citations [76, 117] and is not claimed to be proven by the recall-only experiment; the absence of end-to-end verification accuracy is an evidentiary limitation, not a circular step. The only self-citations (Bin-Hezam and Stevenson 2023/2024; Fletcher and Stevenson 2025; Stevenson and Bin-Hezam 2023) support background on stopping methods and retracted research in RD.2 and RD.7 and are not load-bearing. No step reduces to its own input, so the circularity score is 0.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claim rests on two load-bearing assumptions: that gold annotations are complete (needed for Recall@k to mean anything) and that a small abstract-trained verifier transfers to large corpora. The +Verification method adds one hand-chosen weight (0.5) and one disclosed hyperparameter change (negative sampling 20 to 5). The GPT-4o evidence-source analysis in Section 4.2 depends on an unvalidated classifier over 22 papers. No free parameters are fitted to the central evaluation outcome in the sense of optimizing the reported recall, so circularity burden stays low, but the absence of ablations on the two tunable choices weakens the experimental support.

free parameters (2)
  • score combination weight alpha in +Verification formula = 0.5
    Equation 3: s^{r+v} = 1/2 * (s^r + p^r + p^s). The equal-weight fixed normalization is hand-chosen; no ablation or grid search is reported.
  • MultiVerS negative sampling parameter = 5 (default 20)
    Section 2, Experiment Overview: changed 'to avoid over-fitted verification feedback'; disclosed but no ablation or sensitivity analysis.
assumptions (3)
  • domain assumption Gold evidence annotations in SciFact-Open and Check-COVID are complete and correct
    Recall@k evaluation in Table 1 treats gold documents as the ground truth of evidential relevance; any false negatives in the annotations cap measured recall and would change conclusions.
  • domain assumption Verifier probabilities trained on SciFact abstracts generalize to large open corpora as evidence-utility signals
    Underpins the +Verification experiment and RD.1; the paper adjusts a hyperparameter to make this work but provides no direct transfer validation.
  • ad hoc to paper GPT-4o classification of cited vs self-contained evidence is accurate on 22 papers
    Section 4.2 and Figure 2: the section-location distribution claim rests on unvalidated LLM annotation of a small, convenience-accessible sample; the paper itself notes a misjudged instance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers." pith.science (2026). https://pith.science/paper/IFLAD24X

@misc{pith2026250620844,
  author       = {Pith},
  title        = {Pith review of: The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IFLAD24X}},
  note         = {Machine review of arXiv:2506.20844}
}
read the original abstract

Scientific fact-checking aims to determine the veracity of scientific claims by retrieving and analysing evidence from research literature. The problem is inherently more complex than general fact-checking since it must accommodate the evolving nature of scientific knowledge, the structural complexity of academic literature and the challenges posed by long-form, multimodal scientific expression. However, existing approaches focus on simplified versions of the problem based on small-scale datasets consisting of abstracts rather than full papers, thereby avoiding the distinct challenges associated with processing complete documents. This paper examines the limitations of current scientific fact-checking systems and reveals the many potential features and resources that could be exploited to advance their performance. It identifies key research challenges within evidence retrieval, including (1) evidence-driven retrieval that addresses semantic limitations and topic imbalance (2) time-aware evidence retrieval with citation tracking to mitigate outdated information, (3) structured document parsing to leverage long-range context, (4) handling complex scientific expressions, including tables, figures, and domain-specific terminology and (5) assessing the credibility of scientific literature. Preliminary experiments were conducted to substantiate these challenges and identify potential solutions. This perspective paper aims to advance scientific fact-checking with a specialised IR system tailored for real-world applications.

Figures

Figures reproduced from arXiv: 2506.20844 by the authors.

Figure 1
Figure 1. Pipeline of experiment Approaches. monoT5-3B [64] has demonstrated strong perfor￾mance as a reranker for SciFact, as evidenced by its widespread use as a strong baseline in studies [55, 86] on the BEIR benchmark [89]. It also demonstrated state-of-the-art performance in evidence retrieval for verification [68, 97, 102]. The model assigns a predicted score, 𝑠 𝑟 𝑐,𝑑 , representing the semantic relevance for a document… view at source ↗
Figure 2
Figure 2. Evidence sources in scientific literature. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. An example for parsing the scientific literature [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

128 extracted references · 36 canonical work pages

  1. [1]

    Sahar Abdelnabi, Rakibul Hasan, and Mario Fritz. 2022. Open-domain, content- based, multi-modal fact-checking of out-of-context images via online resources. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition. 14940–14949

  2. [2]

    Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2024. Evaluating Correctness and Faithfulness of Instruction- Following Models for Question Answering. Transactions of the Association for Computational Linguistics 12 (2024), 681–699. doi:10.1162/tacl_a_00667

  3. [3]

    Mubashara Akhtar, Oana Cocarascu, and Elena Simperl. 2022. PubHealthTab: A Public Health Table-based Dataset for Evidence-based Fact Checking. InFindings of the Association for Computational Linguistics: NAACL 2022 , Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.). Association for Computational Linguistics, Seattle, United ...

  4. [4]

    Mubashara Akhtar, Oana Cocarascu, and Elena Simperl. 2023. Reading and Reasoning over Chart Images for Evidence-based Automated Fact-Checking. In Findings of the Association for Computational Linguistics: EACL 2023 , An- dreas Vlachos and Isabelle Augenstein (Eds.). Association for Computational Linguistics, Dubrovnik, Croatia, 399–414. doi:10.18653/v1/20...

  5. [5]

    Mubashara Akhtar, Michael Schlichtkrull, Zhijiang Guo, Oana Cocarascu, Elena Simperl, and Andreas Vlachos. 2023. Multimodal Automated Fact-Checking: A Survey. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Com- putational Linguistics, Singapore, 5430–5448. doi:10....

  6. [6]

    Mubashara Akhtar, Michael Schlichtkrull, and Andreas Vlachos. 2024. Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking. arXiv preprint arXiv:2411.05375 (2024)

  7. [7]

    Mubashara Akhtar, Nikesh Subedi, Vivek Gupta, Sahar Tahmasebi, Oana Co- carascu, and Elena Simperl. 2024. ChartCheck: Explainable Fact-Checking over Real-World Chart Images. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thail...

  8. [8]

    Liesbeth Allein, Marlon Saelens, Ruben Cartuyvels, and Marie-Francine Moens

Show all 128 references
  1. [9]

    Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims. In Proceedings of the 2019 Conference on Empirical Me...

  2. [10]

    Dara Bahri, Yi Tay, Che Zheng, Donald Metzler, and Andrew Tomkins. 2020. Choppy: Cut transformer for ranked list truncation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 1513–1516

  3. [11]

    Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing (EMNLP-IJ...

  4. [12]

    Peters, and Arman Cohan

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer. arXiv:2004.05150 (2020)

  5. [13]

    Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2013. Methods for exploring and mining tables on wikipedia. In Proceedings of the ACM SIGKDD workshop on interactive data exploration and analytics . 18–26

  6. [14]

    Reem Bin-Hezam and Mark Stevenson. 2023. Combining Counting Processes and Classification Improves a Stopping Rule for Technology Assisted Review. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Associ...

  7. [15]

    Reem Bin-Hezam and Mark Stevenson. 2024. RLStop: A Reinforcement Learning Stopping Method for TAR. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2604–2608

  8. [16]

    Jannis Bulian, Jordan Boyd-Graber, Markus Leippold, Massimiliano Ciaramita, and Thomas Diggelmann. 2020. Climate-fever: A dataset for verification of real-world climate claims. InNeurIPS 2020 Workshop on Tackling Climate Change with Machine Learning

  9. [17]

    Peter Buneman, Dennis Dosso, Matteo Lissandrini, and Gianmaria Silvello. 2021. Data citation and the citation graph. Quantitative Science Studies 2, 4 (2021), 1399–1422

  10. [18]

    Rui Cao, Zifeng Ding, Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos

  11. [19]

    Catherine Chen, Zejiang Shen, Dan Klein, Gabriel Stanovsky, Doug Downey, and Kyle Lo. 2023. Are Layout-Infused Language Models Robust to Layout Distribution Shifts? A Case Study with Scientific Documents. In Findings of the Association for Computational Linguistics: ACL 2023 ,...

  12. [20]

    Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17754–17762

  13. [21]

    Qingyu Chen, Yifan Peng, and Zhiyong Lu. 2019. BioSentVec: creating sen- tence embeddings for biomedical texts. In 2019 IEEE International Conference on Healthcare Informatics (ICHI). IEEE, 1–5

  14. [22]

    Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. TabFact : A Large-scale Dataset for Table-based Fact Verification. InInternational Conference on Learning Representations (ICLR). Addis Ababa, Ethiopia

  15. [23]

    Sabrina Chiesurin, Dimitris Dimakopoulos, Marco Antonio Sobrevilla Cabezudo, Arash Eshghi, Ioannis Papaioannou, Verena Rieser, and Ioannis Konstas. 2023. The Dangers of trusting Stochastic Parrots: Faithfulness and Trust in Open- domain Conversational Question Answering. In Fi...

  16. [24]

    Jaemin Cho, Debanjan Mahata, Ozan Irsoy, Yujie He, and Mohit Bansal. 2024. M3docrag: Multi-modal retrieval is what you need for multi-page multi- document understanding. arXiv preprint arXiv:2411.04952 (2024)

  17. [25]

    Gordon V Cormack and Maura R Grossman. 2014. Evaluation of machine- learning protocols for technology-assisted review in electronic discovery. In Proceedings of the 37th international ACM SIGIR conference on Research & devel- opment in information retrieval . 153–162

  18. [26]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021. Overview of the TREC 2020 deep learning track. arXiv:2102.07662 [cs.IR] https://arxiv.org/abs/2102.07662

  19. [27]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees. 2020. Overview of the TREC 2019 deep learning track. arXiv preprint arXiv:2003.07820 (2020)

  20. [28]

    Alphaeus Dmonte, Roland Oruche, Marcos Zampieri, Prasad Calyam, and Is- abelle Augenstein. 2024. Claim Verification in the Age of Large Language Models: A Survey. arXiv preprint arXiv:2408.14317 (2024)

  21. [29]

    Haoyu Dong and Zhiruo Wang. 2024. Large language models for tabular data: Progresses and future directions. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2997– 3000

  22. [30]

    Michael Färber and Adam Jatowt. 2020. Citation recommendation: approaches and datasets. International Journal on Digital Libraries 21, 4 (2020), 375–405

  23. [31]

    Aaron HA Fletcher and Mark Stevenson. 2025. Predicting retracted research: a dataset and machine learning approaches. Research Integrity and Peer Review 10, 1 (2025), 1–10

  24. [32]

    Raymond Fok, Hita Kambhamettu, Luca Soldaini, Jonathan Bragg, Kyle Lo, Marti Hearst, Andrew Head, and Daniel S Weld. 2023. Scim: Intelligent skimming support for scientific papers. In Proceedings of the 28th International Conference on Intelligent User Interfaces . 476–490

  25. [33]

    Maura R Grossman, Gordon V Cormack, and Adam Roegiest. 2016. TREC 2016 Total Recall Track Overview.. In TREC

  26. [34]

    Nianlong Gu, Yingqiang Gao, and Richard HR Hahnloser. 2022. Local citation recommendation with hierarchical-attention text encoder and scibert-based reranking. In European conference on information retrieval . Springer, 274–288

  27. [35]

    Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2020. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. arXiv:arXiv:2007.15779

  28. [36]

    Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, and Xiaoyong Du

  29. [37]

    Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics 10 (2022), 178–206

  30. [38]

    Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly Supervised Table Parsing via Pre-training. In Proceedings of the 58th Annual Meeting of the Association for The Next Phase of Scientific Fact-Checking: Advanc...

  31. [39]

    Jason Hoelscher-Obermaier, Edward Stevinson, Valentin Stauber, Ivaylo Zhelev, Viktor Botev, Ronin Wu, and Jeremy Minton. 2022. Leveraging knowledge graphs to update scientific word embeddings using latent semantic imputation. In Proceedings of the first Workshop on Information...

  32. [40]

    Xuming Hu, Zhaochen Hong, Zhijiang Guo, Lijie Wen, and Philip Yu. 2023. Read it twice: Towards faithfully interpretable fact verification by revisiting evidence. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval ...

  33. [41]

    Po-Wei Huang, Abhinav Ramesh Kashyap, Yanxia Qin, Yajing Yang, and Min-Yen Kan. 2022. Lightweight Contextual Logical Structure Recovery. In Proceedings of the Third Workshop on Scholarly Document Processing , Arman Cohan, Guy Feigenblat, Dayne Freitag, Tirthankar Ghosal, Draho...

  34. [42]

    Tomoki Ikoma and Shigeki Matsubara. 2023. On the Use of Language Models for Function Identification of Citations in Scholarly Papers. InProceedings of the Sec- ond Workshop on Information Extraction from Scientific Publications , Tirthankar Ghosal, Felix Grezes, Thomas Allen, ...

  35. [43]

    Kanoulas, Dan Li, Leif Azzopardi, and René Spijker

    E. Kanoulas, Dan Li, Leif Azzopardi, and René Spijker. 2018. CLEF 2017 Tech- nologically Assisted Reviews in Empirical Medicine Overview. In Conference and Labs of the Evaluation Forum . https://api.semanticscholar.org/CorpusID: 246440739

  36. [44]

    Evangelos Kanoulas, Dan Li, Leif Azzopardi, and Rene Spijker. 2018. CLEF 2018 technologically assisted reviews in empirical medicine overview. In CEUR workshop proceedings, Vol. 2125

  37. [45]

    Evangelos Kanoulas, Dan Li, Leif Azzopardi, and Rene Spijker. 2019. CLEF 2019 technology assisted reviews in empirical medicine overview. In CEUR workshop proceedings, Vol. 2380. 250

  38. [46]

    Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, and Mohit Iyyer

  39. [47]

    Mohammed Abdul Khaliq, Paul Yu-Chun Chang, Mingyang Ma, Bernhard Pflugfelder, and Filip Miletić. 2024. RAGAR, Your Falsehood Radar: RAG- Augmented Reasoning for Political Fact-Checking using Multimodal Large Language Models. In Proceedings of the Seventh Fact Extraction and VE...

  40. [48]

    Jiho Kim, Sungjin Park, Yeonsu Kwon, Yohan Jo, James Thorne, and Edward Choi. 2023. FactKG: Fact Verification via Reasoning on Knowledge Graphs. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jo...

  41. [49]

    Neema Kotonya and Francesca Toni. 2020. Explainable Automated Fact- Checking for Public Health Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for...

  42. [50]

    Xiangci Li, Gully A Burns, and Nanyun Peng. 2021. A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification.. In SDU@ AAAI

  43. [51]

    Xiaoxi Li, Zhicheng Dou, Yujia Zhou, and Fangchao Liu. 2024. Corpuslm: Towards a unified language model on corpus for knowledge-intensive tasks. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 26–37

  44. [52]

    Yen-Chieh Lien, Daniel Cohen, and W Bruce Croft. 2019. An assumption-free approach to the dynamic truncation of ranked lists. In Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval . 79–82

  45. [53]

    Donald AB Lindberg, Betsy L Humphreys, and Alexa T McCray. 1993. The unified medical language system. Yearbook of medical informatics 2, 01 (1993), 41–51

  46. [54]

    Carolyn E Lipscomb. 2000. Medical subject headings (MeSH). Bulletin of the Medical Library Association 88, 3 (2000), 265

  47. [55]

    Qi Liu, Bo Wang, Nan Wang, and Jiaxin Mao. 2025. Leveraging Passage Em- beddings for Efficient Listwise Reranking with Large Language Models. In Pro- ceedings of the ACM on Web Conference 2025 (Sydney NSW, Australia) (WWW ’25). Association for Computing Machinery, New York, NY...

  48. [56]

    Kyle Lo, Joseph Chee Chang, Andrew Head, Jonathan Bragg, Amy X. Zhang, Cas- sidy Trier, Chloe Anastasiades, Tal August, Russell Authur, Danielle Bragg, Erin Bransom, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Yen-Sung Chen, Evie Yu-Yen Cheng, Yvonne Chou, Doug Down...

  49. [57]

    Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. 2023. SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables. In Proceedings of the 2023 Conference on Em- pirical Methods in Natural Language Processing , Hou...

  50. [58]

    María Luengo and David García-Marín. 2020. The performance of truth: politi- cians, fact-checking journalism, and the struggle to tackle COVID-19 misinfor- mation. American Journal of Cultural Sociology 8, 3 (2020), 405

  51. [59]

    Yixiao Ma, Qingyao Ai, Yueyue Wu, Yunqiu Shao, Yiqun Liu, Min Zhang, and Shaoping Ma. 2022. Incorporating retrieval information into the truncation of ranking lists for better legal search. In Proceedings of the 45th International ACM SIGIR Conference on Research and Developme...

  52. [60]

    Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu- Gang Jiang, Jiaqi Wang, Yixin Cao, and Aixin Sun. 2024. MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visuali...

  53. [61]

    Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024. Ranked list truncation for large language model-based re-ranking. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 141–151

  54. [62]

    Isabelle Mohr, Amelie Wührl, and Roman Klinger. 2022. CoVERT: A Corpus of Fact-checked Biomedical COVID-19 Tweets. In Proceedings of the Thirteenth Language Resources and Evaluation Conference , Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher...

  55. [63]

    Yuki Momii, Tetsuya Takiguchi, and Yasuo Ariki. 2024. RAG-Fusion Based Information Retrieval for Fact-Checking. In Proceedings of the Seventh Fact Ex- traction and VERification Workshop (FEVER), Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akh...

  56. [64]

    Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Docu- ment Ranking with a Pretrained Sequence-to-Sequence Model. InFindings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computati...

  57. [65]

    Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis Petrantonakis. 2023. Synthetic misinformers: Generating and com- bating multimodal misinformation. In Proceedings of the 2nd ACM International Workshop on Multimedia AI against Disinformation . 36–44

  58. [66]

    Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis C Petrantonakis. 2024. VERITE: a Robust benchmark for multimodal misinformation detection accounting for unimodal bias. International Journal of Multimedia Information Retrieval 13, 1 (2024), 4

  59. [67]

    or: How I learned to stop worrying and love

    Andrew Parry, Debasis Ganguly, and Manish Chandra. 2024. In-Context Learn- ing" or: How I learned to stop worrying and love" Applied Information Retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 14–25

  60. [68]

    Ronak Pradeep, Xueguang Ma, Rodrigo Nogueira, and Jimmy Lin. 2021. Scientific Claim Verification with VerT5erini. In Proceedings of the 12th International Workshop on Health Text Mining and Information Analysis , Eben Holderness, Antonio Jimeno Yepes, Alberto Lavelli, Anne-Lys...

  61. [69]

    Peng Qi, Zehong Yan, Wynne Hsu, and Mong Li Lee. 2024. SNIFFER: Multi- modal Large Language Model for Explainable Out-of-Context Misinformation Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13052–13062

  62. [70]

    Vatsal Raina and Mark Gales. 2024. Question-Based Retrieval using Atomic Units for Enterprise RAG. In Proceedings of the Seventh Fact Extraction and VER- ification Workshop (FEVER), Michael Schlichtkrull, Yulong Chen, Chenxi White- house, Zhenyun Deng, Mubashara Akhtar, Rami A...

  63. [71]

    Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie. 2024. SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation. arXiv preprint arXiv:2405.08807 (2024)

  64. [72]

    Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. COVID- Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 Pan- demic. In Proceedings of the 59th Annual Meeting of the Association for Com- putational Linguistics and the 11th International Jo...

  65. [73]

    Alireza Salemi and Hamed Zamani. 2024. Evaluating retrieval quality in retrieval- augmented generation. In Proceedings of the 47th International ACM SIGIR Con- ference on Research and Development in Information Retrieval . 2395–2400

  66. [74]

    Alireza Salemi and Hamed Zamani. 2024. Towards a search engine for machines: Unified ranking for multiple retrieval-augmented large language models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 741–751

  67. [75]

    Mourad Sarrouti, Asma Ben Abacha, Yassine Mrabet, and Dina Demner- Fushman. 2021. Evidence-based Fact-Checking of Health-related Claims. In Findings of the Association for Computational Linguistics: EMNLP 2021 , Marie- Francine Moens, Xuanjing Huang, Lucia Specia, and Scott We...

  68. [76]

    Artsiom Sauchuk, James Thorne, Alon Halevy, Nicola Tonellotto, and Fabrizio Silvestri. 2022. On the role of relevance in natural language processing tasks. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1785–1789

  69. [77]

    Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. The Auto- mated Verification of Textual Claims (AVeriTeC) Shared T...

  70. [78]

    Michael Schlichtkrull, Zhijiang Guo, and Andreas Vlachos. 2024. Averitec: A dataset for real-world claim verification with evidence from the web. Advances in Neural Information Processing Systems 36 (2024)

  71. [79]

    Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen- tau Yih, and Sebastian Riedel. 2021. Joint Verification and Reranking for Open Fact Checking Over Tables. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics an...

  72. [80]

    Özge Sevgili, Irina Nikishina, Seid Muhie Yimam, Martin Semmann, and Chris Biemann. 2024. UHH at AVeriTeC: RAG for Fact-Checking with Real- World Claims. In Proceedings of the Seventh Fact Extraction and VERifica- tion Workshop (FEVER) , Michael Schlichtkrull, Yulong Chen, Che...

  73. [81]

    Weld, and Doug Downey

    Zejiang Shen, Kyle Lo, Lucy Lu Wang, Bailey Kuehl, Daniel S. Weld, and Doug Downey. 2022. VILA: Improving Structured Content Extraction from Scien- tific PDFs Using Visual Layout Groups. Transactions of the Association for Computational Linguistics 10 (2022), 376–392. doi:10.1...

  74. [82]

    Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning. PMLR, 31210–31227

  75. [83]

    Qi Shi, Yu Zhang, Qingyu Yin, and Ting Liu. 2021. Logic-level Evidence Retrieval and Graph-based Verification Network for Table-based Fact Verification. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Hu...

  76. [84]

    Ronit Singal, Pransh Patwa, Parth Patwa, Aman Chadha, and Amitava Das

  77. [85]

    Mark Stevenson and Reem Bin-Hezam. 2023. Stopping Methods for Technology- assisted Reviews Based on Point Processes. ACM Transactions on Information Systems 42, 3 (2023), 1–37

  78. [86]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language...

  79. [87]

    Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, and Kuniko Saito. 2023. Slidevqa: A dataset for document visual question answering on multiple images. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 13636–13645

  80. [88]

    Liyan Tang, Philippe Laban, and Greg Durrett. 2024. MiniCheck: Efficient Fact- Checking of LLMs on Grounding Documents. In Proceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Assoc...

  81. [89]

    Evidence-backed Fact Checking using RAG and Few-Shot In-Context Learning with LLMs. In Proceedings of the Seventh Fact Extraction and VERifi- cation Workshop (FEVER), Michael Schlichtkrull, Yulong Chen, Chenxi White- house, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Gu...

  82. [90]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal

  83. [91]

    James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopou- los, and Arpit Mittal. 2018. The Fact Extraction and VERification (FEVER) Shared Task. In Proceedings of the First Workshop on Fact Extraction and VER- ification (FEVER), James Thorne, Andreas Vlachos, Oa...

  84. [92]

    Fangzheng Tian, Debasis Ganguly, and Craig Macdonald. 2025. Is Relevance Propagated from Retriever to Generator in RAG?. In Advances in Information Retrieval: 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy, April 6–10, 2025, Proceedings, Part I (Luc...

  85. [93]

    Image, Tell me your story!

    Jonathan Tonglet, Marie-Francine Moens, and Iryna Gurevych. 2024. “Image, Tell me your story!” Predicting the original meta-context of visual misinfor- mation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Ba...

  86. [94]

    Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Tra...

  87. [95]

    Jordy Van Landeghem, Rubèn Tito, Łukasz Borchmann, Michał Pietruszka, Pawel Joziak, Rafal Powalski, Dawid Jurkiewicz, Mickaël Coustaty, Bertrand Anckaert, Ernest Valveny, et al. 2023. Document understanding dataset and evaluation (dude). In Proceedings of the IEEE/CVF Internat...

  88. [96]

    Vijay Viswanathan, Graham Neubig, and Pengfei Liu. 2021. CitationIE: Lever- aging the Citation Graph for Scientific Information Extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on...

  89. [97]

    Juraj Vladika and Florian Matthes. 2023. Scientific Fact-Checking: A Survey of Resources and Approaches. In Findings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki The Next Phase of Scientific Fact-Checking: Adva...

  90. [98]

    Juraj Vladika and Florian Matthes. 2024. Comparing Knowledge Sources for Open-Domain Scientific Claim Verification. InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , Yvette Graham and Matthew...

  91. [99]

    Juraj Vladika and Florian Matthes. 2024. Improving Health Question Answering with Reliable and Time-Aware Evidence Retrieval. In Findings of the Association for Computational Linguistics: NAACL 2024 , Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association for Computat...

  92. [100]

    David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Sci- entific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webb...

  93. [101]

    Herbert Ullrich, Tomáš Mlynář, and Jan Drchal. 2024. AIC CTU system at AVeriTeC: Re-framing automated fact-checking as a simple RAG task. InProceed- ings of the Seventh Fact Extraction and VERification Workshop (FEVER) , Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Z...

  94. [102]

    David Wadden, Kyle Lo, Lucy Lu Wang, Arman Cohan, Iz Beltagy, and Han- naneh Hajishirzi. 2022. MultiVerS: Improving scientific claim verification with weak supervision and full-document context. In Findings of the Association for Computational Linguistics: NAACL 2022 , Marine ...

  95. [103]

    Dong Wang, Jianxin Li, Tianchen Zhu, Haoyi Zhou, Qishan Zhu, Yuxin Wen, and Hongming Piao. 2022. MtCut: A Multi-Task Framework for Ranked List Truncation. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining . 1054–1062

  96. [104]

    Gengyu Wang, Kate Harwood, Lawrence Chillrud, Amith Ananthram, Melanie Subbiah, and Kathleen McKeown. 2023. Check-COVID: Fact-Checking COVID- 19 News Claims with Scientific Evidence. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan B...

  97. [105]

    Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Rubashevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, Isabelle Augenstein, Iryna Gurevych, and Preslav Nakov. 2024. Factcheck-Bench: Fine-Grained Evaluatio...

  98. [106]

    Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377 (2023)

  99. [107]

    Jevin D West and Carl T Bergstrom. 2021. Misinformation in and about science. Proceedings of the National Academy of Sciences 118, 15 (2021), e1912444117

  100. [108]

    David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Iz Beltagy, Lucy Lu Wang, and Hannaneh Hajishirzi. 2022. SciFact-Open: Towards open-domain scientific claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2022 , Yoav Goldberg, Zornitsa Kozare...

  101. [109]

    Chen Wu, Ruqing Zhang, Jiafeng Guo, Yixing Fan, Yanyan Lan, and Xueqi Cheng. 2021. Learning to truncate ranked lists for information retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 4453–4461

  102. [110]

    Siwei Wu, Yizhi Li, Kang Zhu, Ge Zhang, Yiming Liang, Kaijing Ma, Cheng- hao Xiao, Haoran Zhang, Bohao Yang, Wenhu Chen, Wenhao Huang, Noura Al Moubayed, Jie Fu, and Chenghua Lin. 2024. SciMMIR: Benchmarking Sci- entific Multi-modal Information Retrieval. In Findings of the As...

  103. [111]

    Amelie Wührl and Roman Klinger. 2022. Entity-based Claim Representation Improves Fact-Checking of Medical Content in Tweets. In Proceedings of the 9th Workshop on Argument Mining, Gabriella Lapesa, Jodi Schneider, Yohan Jo, and Sougata Saha (Eds.). International Conference on ...

  104. [112]

    Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang

  105. [113]

    Yunhu Ye, Binyuan Hui, Min Yang, Binhua Li, Fei Huang, and Yongbin Li. 2023. Large language models are versatile decomposers: Decomposing evidence and questions for table-based reasoning. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development...

  106. [114]

    Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. InPro- ceedings of the 58th Annual Meeting of the Association for Computational Linguis- tics, Dan Jurafsky, Joyce Chai, Natalie Schl...

  107. [115]

    Dustin Wright and Isabelle Augenstein. 2021. CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Ass...

  108. [116]

    Xia Zeng, Amani S Abumansour, and Arkaitz Zubiaga. 2021. Automated fact- checking: A survey. Language and Linguistics Compass 15, 10 (2021), e12438

  109. [117]

    Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2023. From Relevance to Utility: Evidence Retrieval with Feedback for Fact Verification. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino,...

  110. [118]

    Tianmai M Zhang and Neil F Abernethy. 2024. Detecting Reference Errors in Sci- entific Literature with Large Language Models. arXiv preprint arXiv:2411.06101 (2024)

  111. [119]

    Zhiwei Zhang, Jiyi Li, Fumiyo Fukumoto, and Yanming Ye. 2021. Abstract, Ra- tionale, Stance: A Joint Model for Scientific Claim Verification. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie- Francine Moens, Xuanjing Huang, Luci...

  112. [120]

    In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

    End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2733– 2743

  113. [123]

    Hamed Zamani and Michael Bendersky. 2024. Stochastic rag: End-to-end retrieval-augmented generation through expected utility maximization. In Pro- ceedings of the 47th International ACM SIGIR Conference on Research and Devel- opment in Information Retrieval . 2641–2646

  114. [128]

    Liwen Zheng, Chaozhuo Li, Xi Zhang, Yu-Ming Shang, Feiran Huang, and Haoran Jia. 2024. Evidence Retrieval is almost All You Need for Fact Verifi- cation. In Findings of the Association for Computational Linguistics ACL 2024 , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds....

  115. [2018]

    FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Pro- ceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Pa- pers), Marilyn Walker, Heng Ji, and Amanda...

  116. [2022]

    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.)

    PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics,...

  117. [2023]

    In Find- ings of the Association for Computational Linguistics: EACL 2023 , Andreas Vla- chos and Isabelle Augenstein (Eds.)

    Implicit Temporal Reasoning for Evidence-Based Fact-Checking. In Find- ings of the Association for Computational Linguistics: EACL 2023 , Andreas Vla- chos and Isabelle Augenstein (Eds.). Association for Computational Linguistics, Dubrovnik, Croatia, 176–189. doi:10.18653/v1/2...

  118. [2024]

    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.)

    One Thousand and One Pairs: A “novel” challenge for long-context language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Mia...

  119. [2025]

    arXiv:2505.17978 [cs.CL] https://arxiv.org/abs/ 2505.17978

    AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web. arXiv:2505.17978 [cs.CL] https://arxiv.org/abs/ 2505.17978

  120. [7864]

    doi:10.18653/v1/2024.emnlp-main.448

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.