Pith. sign in

REVIEW 3 major objections 5 minor 79 references

CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Current large language models cannot reliably fact-check Chinese misinformation on their own, even with chain-of-thought and few-shot prompting, but they can improve human fact-checking accuracy when deployed as assistants.

desk verdict Solid dataset, sloppy stats; the assistive-potential claim doesn't survive its own table. read the letter →

arxiv 2509.03957 v1 pith:ORWKCJAM submitted 2025-09-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords Chinesemisinformationfact-checkingbenchmarklargelanguagemodelshallucinationtaxonomycontamination-freeevaluationhuman-LLMcollaborationchain-of-thoughtpromptingrumordetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds CANDY, a benchmark for testing how well large language models verify Chinese news claims, using about twenty thousand claims with official refutation evidence plus thousands of labeled explanations. It tries to establish that today's models are not reliable enough to run fact-checking autonomously: even the strongest model tested gets only about 76% of contamination-free claims right, and reasoning prompts do not fix this. The paper argues that the main cause is a failure mode it calls factual fabrication, where a model invents specific, confident-sounding details to support a wrong verdict. It then claims the same models are useful in a different role, as assistants: in its human study, people at four education levels became more accurate when they could consult an LLM than when working alone or with web search alone. The contribution is a measurement instrument and a diagnosis: a dataset, an error taxonomy, and evidence about where human oversight still belongs.

What carries the argument

The engine of the benchmark is CANDYSET, a dataset of roughly 20,000 Chinese claims, each with gold-evidence explanations drawn from official refutation platforms, publication dates, and domain labels, split by each model's knowledge cutoff so that contamination-free evaluation is possible. Around it sits a seven-category taxonomy of flawed LLM explanations—faithfulness hallucination, factuality hallucination, and reasoning inadequacy, with factual fabrication as the largest subcategory—that turns a binary verdict into an analyzable failure. The third piece is the human study, which measures the same task as a cooperative rather than autonomous activity: people fact-check alone, with web sea

What would settle it

A pre-registered replication of the human study with dozens of participants per condition, or a contamination-free test that withholds official refutation text, would settle the claims: if Human+Web matches or beats Human+LLM, the assistive claim fails; if models score well above 90% on withheld post-cutoff claims, the unreliability claim weakens.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that LLM fact-checking of Chinese misinformation fails in a specific, diagnosable way: models do not merely return wrong verdicts, they often produce fluent justifications by fabricating evidence. Among roughly five thousand manually annotated flawed explanations, factual fabrication is the largest category, and flawed explanations overwhelmingly accompany wrong verdicts. The best model in the benchmark reaches only 76.2% accuracy on contamination-free claims, and chain-of-thought plus few-shot prompting does not close the gap. Against that negative result, the paper's human study claims a positive one: LLMs can raise human fact-checking acc

Load-bearing premise

The assistive-benefit claim rests on a human study in which only three people carried out each condition; if those few participants are not representative, the advantage of LLM assistance over web search could be a fluke.

Editorial extensions

If this is right

  • Deploying current LLMs as standalone Chinese fact-checkers is not yet safe, particularly for breaking, time-sensitive events outside their training cutoff.
  • Additional reasoning prompts and few-shot examples do not automatically make fact-checking more accurate, and for some smaller models they increase overconfidence.
  • Factual fabrication—models inventing authoritative-sounding details to support a verdict—is the failure mode to target before relying on LLM explanations.
  • The same models can improve human accuracy when framed as assistants, and the gain appears across elementary, middle-school, undergraduate, and master's-level participants.
  • Future evaluations should measure explanation quality, not just final verdicts, because wrong explanations are the clearest signal of untrustworthy conclusions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's interrogation experiment suggests a cheap, testable intervention: reformulating a claim as a question cut factual fabrication sharply in its case study, implying prompt framing could act as a practical guardrail against sycophantic verification.
  • If the contamination-free split is sound, the same design could expose how much of an LLM's apparent knowledge is memorized content; extending it to other languages would test whether fabrication rates are structural rather than specific to Chinese.
  • The taxonomy could support an automatic detector for fabricated explanations, letting systems flag low-confidence verdicts for human review instead of presenting them as final.
  • The assistive benefit may be largest for less-experienced fact-checkers, but because the human study used only three participants per condition, real deployment should re-test with larger samples before treating the gain as robust.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces CANDY, a benchmark for evaluating LLMs on Chinese misinformation fact-checking. It presents CANDYSET, containing ~20,000 claims with gold evidence, 4,891 manually annotated flawed LLM-generated explanations, and a 140-item human study; a seven-category taxonomy of flawed explanations; and evaluations of sixteen LLMs and three LRMs under zero-shot/few-shot and CoT variants. The paper reports that current LLMs are unreliable for fully automated fact-checking--the best model reaches 76.2% contamination-free accuracy--and that LLM-generated explanations frequently contain factual fabrication. It further claims that LLM assistance improves human fact-checking accuracy across educational levels.

Significance. If the claims hold, CANDY would be a valuable resource for Chinese misinformation research: it is one of the first benchmarks to combine multi-domain claims with gold evidence, contamination-free temporal splits, a fine-grained explanation-error taxonomy, and manually annotated explanations. The authors have released the dataset and code, and the annotation process shows substantial inter-annotator agreement (Fleiss' kappa 0.76). The main empirical finding on LLM limitations is supported by a broad model and prompt sweep. However, the assistive-potential half of the central claim rests on a human study with three participants per condition and no inferential statistics, which is not sufficient to establish the stated practical benefits.

major comments (3)
  1. [Section 6] The opening paragraph states that 'a substantial proportion (91.2%) of fact-checking results leading to incorrect conclusions were associated with flawed LLM-generated explanations' and that 'only 8.8% of the results leading to correct conclusions exhibited such flaws.' These percentages are computed from the 4,891 flawed explanations (4,463 wrong + 428 correct), i.e., P(wrong | flawed) and P(correct | flawed), not P(flawed | wrong) or P(flawed | correct). Without the total numbers of correct and incorrect conclusions in the 22,000-explanation sample, the numbers do not support the claim that flawed explanations drive incorrect conclusion. Please recompute with the appropriate denominators or rephrase as conditional-on-flawed rates.
  2. [Section 7 / Table 7] The human study uses n=3 participants per condition and education level, and the reported percentages are not accompanied by confidence intervals or significance tests. For proportions near 0.5-0.8, a three-participant cell carries a standard error on the order of 15-25 percentage points, so the 5-10 point differences between conditions are within sampling noise. The text's statement that 'Human+LLM consistently outperforms Human+Web across all groups' is also contradicted by Table 7: for Undergraduate/After Nov 2023, Human+Web scores 79.7 vs. Human+LLM 78.7; for Master's/Before Nov 2023, Human+Web scores 85.7 vs. Human+LLM 85.0. Because the 140 items are drawn from the same platform and the web-augmented LLM condition can retrieve the gold evidence, the comparison is further confounded. This evidence is insufficient for the abstract's 'considerable potential to augment human performance
  3. [Section 5.2] The text claims GPT-4o exceeds the average Life-domain accuracy by 18.82%. According to Table 6, GPT-4o's Life accuracy is 82.19 and the average is 70.47, a difference of 11.72 percentage points. Please correct the number and avoid the word 'significantly' without a supporting statistical test.
minor comments (5)
  1. [Section 4] 'Fact-Checing Explanation' should be 'Fact-Checking Explanation' in the task enumeration.
  2. [Appendix C.2] The closed-source model list enumerates only seven models (GPT-4o, GPT-4-Turbo, GPT-3.5-Turbo, Gemini-1.5-pro, Baichuan4-Turbo, ChatGLM4, Yi-large), while the text and Table 4 state eight closed-source models; DeepSeek-v3 is missing from the enumeration.
  3. [Appendix D.3 / Table 6] Table 6 includes DeepSeek-V3, DeepSeek-R1, and Qwen-QwQ-Plus in domain-level rows, but Appendix D.3 states that these models are excluded from some domain tables because of sparse samples. Please clarify which evaluation mode Table 6 reports and reconcile the model inclusion.
  4. [Section 7 / Appendix C.1] The human-study instructions do not specify the number of interaction turns, time limits, or whether the interaction logs were recorded. Reporting these details would improve reproducibility.
  5. [Appendix E] The prompt examples contain a duplicated 'Rationales:' placeholder in the Few-shot w/ CoT prompt; please proofread.

Circularity Check

1 steps flagged · score 2.0 of 10

No derivational circularity in the benchmark claims; one mild model-in-the-loop labeling concern involving GPT-4o.

  1. other [Appendix B.2 (Data Normalization) and Section 5.1 (Overall Evaluation)]
    "we use GPT-4o for initial data preprocessing and labeling, followed by manual verification. ... As shown in Table 4, the GPT-4o emerged as the top-performing model, which may underscores its robust utilization of extensive internal knowledge."

    The same model, GPT-4o, was used to help produce the gold labels for CANDYSET and is then evaluated on those labels. Any agreement between GPT-4o and the labels is partly an agreement with its own prior outputs, so the reported evaluation is not a fully external measurement of that model. Manual verification breaks the direct identity, so the top performance is not forced by construction; the circularity is mild and does not affect the paper's other central claims (the taxonomy distribution or the human-study assistive results).

full rationale

The paper is an empirical benchmark study, not a derivational construction, so most possible circularity patterns do not apply. The taxonomy of flawed explanations is pre-defined (Section 3.2) and the frequencies of error types come from manual annotation of 4,891 explanations; the finding that factual fabrication is the most common failure mode is an empirical result, not an artifact of the taxonomy's definitions. The central conclusion that LLMs are unreliable for fully automated Chinese fact-checking rests on accuracy/F1 numbers across 19 models, and the contamination-free split by model cutoff is a reasonable design rather than a circular step. The human-study claim is statistically fragile (3 participants per condition, no significance testing) and Table 7 does not support the text's 'consistently outperforms' statement, but underpowered statistics and internal contradictions are correctness risks, not circularity. The cited prior works by the authors (e.g., Huang et al. 2025 for human-LLM cooperation, Yang et al. 2025 for averaging prompts) are used for framing or methodology and are not load-bearing for the main empirical conclusions. The only mild circularity is that GPT-4o was used in dataset normalization and later evaluated as a top model; manual verification and the use of authoritative platform evidence mitigate this, so it does not rise to a score above 2.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No fitted parameters; the benchmark is empirical. The only numeric choices are design decisions (e.g., 140 items, 3 participants per cell, 200 claims for case studies), which are not fitted to data. No new theoretical entities are introduced; CANDYSET is a dataset, and the taxonomy categories are annotation labels, not physical or theoretical constructs.

assumptions (5)
  • domain assumption Gold evidence and labels from Chinese rumor-refutation platforms are accurate and representative of real-world misinformation.
    All accuracy scores depend on the correctness of these labels; only a 3% sample was manually re-verified (Appendix B.4), and GPT-4o was used for initial labeling (Appendix B.2).
  • domain assumption Model knowledge cut-off dates accurately bound training data; claims published after the cutoff are unseen.
    The contamination-free evaluation (Section 5) relies on this; if models were trained on later data, the reported gap is understated (Table 4).
  • ad hoc to paper The 7-category taxonomy is a valid and exhaustive way to classify flawed explanations.
    The taxonomy is introduced by the authors (Section 3.2); Fleiss Kappa 0.76 shows agreement but not exclusivity or completeness.
  • ad hoc to paper The 140-item human study measures fact-checking ability reliably and participants are representative of their education level.
    Only 3 participants per condition (Appendix C.1); no test-retest reliability or power analysis.
  • ad hoc to paper Augmented claims (via entity replacement or negation) retain valid gold evidence.
    Appendix B.3 describes modifying claims without stating that gold evidence was updated; if not, some labels are wrong.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking." pith.science (2026). https://pith.science/paper/ORWKCJAM

@misc{pith2026250903957,
  author       = {Pith},
  title        = {Pith review of: CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ORWKCJAM}},
  note         = {Machine review of arXiv:2509.03957}
}
read the original abstract

The effectiveness of large language models (LLMs) to fact-check misinformation remains uncertain, despite their growing use. To this end, we present CANDY, a benchmark designed to systematically evaluate the capabilities and limitations of LLMs in fact-checking Chinese misinformation. Specifically, we curate a carefully annotated dataset of ~20k instances. Our analysis shows that current LLMs exhibit limitations in generating accurate fact-checking conclusions, even when enhanced with chain-of-thought reasoning and few-shot prompting. To understand these limitations, we develop a taxonomy to categorize flawed LLM-generated explanations for their conclusions and identify factual fabrication as the most common failure mode. Although LLMs alone are unreliable for fact-checking, our findings indicate their considerable potential to augment human performance when deployed as assistive tools in scenarios. Our dataset and code can be accessed at https://github.com/SCUNLP/CANDY

Figures

Figures reproduced from arXiv: 2509.03957 by the authors.

Figure 1
Figure 1. Fact-checking accuracy when handling au [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Distribution of flawed explanations in contam [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Distributions of flawed LLM-generated explanations based on our taxonomy (value statistics in Figure [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Specific examples for understanding taxonomy. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Data gathering pipeline. our data gathering pipeline includes 3 steps: 1) Data collection and pre-processing. [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Decision Tree for Annotation (Baidu)5 ; (3) human judgment assisted by a large language model (GPT-4o); and (4) human judg￾ment assisted by GPT-4o with web-augmented re￾trieval. Detailed prompts for each condition can be found in Appendix E. The LLM experiment plat￾for…
Figure 8
Figure 8. Figure 8: The influence of claim framing strategies on [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]
Figure 9
Figure 9. Figure 9: Specific value statistics on flawed LLM-generated explanations based on our taxonomy. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: overall distributions of flawed LLM-generated explanations. [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: The influence of claim framing strategies on fact-checking outputs. (In English: Fig. [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

79 extracted references · 73 canonical work pages

  1. [1]

    AI, :, Alex Young, Bei Chen, Chao Li, Chen- gen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, and 13 others. 2024. Yi: Open foundation models by 01.ai. Preprint, arXiv:2403.04652. Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Ch...

  2. [2]

    Official institutions have empha- sized the incident, and it has been widely covered by authoritative media such as CCTV and BBC

    Response generation. 3) Human annotation. Platform English Name Link Count 中国互联网联合辟谣平台 China Internet United Rumor Refutation Platform https://www.piyao.org.cn/ 2172 新华社 Xinhua News Agency https://www.xinhuanet.com/ 1255 科普中国 Science Popularization China https://www.kepuchina.cn/ 595 央视新闻 CCTV News https://news.cctv.com/ 497 人民网科普 People’s Daily Online Sc...

  3. [3]

    Few-shot w/o CoT (Dong et al., 2022a), where LLMs are given a few examples to guide their conclusions

  4. [4]

    In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3362–3376

    Chef: A pilot chinese dataset for evidence- based fact-checking. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3362–3376. Associa- tion for Computational Linguistics. Chen Huang, Yang Deng, Wenqiang Lei, Jiancheng Lv, Tat-Seng Chua, and Jimmy Huang. ...

  5. [5]

    LTCR: Long-Text Chinese Rumor Detection Dataset

    Ltcr: Long-text chinese rumor detection dataset. Preprint, arXiv:2306.07201. Preslav Nakov, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barrón-Cedeño, Paolo Papotti, Shaden Shaar, and Giovanni Da San Martino

  6. [7]

    From Chaos to Clarity: Claim Normalization to Empower Fact-Checking

    From chaos to clarity: Claim normaliza- tion to empower fact-checking. arXiv preprint arXiv:2310.14338. Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, and 1 others. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv pre...

  7. [9]

    In Interna- tional AAAI Conference on Web and Social Media (ICWSM)

    Know it to defeat it: Exploring health ru- mor characteristics and debunking efforts on chi- nese social media during covid-19 crisis. In Interna- tional AAAI Conference on Web and Social Media (ICWSM). Xinwei Yang, Zhaofeng Liu, Chen Huang, Jiashuai Zhang, Tong Zhang, Yifan Zhang, and Wenqiang Lei

  8. [12]

    Zero-shot w/o CoT , where LLMs are prompted to directly draw conclusions

Show all 79 references
  1. [13]

    Zero-shot w/ CoT (Wei et al., 2022), where LLMs first perform a factual analysis, explain- ing their reasoning before making a conclu- sion

  2. [15]

    claim":

    Few-shot w/ CoT (Dong et al., 2022b), where LLMs, after analyzing examples of misinfor- mation, provide conclusions along with expla- nations. For the fact-checking explanation task, we use a few-shot CoT approach, prompting the LLM to generate five distinct explanatory ration...

  3. [19]

    claim":

    工伤认定不仅涉及劳动者权益保护,也是企业社会责任的重要体现,有助于维护社会稳定和谐。 结论:非谣言 输出格式如下: 理由: 结论: 现在,我将提供一个新的事件信息,请你根据以上格式给出结论和分析。 事件信息: "claim": "claim", "publish_date": "publish_date", Few-shot and w CoT Prompt for Fact-Checking Conclusion (English) You are an extremely strict fact-checking expert. You wil...

  4. [20]

    根据多项研究,人类的肝脏和肾脏自然具有排毒功能,额外摄入竹炭并未得到临床验证的排毒效果。 3.竹炭食品可能会影响营养物质的吸收,例如维生素和矿物质,也可能带来健康风险。

  5. [21]

    一些国家和地区对竹炭的食品使用有法律限制,需遵循相关法规。

  6. [22]

    claim":

    目前没有权威机构发布证实竹炭食品对美容或健康有益的声明。 用户输入: "claim": " 没签劳动合同的职工受伤后可以申请工伤认定。", "publish_date": "2023-12-5" 结论:谣言 回复: 理由: 1.根据《工伤保险条例》,职工与用人单位建立劳动关系后,即使未签订书面劳动合同,在工作时间和工作 场所内因工作原因受到事故伤害的,应认定为工伤。 2.未签订劳动合同的职工在发生工伤时,可通过提供其他证明材料(如工资支付凭证、工作证等)来证明存 在劳动关系,进而申请工伤认定。 3.劳动和社会保障部门负责工伤认定工作,会根据实际情况...

  7. [23]

    claim":

    工伤认定不仅涉及劳动者权益保护,也是企业社会责任的重要体现,有助于维护社会稳定和谐。 结论:非谣言 输出格式如下: 理由: 结论: 现在,我将提供一个新的事件信息,请你根据以上格式给出结论和分析。 事件信息: "claim": "claim", "publish_date": "publish_date", Prompt for Fact-Checking Explanation (English) You are an extremely strict fact-checking expert. You will receive event in...

  8. [24]

    please determine which ones are true and which are rumors by yourself

    Independent Human Judgment: Below are 140 news claims. please determine which ones are true and which are rumors by yourself

  9. [25]

    Please use the Baidu search engine to determine which ones are true and which are rumors

    Human Judgment Assisted by Internet Search(Baidu): Below are 140 news claims. Please use the Baidu search engine to determine which ones are true and which are rumors

  10. [26]

    Please interact with a large language model to determine which ones are true and which are rumors

    Human Judgment Assisted by a Large Language Model (GPT-4o): Below are 140 news claims. Please interact with a large language model to determine which ones are true and which are rumors

  11. [27]

    China’s V oice

    Human Judgment Assisted by LLM (GPT-4o) with Web-augmented Retrieval: Below are 140 news claims. Please interact with a large language model with web-augmented retrieval to determine which ones are true and which are rumors. Instruction for Human Study (Chinese) 独立判断: 以下是140条新...

  12. [28]

    Antibiotics are specifically used to treat bacterial infections, while COVID-19 is caused by a coronavirus, making antibiotics ineffective against viral infections.,

  13. [29]

    Both the World Health Organization (WHO) and the Centers for Disease Control and Prevention (CDC) have explicitly stated that antibiotics should not be used to prevent or treat COVID-19.,

  14. [30]

    Overuse of antibiotics can lead to increased antibiotic resistance, which poses a serious threat to public health.,

  15. [31]

    Treating COVID-19 requires the use of antiviral medications and supportive care based on the specific condition, rather than antibiotics.,

  16. [32]

    Ya Ya’s" stay in the United States was for 15 years., publish_date: 2023/4/24, label: 1, source: null, gold evidence:

    Only in cases where bacterial complications occur might doctors consider the use of antibiotics, based on the specific situation.], label_p: 1, Error Type: C:Context Inconsistency id: 462, claim: "Ya Ya’s" stay in the United States was for 15 years., publish_date: 2023/4/24, l...

  17. [33]

    Ya Ya" is a giant panda born in China. In April 2003,

    "Ya Ya" is a giant panda born in China. In April 2003, "Ya Ya" and another panda, "Le Le," were sent to the Memphis Zoo in Tennessee, USA, under a 10-year loan agreement.,

  18. [34]

    Ya Ya" and

    According to the agreement between China and the United States, after the initial 10-year term expired in 2013, the loan agreement for "Ya Ya" and "Le Le" was renewed for another 10 years, extending their stay until April 7, 2023, making the total stay 15 years.,

  19. [35]

    This agreement’s details were officially confirmed by the Chinese Association of Wildlife Conservation and the Memphis Zoo.,

  20. [36]

    Ya Ya" and

    During "Ya Ya" and "Le Le’s" stay in the United States, both countries collaborated on panda conservation and research, showcasing mutual efforts in wildlife protection.,

  21. [37]

    Le Le" sadly passed away, raising concerns about

    In early 2023, "Le Le" sadly passed away, raising concerns about "Ya Ya’s" health, which eventually led to the decision to bring "Ya Ya" back to China earlier than planned. However, this did not alter the originally agreed 15-year term.], label_p: 0, Error Type: B:Logical Inco...

  22. [38]

    抗生素专门用于治疗细菌感染,而COVID-19是由冠状病毒引起的,抗生素对病毒感染无效。,

  23. [39]

    世界卫生组织 (WHO)和美国疾病控制与预防中心(CDC)都明确指出,抗生素不能用于预防或治 疗COVID-19。,

  24. [40]

    过度使用抗生素可能导致抗生素耐药性增加,这对公共健康构成威胁。,

  25. [41]

    丫丫”在美国的旅居期限为15年。, publish_date: 2023/4/24, label: 1, source: null, gold evidence: “丫丫

    治疗COVID-19需根据具体病情使用抗病毒药物和支持性治疗,而非抗生素。, 5.只有在出现细菌感染并发症时,医生才可能根据具体情况考虑使用抗生素治疗。], label_p: 1, Error Type: C:Context Inconsistency id: 462, claim: “丫丫”在美国的旅居期限为15年。, publish_date: 2023/4/24, label: 1, source: null, gold evidence: “丫丫”旅居美国协议期限为10年,2013年到期后又延长了10年。, domain: 社会, mode...

  26. [42]

    丫丫”是一只在中国出生的大熊猫,于2003年4月与另一只大熊猫“乐乐

    “丫丫”是一只在中国出生的大熊猫,于2003年4月与另一只大熊猫“乐乐”一起被送往美国田纳西州孟菲斯动 物园,开始了为期10年的租借期。,

  27. [43]

    根据中美双方协议,2013年租借期满后,两国再次签署了延长协议,将“丫丫”和“乐乐”的旅美期限延长 至2023年4月7日,总计15年。,

  28. [44]

    这一协议的内容得到了中国野生动物保护协会和美国孟菲斯动物园的官方确认。,

  29. [45]

    在“丫丫”和“乐乐”的旅美期间,两国在大熊猫保护与科研合作方面进行了多次交流与合作,体现了双方在 野生动物保护领域的共同努力。,

  30. [46]

    乐乐”不幸去世,引发了关于“丫丫

    2023 年初,“乐乐”不幸去世,引发了关于“丫丫”健康状况的关注,最终促成了“丫丫”提前返回中国的决 定,但这并不改变原定的15年旅居期限的事实。], label_p: 0, Error Type: B:Logical Inconsistency English Example of Factuality Hallucination id: 199, claim: A criminal case occurred in Jilin Province, where a butcher killed 12 people due to a land d...

  31. [47]

    A criminal case indeed occurred in Jilin Province, and it has drawn widespread local attention.,

  32. [48]

    According to the police report, the suspect developed resentment due to a land dispute, leading to this tragedy.,

  33. [49]

    The number of victims was 12, a figure confirmed by the authorities.,

  34. [50]

    The suspect was a butcher, a detail also confirmed by the authorities.,

  35. [51]

    The case is under further investigation, and the responsible parties will be punished according to the law.], label_p: 0, Error Type: D:Factual Fabrication id: 1394, claim: Lunar volcanic activity stopped 3 billion years ago., publish_date: 2023/1/20, label: 1, source: Science...

  36. [52]

    Scientists have found through the analysis of lunar rock samples brought back by the Apollo missions that volcanic activity on the moon indeed peaked around 3 billion years ago and then gradually declined.,

  37. [53]

    So far, there is no conclusive evidence suggesting that the moon had active volcanic activity after 3 billion years ago.,

  38. [54]

    Observations and explorations of the lunar surface, including the density and distribution of impact craters, support the view that lunar volcanic activity ceased billions of years ago.,

  39. [55]

    Compared to Earth, the moon’s geological activity is very weak, and its rapid internal heat loss led to an earlier end to volcanic activity.,

  40. [56]

    Research by NASA and other international space agencies supports the presence of ancient lunar volcanic activity, which had essentially ended by 3 billion years ago.], label_p: 0, Error Type: E:Factual Inconsistency Chinese Example of Factuality Hallucination id: 199, claim:吉林...

  41. [57]

    被害人数为12人,这一数字得到了官方确认。,

  42. [58]

    犯罪嫌疑人为一名屠夫,这一点也得到了官方证实。,

  43. [59]

    此案件正在进一步调查中,相关责任人将依法受到严惩。], label_p: 0, Error Type: D:Factual Fabrication id: 1394, claim: 月球的火山活动在 30 亿年前就停止了。, publish_date: 2023/1/20, label: 1, source: 科普中国, gold evidence:发现月球的火山活动其实一直持续到了 20.3亿年前,并不是之前认为的 30 亿年前就死了。, domain: 科学, model: Baichuan4-Turbo, analysis: [

  44. [60]

    科学家们通过分析阿波罗任务带回的月球岩石样本发现,月球上的火山活动确实在约30亿年前达到了高 峰,之后逐渐减少。,

  45. [61]

    迄今为止,尚未有确凿证据表明月球在30亿年后仍有活跃的火山活动。, 3.通过对月球表面的观测和探测,包括撞击坑的密度和分布等特征,支持月球火山活动在数十亿年前就已经 停止的观点。,

  46. [62]

    月球的地质活动与地球相比非常微弱,其内部热量散失快,导致火山活动较早结束。,

  47. [63]

    美国宇航局(NASA)和其他国际空间机构 的研究均支持月球在远古时期存在火山活动,而这些活动 在30亿年前已经基本结束。], label_p: 0, Error Type: E:Factual Inconsistency English Example of Reasoning Inadequacy id: 2128, claim: Smelling sachets or essential oils can help improve reduced sense of smell after COVID-19 recovery., publish_...

  48. [64]

    Smelling sachets or essential oils lacks scientific evidence as a treatment for reduced sense of smell.,

  49. [65]

    Reduced sense of smell may indicate underlying pathological issues and requires professional diagnosis.,

  50. [66]

    Home remedies like sachets or essential oils might delay necessary medical treatment.,

  51. [67]

    Professional medical evaluation is critical for addressing symptoms effectively.,

  52. [68]

    Unverified remedies should be approached with caution to avoid potential health risks.], label_p: 1, Error Type: G:Under Informativeness id: 28, claim: Wuxi police recently arrested human traffickers., publish_date: 2024/10/15, label: 1, source: Chongqing Rumor Refutation, gol...

  53. [69]

    Wuxi police have a history of combating human trafficking and solving related cases.,

  54. [70]

    Combating human trafficking is a key priority for China’s security agencies.,

  55. [71]

    China’s Ministry of Public Security organizes nationwide operations against human trafficking.,

  56. [72]

    Media and police frequently report on human trafficking arrests, including in the Wuxi region.,

  57. [73]

    Human trafficking is a global problem, and China has implemented effective measures to address it.], label_p: 0, Error Type: F:Overgeneralized Reasoning Chinese Example of Reasoning Inadequacy id: 2128, claim: 新冠康复后嗅觉减退可以通过闻香囊、精油等锻炼。, publish_date: 2023/1/18, label: 0, source:...

  58. [74]

    闻香囊、精油等并不能缓解嗅觉减退的症状。,

  59. [75]

    嗅觉减退是一种病理状态,通常由多种原因引起,如感冒、鼻窦炎、脑部疾病等,需要及时就医进行诊断 和治疗。,

  60. [76]

    闻香囊、精油等虽然在一定程度上可以帮助人们放松情绪,但并没有科学证据表明它们可以治疗或改善嗅 觉减退症状。,

  61. [77]

    嗅觉减退通常需要通过专业的医学检查和治疗来解决,自行使用闻香囊、精油等方法不仅无效,还可能延 误病情。,

  62. [78]

    在出现嗅觉减退等健康问题时,应及时就医,接受专业医生的诊断和治疗,避免轻信未经科学验证的偏方 或谣言。], label_p: 1, Error Type: G:Under Informativeness id: 28, claim: 巫溪警方最近抓捕了人贩子。, publish_date: 2024/10/15, label: 1, source: 重庆辟谣, gold evidence: 经巫溪警方核实,当地近期并未发生贩卖儿童案件,也未组织抓捕人贩子行动,网传“抓人贩 子”系谣言。, domain: 社会, model: GPT-4o, analysis: [

  63. [79]

    巫溪警方在多次行动中积极打击人口贩卖犯罪,过去也曾多次成功破获相关案件。,

  64. [80]

    中国公安机关一直将打击人口贩卖作为重要任务,并定期通报相关案件的抓捕情况。,

  65. [81]

    根据中国公安部的公开信息,各地警方定期开展专项行动以打击人口拐卖犯罪。, 4.媒体报道和警方通报中经常会有关于人口贩子的抓捕新闻,巫溪地区也不例外。,

  66. [82]

    人口贩卖是全球性问题,中国政府和警方对此采取了积极措施,并取得了一定成效。], label_p: 0, Error Type: F:Overgeneralized Reasoning

  67. [2021]

    arXiv preprint arXiv:2103.07769

    Automated fact-checking for assisting human fact-checkers. arXiv preprint arXiv:2103.07769. Q. Nan, J. Cao, Y . Zhu, Y . Wang, and J. Li. 2021. Md- fend: Multi-domain fake news detection. In Pro- ceedings of the 30th ACM International Conference on Information & Knowledge Mana...

  68. [2022]

    arXiv preprint arXiv:2210.13865

    Missing counter-evidence renders nlp fact- checking unrealistic for misinformation. arXiv preprint arXiv:2210.13865. Jian Guan, Jesse Dodge, David Wadden, Minlie Huang, and Hao Peng. 2023. Language models hallucinate, but may excel at fact verification. arXiv preprint arXiv:23...

  69. [2023]

    In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing, pages 530–543, Singapore

    Selectively answering ambiguous questions. In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing, pages 530–543, Singapore. Association for Computational Linguistics. Yang Deng, Lizi Liao, Liang Chen, Hongru Wang, Wenqiang Lei, and Tat-Sen...

  70. [2024]

    arXiv preprint arXiv:2407.02351

    Generative large language models in auto- mated fact-checking: A survey. arXiv preprint arXiv:2407.02351. Binjie Wang, Ethan Chern, and Pengfei Liu. 2023. Chi- nesefacteval: A factuality benchmark for chinese llms. Technical report, GAIR-NLP. Bo Wang, Jing Ma, Hongzhan Lin, Zh...

  71. [2025]

    Daily Popular Science

    ELABORATION: A comprehensive bench- mark on human-LLM competitive programming. In Proceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 59–104, Vienna, Austria. Asso- ciation for Computational Linguistics. T. Y...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.