{"id":"4dc0ee7f-8035-410a-a62e-3f7f1fe68c74","arxiv_id":"2605.30241","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"CommunityFact provides a new dynamic benchmark showing web access improves LLM misinformation detection but source selection remains misaligned with human Community Notes raters.","lead":"The paper introduces CommunityFact, a refreshable benchmark containing 15,992 claims across five languages and two domains to test misinformation detection. A smart generalist might read it to see how current LLMs perform on real-world multilingual fact-checking and how their web-search behavior differs from human raters.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Human Community Notes convergence treated as reliable reference for LLM misalignment lacks validation against external accuracy criteria","rationale":"The reader's weakest_assumption directly identifies the load-bearing assumption for the misalignment result. No other technical detail (dataset construction, retrieval mechanics, or language-domain slicing) can be evaluated as more critical without first securing the validity of the human reference. Full-text methods might contain such a validation, but none is visible from the supplied abstract or reader summary.","tokens_in":1678,"tokens_out":319,"duration_ms":17505,"concrete_test":"Sample 200 claims from the benchmark; obtain independent expert annotations (3+ domain specialists per claim) of the top-3 most relevant primary sources; compute overlap between expert sources and Community Notes converged sources versus overlap with LLM-retrieved sources; if expert–human overlap is not reliably higher than expert–LLM overlap, the reference standard is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim rests on measuring 'systematic misalignment' between web-enabled LLM source-selection policies and the sources on which human Community Notes raters converge. This comparison is only interpretable as a performance gap if the converged human sources constitute a trustworthy external standard (i.e., they are more accurate, complete, or authoritative than the alternatives LLMs retrieve). The paper provides no independent check that convergence reflects factual superiority rather than rater-pool composition, claim visibility, or consensus dynamics; without that, the misalignment metric risks conflating deviation from crowd behavior with deviation from truth.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents CommunityFact, a refreshable benchmark with 15,992 standalone claims across five languages and two domains, designed to evaluate LLMs on misinformation detection in dynamic, multilingual, real-world settings. It reports evaluations of ten LLMs under closed-input, thinking, and web-search conditions, claiming that closed-input verification is challenging, web access yields the largest performance gains, web-enabled LLMs exhibit systematic misalignment in source-selection policies relative to sources converged upon by human Community Notes raters (with the gap closable via model-specific retrieval expansion or pruning), and that substantial variation exists across language-domain slices and evidence ecosystems. The work further positions the benchmark as a potential training signal for claim-conditioned source suggesters.","tokens_in":1802,"tokens_out":614,"duration_ms":15127,"significance":"If the core empirical findings on web-access gains and source misalignment hold after addressing validation gaps, the benchmark could meaningfully advance evaluation practices for LLMs in fast-moving, multilingual misinformation contexts and open avenues for using Community Notes data in training. The dynamic and redistributable design addresses limitations of static benchmarks, and the multilingual/multi-domain coverage is a strength. However, the interpretive weight placed on human rater convergence as a reference standard without external accuracy validation limits the immediate significance of the misalignment claims.","major_comments":[{"comment":"Abstract (and implied results section): the claim that web-enabled LLMs' source-selection policies are 'systematically misaligned' with sources on which human Community Notes raters converge treats rater convergence as a reliable external reference standard for measuring misalignment. No independent validation (e.g., against external fact-checking accuracy, expert adjudication, or ground-truth labels) is described to establish that converged sources are more accurate, complete, or authoritative than LLM-retrieved alternatives; without this, the misalignment metric risks conflating deviation from crowd consensus with deviation from truth and is load-bearing for the performance-gap interpretation.","section":"Abstract"},{"comment":"Abstract: the reported results on closed-input verification difficulty, web-access gains, and cross-slice variation lack any description of the underlying statistical methods, confidence intervals, error analysis, or controls for claim difficulty, making it impossible to assess whether the stated patterns are robust or driven by the specific 15,992-claim sample.","section":"Abstract"}],"minor_comments":[{"comment":"The manuscript provides no details on claim sourcing process, annotation guidelines, exclusion criteria, inter-rater agreement for the human Community Notes convergence, or how the benchmark ensures redistributability while remaining dynamic.","section":null},{"comment":"No information is given on the ten LLMs evaluated (specific models, sizes, or versions), the exact prompting or retrieval setups for web-enabled conditions, or how 'thinking' and 'web-search' capabilities were operationalized.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive comments. We address each major comment below, indicating planned revisions where appropriate.","responses":[{"response":"The manuscript frames the source-selection analysis as a comparison against the specific sources on which human Community Notes raters converge, consistent with CommunityFact's construction from Community Notes data as a real-world human signal. We do not claim or provide evidence that these converged sources are more accurate, complete, or authoritative than LLM-selected alternatives, nor do we equate human consensus with ground truth. The reported misalignment describes a policy difference, and the performance gap refers to deviation from observed human rater behavior. We will revise the abstract and relevant sections to explicitly qualify the reference standard as human rater convergence without external accuracy validation, thereby avoiding any implication that deviation equals deviation from truth.","revision_made":"partial","referee_comment":"[Abstract] Abstract (and implied results section): the claim that web-enabled LLMs' source-selection policies are 'systematically misaligned' with sources on which human Community Notes raters converge treats rater convergence as a reliable external reference standard for measuring misalignment. No independent validation (e.g., against external fact-checking accuracy, expert adjudication, or ground-truth labels) is described to establish that converged sources are more accurate, complete, or authoritative than LLM-retrieved alternatives; without this, the misalignment metric risks conflating deviation from crowd consensus with deviation from truth and is load-bearing for the performance-gap interpretation."},{"response":"We agree that the abstract does not describe the statistical methods, confidence intervals, error analysis, or controls. The full manuscript reports the core quantitative results, but to improve transparency we will revise the abstract to include a concise description of the statistical approaches (including confidence intervals and any controls for claim difficulty, language, and domain) and will ensure the results section provides explicit error analysis and robustness details.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the reported results on closed-input verification difficulty, web-access gains, and cross-slice variation lack any description of the underlying statistical methods, confidence intervals, error analysis, or controls for claim difficulty, making it impossible to assess whether the stated patterns are robust or driven by the specific 15,992-claim sample."}],"tokens_in":1434,"tokens_out":491,"duration_ms":29612,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main contribution is CommunityFact itself: nearly 16k standalone claims across five languages and two domains, built to be refreshable rather than static. That addresses a real gap in misinformation evaluation, where most existing sets age quickly and stay monolingual. The evaluation of ten LLMs under closed, thinking, and web-search conditions is straightforward and shows the expected pattern that web access helps most while closed-input verification stays hard. Variation across language-domain slices is also reported, which is worth having on record.\n\nThe headline result on source-selection misalignment is the weaker part. The paper measures how web-enabled models pick sources differently from the ones human Community Notes raters converge on, then calls the difference systematic misalignment. This only works as a performance gap if the converged human sources are more accurate or complete than the alternatives. The abstract gives no sign of an independent check against external accuracy criteria, ground-truth labels, or expert review. Without that, the metric risks measuring deviation from crowd behavior rather than deviation from truth. The stress-test note correctly flags this as the load-bearing assumption.\n\nThe work is aimed at researchers who build or evaluate misinformation detectors and want a larger, more current testbed. Readers focused on benchmark construction will find the coverage and redistributability goals concrete. The suggestion to treat Community Notes as a training signal for source suggesters is a reasonable downstream idea, though it inherits the same reference-standard issue.\n\nIt deserves peer review. The benchmark data and scale are new enough to justify referee time, and the empirical LLM results can be assessed directly even if the misalignment interpretation needs tightening or additional validation.","headline":"CommunityFact adds a useful new multilingual benchmark with real scale, but its claim of systematic LLM misalignment with Community Notes rests on an unvalidated assumption that rater convergence tracks accuracy.","tokens_in":2288,"tokens_out":407,"would_cite":false,"duration_ms":14551,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Web-enabled LLMs select sources for verifying claims that differ systematically from the sources human Community Notes raters converge on.","keywords":["misinformation detection","benchmark","large language models","Community Notes","multilingual evaluation","web search","source selection"],"falsifier":"A fresh collection of claims on which web-enabled LLMs that apply retrieval expansion or pruning still select sources that diverge from the sources human raters converge on.","tokens_in":2591,"feed_emoji":"🔍","tokens_out":645,"duration_ms":17108,"temperature":0.7,"pith_summary":"The paper presents CommunityFact as a refreshable benchmark containing 15,992 claims in five languages and two domains to test how well models handle real-world misinformation. It runs ten LLMs through closed-input, thinking, and web-search conditions and tracks both accuracy and the specific sources each model retrieves. The main result is that web access produces the largest accuracy gains, yet the sources chosen by these models diverge from the sources that human raters settle on in Community Notes. The divergence shrinks when models apply their own retrieval-expansion or pruning rules. The benchmark is also positioned to supply training data for systems that suggest sources conditioned on a new claim.","feed_headline":"LLMs pick different sources than humans for fact-checking","feed_subtitle":"Benchmark of 16k claims shows web-enabled models diverge from Community Notes raters, with gaps reduced by retrieval tweaks.","key_machinery":"The CommunityFact benchmark, which pairs each claim with the sources converged upon by human Community Notes raters to quantify misalignment in LLM source-selection policies.","core_discovery":"CommunityFact supplies a dynamic collection of 15,992 standalone claims across five languages and two domains. Evaluation of ten LLMs shows closed-input verification remains hard, web access yields the largest gains, and web-enabled models' source-selection policies are systematically misaligned with the sources that human Community Notes raters converge upon; this gap narrows through model-specific retrieval expansion or pruning. Substantial variation appears across language-domain slices and across the evidence ecosystems used by different web-enabled systems. The resource further frames Community Notes data as a training signal for claim-conditioned source suggesters.","pith_inferences":["Models trained on these alignments could generalize source suggestions to claims outside the current benchmark.","Dynamic updates to the benchmark would let researchers track whether retrieval policies improve or drift over time.","The observed language-domain variation suggests targeted retrieval tuning per slice may be needed before broad deployment."],"forward_implications":["Closed-input verification stays difficult across the tested models.","Web access produces larger accuracy gains than internal thinking alone.","Source-selection misalignment varies by language and domain slice.","Different web-enabled systems draw from distinct evidence ecosystems.","Community Notes data can train claim-conditioned source suggesters for novel claims."],"fun_headline_variants":["LLMs diverge from humans in source selection for fact checks","CommunityFact reveals models pick different sources than raters","Web models misalign with Community Notes on evidence sources","Benchmark shows LLM source policies differ from human convergence"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Agreement among human Community Notes raters supplies a reliable external reference standard for judging whether an LLM has selected appropriate sources.","fun_headline_variants_meta":{"raw":{"variants":["LLMs diverge from humans in source selection for fact checks","CommunityFact reveals models pick different sources than raters","Web models misalign with Community Notes on evidence sources","Benchmark shows LLM source policies differ from human convergence"]},"model":"grok-4.3","cost_usd":0.006661,"raw_usage":{"total_tokens":3101,"prompt_tokens":658,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":66612000,"prompt_tokens_details":{"text_tokens":658,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2382,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":658,"tokens_out":61,"duration_ms":17697,"temperature":1.0,"reasoning_tokens":2382,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T07:30:51.690889+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A fresh collection of claims on which web-enabled LLMs that apply retrieval expansion or pruning still select sources that diverge from the sources human raters converge on.","supporting_citations":[],"review_version":1}