{"id":"4cfa4e32-b44a-4b1e-8112-e0cc01acd8d3","arxiv_id":"2508.00109","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The document's abstract and body are two different papers, so the FACTORY benchmark claims are unsupported by the supplied full text.","lead":"The submission's metadata and abstract describe FACTORY, a human-verified prompt set for long-form factuality by Chen et al., but the full text is an unrelated statistics paper, funOCLUST, by Clark and McNicholas. Because the body contains none of the FACTORY materials or results, the abstract's claims cannot be checked from this document.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract and body are different papers (FACTORY vs funOCLUST), so the benchmark's claims rest on absent content; the 40%-vs-10% nonfactual claim rate cannot be checked.","rationale":"The reader's weakest_assumption—that the body text belongs to the FACTORY paper—is exactly the load-bearing concern. The abstract promises a benchmark and experimental results; the body delivers an unrelated functional-data clustering paper. Without the prompt set, annotation protocol, model list, and comparison datasets, there is no way to test the central claim that FACTORY induces roughly 40% nonfactual claims versus about 10% on other datasets. I agree with the UNVERDICTED verdict and LOW confidence: the submission fails at the level of completeness, not merely at the level of a debatable assumption or a contested statistic. If the source is later recovered and the body matches the abstract, a substantive review of the benchmark's human verification and evaluation fairness would be needed. But as submitted, the only honest disposition is 'unverified.'","tokens_in":13802,"tokens_out":1878,"duration_ms":20288,"concrete_test":"Fetch the arXiv source files for 2508.00109 and inspect the main .tex or PDF body. Determine whether the body text after the abstract is genuinely the funOCLUST paper or whether the submitted full text is a corrupted/mismatched artifact. If the body is funOCLUST as submitted, then the FACTORY claims are unsupported and the paper must remain unverifiable; if an actual FACTORY body exists in the source, retrieve it and re-review the benchmark content, especially the human-verification protocol and the 40%-vs-10% comparison methodology.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The submission's title, author list, and abstract describe FACTORY, a human-verified long-form factuality prompt set, but the full text is entirely 'funOCLUST: Clustering Functional Data with Outliers' by Clark and McNicholas. The body contains no prompt set, no human-annotation protocol, no model outputs, no evaluation on six LLMs, and no comparison datasets supporting the stated 'approximately 40% versus 10%' nonfactual claim rates. The abstract's central claim is that FACTORY is a more challenging and more reliable benchmark than existing long-form factuality datasets. For that claim to hold, the submitted document would need to include the FACTORY prompt set, evidence of human verification quality (e.g., annotator agreement or adjudication procedures), and a transparent evaluation setup that justifies the 40%/10% comparison. None of this is present. The body's internal statistical content—e.g., the shifted-and-scaled beta distribution for subset log-likelihoods in funOCLUST Lemma 1—is a complete non sequitur with respect to the abstract. This is a missing-support failure at the submission level: even if the funOCLUST mathematics were entirely correct, it would supply no evidence for any FACTORY claim. The paper therefore cannot be verified as submitted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission, arXiv:2508.00109, presents in its title and abstract a benchmark paper called FACTORY: a large-scale, human-verified prompt set for long-form factuality, claiming that state-of-the-art language models generate about 40% nonfactual claims on FACTORY versus about 10% on existing datasets. The full text, however, is entirely a different manuscript, 'funOCLUST: Clustering Functional Data with Outliers' by Clark and McNicholas, which develops a robust clustering algorithm for functional data and contains no prompt set, no human annotation protocol, no language-model evaluation, and no comparison of nonfactual rates. As a result, every central claim in the abstract is unsupported by the submitted document.","tokens_in":14009,"tokens_out":3882,"duration_ms":36634,"significance":"If the FACTORY claims were substantiated, the contribution would be significant: a human-verified benchmark that is fact-seeking, answerable, and unambiguous, together with evidence that it is substantially more challenging than prior datasets, would be a useful resource for long-form factuality evaluation. The reported 40% versus 10% gap is the kind of falsifiable, head-to-head comparison that would matter to the field. However, none of the supporting artifacts or analyses appear in the manuscript: the prompt set, verification protocol, annotator agreement, model list, response outputs, and dataset comparison are all absent. The body's funOCLUST material is mathematically unrelated and cannot substitute for the missing evidence. The submission therefore cannot be verified in its current form.","major_comments":[{"comment":"The document's title and abstract describe FACTORY, a human-verified long-form factuality prompt set, but the full text is entirely the funOCLUST paper on functional-data clustering. None of the benchmark claims in the abstract—neither the construction of the prompt set, nor the human-verification procedure, nor the evaluation on six language models—appears anywhere in the body. This mismatch removes the evidentiary basis for every central claim of the submission.","section":"Abstract vs. Full text"},{"comment":"The abstract claims the prompt set is 'large-scale, human-verified' and developed 'model-in-the-loop' with human refinement, but the manuscript contains no prompt set, no annotation instructions, no annotator-agreement statistics, and no adjudication protocol. Without these materials, the reliability claim that is the paper's stated motivation cannot be assessed.","section":"Abstract; missing benchmark materials"},{"comment":"The headline quantitative result, that SOTA models produce approximately 40% nonfactual claims on FACTORY versus 10% on other datasets, is stated with no supporting experiments: no model names or versions, no generation settings, no response samples, no factuality annotation scheme, and no description of the comparison datasets. This is load-bearing because the paper's contribution is precisely the challenge gap and reliability relative to existing benchmarks.","section":"Abstract; missing evaluation"},{"comment":"The only technical content in the body is the funOCLUST extension. Lemma 1's proof delegates to Theorem 1 of Clark and McNicholas (2024), and Remark 1 explicitly conditions the derived distribution on a finite Gaussian mixture with no outlying coefficients. These are internal limitations of the statistical contribution, but they do not bear on the FACTORY claims; even a fully corrected funOCLUST paper would not supply the missing benchmark evidence.","section":"§3.2, Lemma 1 and Remark 1"}],"minor_comments":[{"comment":"The body contains encoding artifacts such as the header 'R AMi`Q/m+iBQM' and rot13-like section headings, which make the document difficult to read and suggest a conversion error.","section":"Throughout"},{"comment":"The author list and references in the body belong to the funOCLUST paper; the name FACTORY never appears after the abstract. The submission metadata should be reconciled with the actual content.","section":"Metadata"},{"comment":"The abstract says '6 state-of-the-art language models' but names none of them; if the correct manuscript is resubmitted, the models, versions, and prompts should be specified.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"The submitted document appears to be a packaging error rather than a deliberate scientific claim: the abstract describes one paper and the body is an unrelated paper. I recommend the editor return the submission to the authors to verify the uploaded file, and if the intended FACTORY manuscript is provided, treat this report as a preliminary assessment requiring a new review cycle."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi colleague,\n\nQuick take: this submission is broken. The title page and abstract describe FACTORY, a human-verified prompt set for long-form factuality, with the claim that state-of-the-art models produce roughly 40% nonfactual claims on it versus about 10% on existing datasets. The full text is an unrelated statistics paper, funOCLUST, by Clark and McNicholas, about clustering functional data with outliers. None of the FACTORY content appears in the body. As submitted, there is no benchmark, no prompts, no annotation protocol, no model evaluations. The 40% versus 10% number cannot be checked. This is a load-bearing missing-content failure, not a minor quibble.\n\nWhat is actually new and better: the funOCLUST body itself looks like a genuine and reasonably careful extension of the OCLUST algorithm to functional data via B-spline coefficients, with a derived shifted-and-scaled beta distribution for subset log-likelihoods, simulations against five competitors, and two real-data examples. The math appears to follow from Theorem 1 of Clark and McNicholas (2024), and the paper is transparent about inheriting those assumptions. If this were sent to me as a standalone statistics manuscript, I would call it a modest but sound contribution, likely worth a referee. But it is not the paper the abstract advertises.\n\nSoft spots: besides the mismatch, the body's own evaluation reclassifies outliers before computing ARI, which flatters the method; the 51% false-negative rate in one simulation scenario is honestly reported but shows limits. These are secondary concerns for a separate review.\n\nMy recommendation: do not send this to peer review. Desk reject or return to the authors with instructions to resubmit the correct full text. If the FACTORY paper is real, this version gives no basis to evaluate it. The citation pattern in the body is fine for a statistics paper; that is not the issue.\n\nWe cannot judge the thinking behind FACTORY because the content is absent. This is not a paper; it is an abstract stapled to someone else's manuscript.","headline":"The abstract describes a factuality benchmark, the body is a functional-data clustering paper; the submission is not a reviewable paper.","tokens_in":14580,"tokens_out":2452,"would_cite":false,"duration_ms":23523,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper announces FACTORY, a human-verified prompt set for long-form factuality, and claims that state-of-the-art language models make about 40% nonfactual claims on it compared with about 10% on existing benchmarks.","keywords":["long-form factuality","benchmark","human verification","prompt set","language models","claim-level evaluation","nonfactual claims","long-tailed facts"],"falsifier":"Open the submitted artifact and verify whether the body contains FACTORY's prompt set and human evaluation; it does not, which is an immediate negative check. If the missing materials are restored, the decisive test is independent human re-annotation of a random sample of FACTORY prompts: if many prompts are unanswerable or ambiguous, or a fresh run yields near-10% nonfactual claims, the central claim is false.","tokens_in":13577,"feed_emoji":"🎯","tokens_out":7133,"duration_ms":64366,"temperature":0.7,"pith_summary":"The paper announces FACTORY, a large-scale, human-verified set of prompts for evaluating long-form factuality. The claim is that existing benchmarks lack human verification and therefore overstate model reliability: on FACTORY, state-of-the-art language models produce about 40% nonfactual claims, compared with about 10% on other datasets. The prompts are intended to be fact-seeking, answerable, and unambiguous, and to force models to reason across long-tailed facts. A sympathetic reader should care because, if true, the familiar ten-percent error numbers would be too optimistic and the field would need a harder, more reliable evaluation. The full text supplied with this submission is a different paper on functional-data clustering, so FACTORY's methods and evidence are not present in the body.","feed_headline":"FACTORY benchmark: 40% of AI claims are not factual","feed_subtitle":"Human-verified long-form factuality benchmark reports about four times more model errors than prior datasets.","key_machinery":"The object that carries the abstract's argument is FACTORY itself: a prompt set built with a model-in-the-loop approach and refined by humans, designed so every prompt is fact-seeking, answerable, and unambiguous. The measuring instrument is claim-level human evaluation, which decomposes each model response into claims and labels each claim factual or not, with the headline number being the share of nonfactual claims. The supplied body text, however, makes no use of this machinery; its machinery is funOCLUST, a two-stage algorithm that decomposes curves with cubic B-splines and then runs an outlier-trimming Gaussian mixture clusterer on the coefficients, using a shifted-and-scaled beta distribution for subset log-likelihoods to decide when to stop trimming. For the abstract's argument, the load-bearing machinery is the human-verified prompt set and the human evaluation protocol, neither of which appears in the supplied full text.","core_discovery":"On the paper's own terms, the central discovery is that long-form factuality can be measured more honestly with a human-verified prompt set: when six state-of-the-art models are evaluated claim by claim by humans, roughly 40% of their claims on FACTORY are not factual, against roughly 10% on existing datasets. The intended cause is prompt quality, because FACTORY prompts are selected to be fact-seeking, answerable, and unambiguous and to draw on long-tailed facts that models cannot answer by rote. The supplied body text does not contain the FACTORY study; it describes funOCLUST, a clustering algorithm for functional data, so this discovery is asserted in the abstract but not demonstrated in the manuscript.","pith_inferences":["If the 40% versus 10% gap survives independent audit, it would suggest that earlier benchmarks are saturated: reported advances may mostly reflect easier prompt distributions, prompting a re-reading of past comparisons.","The human-verified design implies that automatic factuality metrics that match claims to retrieval sources may miss long-tail falsehoods, so hybrid human-and-automatic claim adjudication is a natural next measurement to test.","A controlled test of the benchmark's contribution would build an unverified version of FACTORY with the same prompts but no human filtering and compare error rates; a large drop would isolate human verification as the driver.","Because the submitted body is a different paper, the immediate extension is procedural: reconcile the artifact so the benchmark's prompts, annotation instructions, and model outputs are publicly inspectable."],"forward_implications":["If FACTORY's 40% figure holds, current long-form factuality results showing near-10% nonfactual rates are optimistic by a factor of about four.","A human-verified, answerable, unambiguous prompt set gives model developers a more reliable signal for where factual generation fails.","The benchmark's emphasis on long-tailed facts implies that progress on FACTORY would require models to retrieve or reason about uncommon knowledge, not just popular topics.","Because the comparison is claim-level, differences between models on FACTORY would be attributable to factual accuracy rather than to prompt ambiguity or unanswerability."],"supporting_citations":[],"fun_headline_variants":["FACTORY: 40% of AI claims are false","AI models flunk long-form factuality test: 40% errors","Human-verified prompts expose 40% AI claims as untrue","FACTORY: 4 in 10 AI claims aren't factual"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that human verification really makes FACTORY prompts fact-seeking, answerable, and unambiguous and that the 40% versus 10% comparison was measured under the same protocol; the supplied body text is a different paper, so neither premise can be checked from this manuscript.","fun_headline_variants_meta":{"raw":{"variants":["FACTORY: 40% of AI claims are false","AI models flunk long-form factuality test: 40% errors","Human-verified prompts expose 40% AI claims as untrue","FACTORY: 4 in 10 AI claims aren't factual"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000873,"raw_usage":{"total_tokens":3721,"prompt_tokens":830,"completion_tokens":2891,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":2816}},"tokens_in":446,"tokens_out":2891,"duration_ms":19963,"temperature":1.0,"reasoning_tokens":2816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:22:04.981815+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the submitted artifact and verify whether the body contains FACTORY's prompt set and human evaluation; it does not, which is an immediate negative check. If the missing materials are restored, the decisive test is independent human re-annotation of a random sample of FACTORY prompts: if many prompts are unanswerable or ambiguous, or a fresh run yields near-10% nonfactual claims, the central claim is false.","supporting_citations":[],"review_version":1}