Pith. sign in

REVIEW 4 major objections 3 minor 1 cited by

FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality

T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper announces FACTORY, a human-verified prompt set for long-form factuality, and claims that state-of-the-art language models make about 40% nonfactual claims on it compared with about 10% on existing benchmarks.

desk verdict The abstract describes a factuality benchmark, the body is a functional-data clustering paper; the submission is not a reviewable paper. read the letter →

arxiv 2508.00109 v1 pith:6ZWAAYZG submitted 2025-07-31 cs.CL cs.AI

classification cs.CLcs.AI
keywords long-formfactualitybenchmarkhumanverificationpromptsetlanguagemodelsclaim-levelevaluationnonfactualclaimslong-tailedfacts
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper announces FACTORY, a large-scale, human-verified set of prompts for evaluating long-form factuality. The claim is that existing benchmarks lack human verification and therefore overstate model reliability: on FACTORY, state-of-the-art language models produce about 40% nonfactual claims, compared with about 10% on other datasets. The prompts are intended to be fact-seeking, answerable, and unambiguous, and to force models to reason across long-tailed facts. A sympathetic reader should care because, if true, the familiar ten-percent error numbers would be too optimistic and the field would need a harder, more reliable evaluation. The full text supplied with this submission is a different paper on functional-data clustering, so FACTORY's methods and evidence are not present in the body.

What carries the argument

The object that carries the abstract's argument is FACTORY itself: a prompt set built with a model-in-the-loop approach and refined by humans, designed so every prompt is fact-seeking, answerable, and unambiguous. The measuring instrument is claim-level human evaluation, which decomposes each model response into claims and labels each claim factual or not, with the headline number being the share of nonfactual claims. The supplied body text, however, makes no use of this machinery; its machinery is funOCLUST, a two-stage algorithm that decomposes curves with cubic B-splines and then runs an outlier-trimming Gaussian mixture clusterer on the coefficients, using a shifted-and-scaled beta distribution for subset log-likelihoods to decide when to stop trimming. For the abstract's argument, the load-bearing machinery is the human-verified prompt set and the human evaluation protocol, neither of which appears in the supplied full text.

What would settle it

Open the submitted artifact and verify whether the body contains FACTORY's prompt set and human evaluation; it does not, which is an immediate negative check. If the missing materials are restored, the decisive test is independent human re-annotation of a random sample of FACTORY prompts: if many prompts are unanswerable or ambiguous, or a fresh run yields near-10% nonfactual claims, the central claim is false.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that long-form factuality can be measured more honestly with a human-verified prompt set: when six state-of-the-art models are evaluated claim by claim by humans, roughly 40% of their claims on FACTORY are not factual, against roughly 10% on existing datasets. The intended cause is prompt quality, because FACTORY prompts are selected to be fact-seeking, answerable, and unambiguous and to draw on long-tailed facts that models cannot answer by rote. The supplied body text does not contain the FACTORY study; it describes funOCLUST, a clustering algorithm for functional data, so this discovery is asserted in the abstract but not demonstrated in the manuscript.

Load-bearing premise

The load-bearing premise is that human verification really makes FACTORY prompts fact-seeking, answerable, and unambiguous and that the 40% versus 10% comparison was measured under the same protocol; the supplied body text is a different paper, so neither premise can be checked from this manuscript.

Editorial extensions

If this is right

  • If FACTORY's 40% figure holds, current long-form factuality results showing near-10% nonfactual rates are optimistic by a factor of about four.
  • A human-verified, answerable, unambiguous prompt set gives model developers a more reliable signal for where factual generation fails.
  • The benchmark's emphasis on long-tailed facts implies that progress on FACTORY would require models to retrieve or reason about uncommon knowledge, not just popular topics.
  • Because the comparison is claim-level, differences between models on FACTORY would be attributable to factual accuracy rather than to prompt ambiguity or unanswerability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 40% versus 10% gap survives independent audit, it would suggest that earlier benchmarks are saturated: reported advances may mostly reflect easier prompt distributions, prompting a re-reading of past comparisons.
  • The human-verified design implies that automatic factuality metrics that match claims to retrieval sources may miss long-tail falsehoods, so hybrid human-and-automatic claim adjudication is a natural next measurement to test.
  • A controlled test of the benchmark's contribution would build an unverified version of FACTORY with the same prompts but no human filtering and compare error rates; a large drop would isolate human verification as the driver.
  • Because the submitted body is a different paper, the immediate extension is procedural: reconcile the artifact so the benchmark's prompts, annotation instructions, and model outputs are publicly inspectable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The submission, arXiv:2508.00109, presents in its title and abstract a benchmark paper called FACTORY: a large-scale, human-verified prompt set for long-form factuality, claiming that state-of-the-art language models generate about 40% nonfactual claims on FACTORY versus about 10% on existing datasets. The full text, however, is entirely a different manuscript, 'funOCLUST: Clustering Functional Data with Outliers' by Clark and McNicholas, which develops a robust clustering algorithm for functional data and contains no prompt set, no human annotation protocol, no language-model evaluation, and no comparison of nonfactual rates. As a result, every central claim in the abstract is unsupported by the submitted document.

Significance. If the FACTORY claims were substantiated, the contribution would be significant: a human-verified benchmark that is fact-seeking, answerable, and unambiguous, together with evidence that it is substantially more challenging than prior datasets, would be a useful resource for long-form factuality evaluation. The reported 40% versus 10% gap is the kind of falsifiable, head-to-head comparison that would matter to the field. However, none of the supporting artifacts or analyses appear in the manuscript: the prompt set, verification protocol, annotator agreement, model list, response outputs, and dataset comparison are all absent. The body's funOCLUST material is mathematically unrelated and cannot substitute for the missing evidence. The submission therefore cannot be verified in its current form.

major comments (4)
  1. [Abstract vs. Full text] The document's title and abstract describe FACTORY, a human-verified long-form factuality prompt set, but the full text is entirely the funOCLUST paper on functional-data clustering. None of the benchmark claims in the abstract—neither the construction of the prompt set, nor the human-verification procedure, nor the evaluation on six language models—appears anywhere in the body. This mismatch removes the evidentiary basis for every central claim of the submission.
  2. [Abstract; missing benchmark materials] The abstract claims the prompt set is 'large-scale, human-verified' and developed 'model-in-the-loop' with human refinement, but the manuscript contains no prompt set, no annotation instructions, no annotator-agreement statistics, and no adjudication protocol. Without these materials, the reliability claim that is the paper's stated motivation cannot be assessed.
  3. [Abstract; missing evaluation] The headline quantitative result, that SOTA models produce approximately 40% nonfactual claims on FACTORY versus 10% on other datasets, is stated with no supporting experiments: no model names or versions, no generation settings, no response samples, no factuality annotation scheme, and no description of the comparison datasets. This is load-bearing because the paper's contribution is precisely the challenge gap and reliability relative to existing benchmarks.
  4. [§3.2, Lemma 1 and Remark 1] The only technical content in the body is the funOCLUST extension. Lemma 1's proof delegates to Theorem 1 of Clark and McNicholas (2024), and Remark 1 explicitly conditions the derived distribution on a finite Gaussian mixture with no outlying coefficients. These are internal limitations of the statistical contribution, but they do not bear on the FACTORY claims; even a fully corrected funOCLUST paper would not supply the missing benchmark evidence.
minor comments (3)
  1. [Throughout] The body contains encoding artifacts such as the header 'R AMi`Q/m+iBQM' and rot13-like section headings, which make the document difficult to read and suggest a conversion error.
  2. [Metadata] The author list and references in the body belong to the funOCLUST paper; the name FACTORY never appears after the abstract. The submission metadata should be reconciled with the actual content.
  3. [Abstract] The abstract says '6 state-of-the-art language models' but names none of them; if the correct manuscript is resubmitted, the models, versions, and prompts should be specified.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation in the funOCLUST body; the FACTORY abstract is unsupported by the text, but that is a missing-support issue, not circular reasoning.

full rationale

The submitted full text is funOCLUST, a functional-data clustering paper, not the FACTORY benchmark described in the abstract. The abstract's claims about human verification, 40% vs 10% nonfactual claim rates, and six LLMs cannot be checked because the body contains no prompt set, annotation protocol, or model evaluation; this is an integrity/verifiability failure, not a circular-derivation failure. Within funOCLUST, Lemma 1 explicitly extends Theorem 1 of Clark and McNicholas (2024) by substituting B-spline fitted coefficients, Sigma_g + sigma^2(B^T B)^-1 for Sigma_g, and p = K + 4. That is a legitimate corollary of a published theorem, not a result whose conclusion is assumed in its premises. The outlier-trimming rule compares subset log-likelihood differences to a derived null distribution, which is a goodness-of-fit target rather than a fitted parameter renamed as a prediction. The simulation study, main-effect analysis, and real-data examples (Melbourne pedestrian data, Barcelona NOx data) provide external, independent evidence for the method's behavior. No step in the derivation chain reduces by construction to its own inputs.

Assumptions & free parameters 0 free parameters · 2 assumptions · 1 invented entities

For the stated FACTORY claim, the ledger is nearly empty because the supporting material is missing. The two axioms are the abstract's implicit assumptions about human labeling and prompt quality. The body's funOCLUST parameters and model choices do not bear on the FACTORY claim.

assumptions (2)
  • domain assumption Human verification provides a valid and consistent ground truth for whether a claim is factual.
    The abstract's reliability claim depends on human annotations being correct; no annotation protocol, sample size, or inter-annotator agreement is supplied.
  • domain assumption Model-in-the-loop generation followed by human refinement produces prompts that are fact-seeking, answerable, and unambiguous.
    These properties are asserted in the abstract but no prompt-construction details or quality checks are in the supplied document.
invented entities (1)
  • FACTORY prompt set
    purpose: Challenging human-verified benchmark for long-form factuality
    The prompt set is claimed in the abstract but absent from the body, so it has no verifiable existence in this submission.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality." pith.science (2026). https://pith.science/paper/6ZWAAYZG

@misc{pith2026250800109,
  author       = {Pith},
  title        = {Pith review of: FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6ZWAAYZG}},
  note         = {Machine review of arXiv:2508.00109}
}
read the original abstract

Long-form factuality evaluation assesses the ability of models to generate accurate, comprehensive responses to short prompts. Existing benchmarks often lack human verification, leading to potential quality issues. To address this limitation, we introduce FACTORY, a large-scale, human-verified prompt set. Developed using a model-in-the-loop approach and refined by humans, FACTORY includes challenging prompts that are fact-seeking, answerable, and unambiguous. We conduct human evaluations on 6 state-of-the-art language models using FACTORY and existing datasets. Our results show that FACTORY is a challenging benchmark: approximately 40% of the claims made in the responses of SOTA models are not factual, compared to only 10% for other datasets. Our analysis identifies the strengths of FACTORY over prior benchmarks, emphasizing its reliability and the necessity for models to reason across long-tailed facts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

    cs.CL 2026-01 conditional novelty 5.0 of 10

    On a new extended needle-in-a-haystack benchmark, explicit anti-hallucination prompts and dispersed fact placement cause some long-context LLMs to over-refuse or collapse in accuracy, while others remain robust.

Reference graph

Works this paper leans on

14 extracted references · 14 canonical work pages · cited by 1 Pith paper

  1. [1]

    Clark 1 and Paul D

    funOCLUST: Clustering Functional Data with Outliers Katharine M. Clark 1 and Paul D. McNicholas 2 1Department of Mathematics & Statistics, Trent University, Ontario, Canada. 2Department of Mathematics & Statistics, McMaster University, Ontario, Canada. Abstract Functional data present unique challenges for clustering due to their infinite-dimensional natu...

  2. [3]

    2 kXk 6mM+iBQM H .2+QKTQbBiBQM In multivariate data analysis, a random sample ( x1; : : :xn) is typically a collection of vectors in Rp

    fbeta 2nh (nh 1)2 (dj c) p 2 ; nh p 1 2 (2) for c < d j < (nh1)2 2nh + c; nh > p + 1, where c = log ^ h + p 2 log(2 ) + 1 2 logjShj, nh is the number of points in cluster h, ^ h = nh=n, Sh = 1 nh 1 nX i=1 zih(xi xh)(xi xh)> is the sample covariance matrix of cluster h, and xh = 1 nh Pn i=1 zihxi: The trimmed model chosen minimizes the Kullback-Leibler (KL...

  3. [4]

    i = B(t) i +

    is a collection of isolation trees which randomly split the functional space into partitions, iteratively, until each function is isolated. Functions isolated in fewer iterations tend to be more outlying. An outlier score is calculated using a forest of these isolation trees. Finally, Hubert et al. (2015) determine functional outlyingness by first calcula...

  4. [5]

    As in Clark and McNicholas (2024), these assumptions can be relaxed in practice

    f#2i 2nh (nh 1)2 (dj c) K + 4 2 ; nh K 5 2 (7) 7Q` c < d j < (nh1)2 2nh + c; nh > K + 5- r?2`2 c = log ^ h + K+4 2 log(2 ) + 1 2 logjShj- nh Bb i?2 MmK#2` Q7 TQBMib BM +Hmbi2` h- ^ h = nh=n- Sh = 1 nh 1 nX i=1 zih(bi bh)(bi bh)> Bb i?2 b KTH2 +Qp `B M+2 K i`Bt Q7 +Hmbi2` h- bh = 1 nh Pn i=1 zihbi; M/ K Bb i?2 MmK#2` Q7 BMi2`BQ` FMQib BM i?2 +m#B+ bTHBM2 #...

  5. [6]

    This transformation carries over to the functional representation: B(t)b = B(t)(a + c) = a (t) + B(t)c; 7 where (t) = B(t) denotes the mean function

    Thus, any nontrivial transformation of increases the MSD. This transformation carries over to the functional representation: B(t)b = B(t)(a + c) = a (t) + B(t)c; 7 where (t) = B(t) denotes the mean function. Hence, deviations of b from correspond directly to deviations of the associated function from the mean function. Because lXnb lX / DM (b ); observati...

  6. [7]

    Additionally, these outliers often exhibit unusual combinations of traffic patterns

    Many of the outlying observations have higher-than-average traffic in the afternoon (noon to 5pm) or early morning (midnight to 4am). Additionally, these outliers often exhibit unusual combinations of traffic patterns. For example, traffic volume around noon and 7pm may resemble a typical workday, while the volume around 4pm is more similar to what is observ...

  7. [11]

    (North) sensor are plotted in Figure 6, with observations coloured according to whether the data correspond to a workday or non-workday (i.e., a weekend or public holiday)

    Data captured for the year 2017 from the Chinatown–Swanston St. (North) sensor are plotted in Figure 6, with observations coloured according to whether the data correspond to a workday or non-workday (i.e., a weekend or public holiday). There are 365 observations, each with values collected at 24 time points. Of the 365 days captured, 249 are workdays and...

  8. [1982]

    outliers

    as the distance option. 9Xk *Hmbi2`BM; _2bmHib Each clustering algorithm is evaluated using the adjusted Rand index (ARI; Hubert and Arabie, 1985). The ARI compares two partitions — in this case, real and predicted classes. 11 The ARI equals one under perfect classification and has expected value zero under random class assignment. For funOCLUST and tkmea...

Show all 14 references
  1. [2005]

    Solid black curves indicate the group-wise hourly mean

    Curves are coloured according to whether each day is a workday (W) or non-workday (NW). Solid black curves indicate the group-wise hourly mean. 0.51 to 0.86 places its best-performing models on par with existing approaches while demon- strating the importance of the choice of ...

  2. [2019]

    have been pro- posed, but these methods remain focused on subspace-specific clustering. While subspace methods can be highly effective, there are settings in which the clustering structure requires the entire functional domain, making dimension reduction less desirable. Non-pa...

  3. [2021]

    The BIC selects the optimal model

    and h6mM>..* (Anton and Smith , 2023), respectively, each with 20 random k-means starts, all models, and with the Catell-scree threshold in {0.05, 0.1, 0.2, 0.6}. The BIC selects the optimal model. The mocca algorithm is run with the 7/ JQ++ package, setting the number of knot...

  4. [2022]

    for all sensors between the years 2009 and

  5. [2023]

    (2019), providing a natural basis for comparison with existing methods

    and Rivera-García et al. (2019), providing a natural basis for comparison with existing methods. First used in Febrero et al. (2008), the data reflect hourly measurements of nitric oxide (NO) and nitrogen dioxide (NO

  6. [2025]

    for _ (R Core Team , 2025), with X = fb1; : : : ;bng, G clusters, and F maximum outliers. 5: 2M/ T`Q+2/m`2 8 _2K `F kX h?2 +?QB+2 Q7 F Bb mb2`@bT2+B}2/ M/ /2MQi2b i?2 K tBKmK MmK#2` Q7 TQBMib `2KQp2/ /m`BM; i?2 QmiHB2` b2 `+?X AM i?2 #b2M+2 Q7 /QK BM FMQrH2/;2- F=2 Bb +QMb2`p ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.