REVIEW 4 major objections 3 minor 1 cited by
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper announces FACTORY, a human-verified prompt set for long-form factuality, and claims that state-of-the-art language models make about 40% nonfactual claims on it compared with about 10% on existing benchmarks.
desk verdict The abstract describes a factuality benchmark, the body is a functional-data clustering paper; the submission is not a reviewable paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object that carries the abstract's argument is FACTORY itself: a prompt set built with a model-in-the-loop approach and refined by humans, designed so every prompt is fact-seeking, answerable, and unambiguous. The measuring instrument is claim-level human evaluation, which decomposes each model response into claims and labels each claim factual or not, with the headline number being the share of nonfactual claims. The supplied body text, however, makes no use of this machinery; its machinery is funOCLUST, a two-stage algorithm that decomposes curves with cubic B-splines and then runs an outlier-trimming Gaussian mixture clusterer on the coefficients, using a shifted-and-scaled beta distribution for subset log-likelihoods to decide when to stop trimming. For the abstract's argument, the load-bearing machinery is the human-verified prompt set and the human evaluation protocol, neither of which appears in the supplied full text.
What would settle it
Open the submitted artifact and verify whether the body contains FACTORY's prompt set and human evaluation; it does not, which is an immediate negative check. If the missing materials are restored, the decisive test is independent human re-annotation of a random sample of FACTORY prompts: if many prompts are unanswerable or ambiguous, or a fresh run yields near-10% nonfactual claims, the central claim is false.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that long-form factuality can be measured more honestly with a human-verified prompt set: when six state-of-the-art models are evaluated claim by claim by humans, roughly 40% of their claims on FACTORY are not factual, against roughly 10% on existing datasets. The intended cause is prompt quality, because FACTORY prompts are selected to be fact-seeking, answerable, and unambiguous and to draw on long-tailed facts that models cannot answer by rote. The supplied body text does not contain the FACTORY study; it describes funOCLUST, a clustering algorithm for functional data, so this discovery is asserted in the abstract but not demonstrated in the manuscript.
Load-bearing premise
The load-bearing premise is that human verification really makes FACTORY prompts fact-seeking, answerable, and unambiguous and that the 40% versus 10% comparison was measured under the same protocol; the supplied body text is a different paper, so neither premise can be checked from this manuscript.
Editorial extensions
If this is right
- If FACTORY's 40% figure holds, current long-form factuality results showing near-10% nonfactual rates are optimistic by a factor of about four.
- A human-verified, answerable, unambiguous prompt set gives model developers a more reliable signal for where factual generation fails.
- The benchmark's emphasis on long-tailed facts implies that progress on FACTORY would require models to retrieve or reason about uncommon knowledge, not just popular topics.
- Because the comparison is claim-level, differences between models on FACTORY would be attributable to factual accuracy rather than to prompt ambiguity or unanswerability.
Reading between the lines
- If the 40% versus 10% gap survives independent audit, it would suggest that earlier benchmarks are saturated: reported advances may mostly reflect easier prompt distributions, prompting a re-reading of past comparisons.
- The human-verified design implies that automatic factuality metrics that match claims to retrieval sources may miss long-tail falsehoods, so hybrid human-and-automatic claim adjudication is a natural next measurement to test.
- A controlled test of the benchmark's contribution would build an unverified version of FACTORY with the same prompts but no human filtering and compare error rates; a large drop would isolate human verification as the driver.
- Because the submitted body is a different paper, the immediate extension is procedural: reconcile the artifact so the benchmark's prompts, annotation instructions, and model outputs are publicly inspectable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, arXiv:2508.00109, presents in its title and abstract a benchmark paper called FACTORY: a large-scale, human-verified prompt set for long-form factuality, claiming that state-of-the-art language models generate about 40% nonfactual claims on FACTORY versus about 10% on existing datasets. The full text, however, is entirely a different manuscript, 'funOCLUST: Clustering Functional Data with Outliers' by Clark and McNicholas, which develops a robust clustering algorithm for functional data and contains no prompt set, no human annotation protocol, no language-model evaluation, and no comparison of nonfactual rates. As a result, every central claim in the abstract is unsupported by the submitted document.
Significance. If the FACTORY claims were substantiated, the contribution would be significant: a human-verified benchmark that is fact-seeking, answerable, and unambiguous, together with evidence that it is substantially more challenging than prior datasets, would be a useful resource for long-form factuality evaluation. The reported 40% versus 10% gap is the kind of falsifiable, head-to-head comparison that would matter to the field. However, none of the supporting artifacts or analyses appear in the manuscript: the prompt set, verification protocol, annotator agreement, model list, response outputs, and dataset comparison are all absent. The body's funOCLUST material is mathematically unrelated and cannot substitute for the missing evidence. The submission therefore cannot be verified in its current form.
major comments (4)
- [Abstract vs. Full text] The document's title and abstract describe FACTORY, a human-verified long-form factuality prompt set, but the full text is entirely the funOCLUST paper on functional-data clustering. None of the benchmark claims in the abstract—neither the construction of the prompt set, nor the human-verification procedure, nor the evaluation on six language models—appears anywhere in the body. This mismatch removes the evidentiary basis for every central claim of the submission.
- [Abstract; missing benchmark materials] The abstract claims the prompt set is 'large-scale, human-verified' and developed 'model-in-the-loop' with human refinement, but the manuscript contains no prompt set, no annotation instructions, no annotator-agreement statistics, and no adjudication protocol. Without these materials, the reliability claim that is the paper's stated motivation cannot be assessed.
- [Abstract; missing evaluation] The headline quantitative result, that SOTA models produce approximately 40% nonfactual claims on FACTORY versus 10% on other datasets, is stated with no supporting experiments: no model names or versions, no generation settings, no response samples, no factuality annotation scheme, and no description of the comparison datasets. This is load-bearing because the paper's contribution is precisely the challenge gap and reliability relative to existing benchmarks.
- [§3.2, Lemma 1 and Remark 1] The only technical content in the body is the funOCLUST extension. Lemma 1's proof delegates to Theorem 1 of Clark and McNicholas (2024), and Remark 1 explicitly conditions the derived distribution on a finite Gaussian mixture with no outlying coefficients. These are internal limitations of the statistical contribution, but they do not bear on the FACTORY claims; even a fully corrected funOCLUST paper would not supply the missing benchmark evidence.
minor comments (3)
- [Throughout] The body contains encoding artifacts such as the header 'R AMi`Q/m+iBQM' and rot13-like section headings, which make the document difficult to read and suggest a conversion error.
- [Metadata] The author list and references in the body belong to the funOCLUST paper; the name FACTORY never appears after the abstract. The submission metadata should be reconciled with the actual content.
- [Abstract] The abstract says '6 state-of-the-art language models' but names none of them; if the correct manuscript is resubmitted, the models, versions, and prompts should be specified.
Circularity Check
No circular derivation in the funOCLUST body; the FACTORY abstract is unsupported by the text, but that is a missing-support issue, not circular reasoning.
full rationale
The submitted full text is funOCLUST, a functional-data clustering paper, not the FACTORY benchmark described in the abstract. The abstract's claims about human verification, 40% vs 10% nonfactual claim rates, and six LLMs cannot be checked because the body contains no prompt set, annotation protocol, or model evaluation; this is an integrity/verifiability failure, not a circular-derivation failure. Within funOCLUST, Lemma 1 explicitly extends Theorem 1 of Clark and McNicholas (2024) by substituting B-spline fitted coefficients, Sigma_g + sigma^2(B^T B)^-1 for Sigma_g, and p = K + 4. That is a legitimate corollary of a published theorem, not a result whose conclusion is assumed in its premises. The outlier-trimming rule compares subset log-likelihood differences to a derived null distribution, which is a goodness-of-fit target rather than a fitted parameter renamed as a prediction. The simulation study, main-effect analysis, and real-data examples (Melbourne pedestrian data, Barcelona NOx data) provide external, independent evidence for the method's behavior. No step in the derivation chain reduces by construction to its own inputs.
Assumptions & free parameters
assumptions (2)
- domain assumption Human verification provides a valid and consistent ground truth for whether a claim is factual.
- domain assumption Model-in-the-loop generation followed by human refinement produces prompts that are fact-seeking, answerable, and unambiguous.
invented entities (1)
-
FACTORY prompt set
Cite this review
Pith. "Pith review of FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality." pith.science (2026). https://pith.science/paper/6ZWAAYZG
@misc{pith2026250800109,
author = {Pith},
title = {Pith review of: FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality},
year = {2026},
howpublished = {\url{https://pith.science/paper/6ZWAAYZG}},
note = {Machine review of arXiv:2508.00109}
}
read the original abstract
Long-form factuality evaluation assesses the ability of models to generate accurate, comprehensive responses to short prompts. Existing benchmarks often lack human verification, leading to potential quality issues. To address this limitation, we introduce FACTORY, a large-scale, human-verified prompt set. Developed using a model-in-the-loop approach and refined by humans, FACTORY includes challenging prompts that are fact-seeking, answerable, and unambiguous. We conduct human evaluations on 6 state-of-the-art language models using FACTORY and existing datasets. Our results show that FACTORY is a challenging benchmark: approximately 40% of the claims made in the responses of SOTA models are not factual, compared to only 10% for other datasets. Our analysis identifies the strengths of FACTORY over prior benchmarks, emphasizing its reliability and the necessity for models to reason across long-tailed facts.
Forward citations
Cited by 1 Pith paper
-
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
On a new extended needle-in-a-haystack benchmark, explicit anti-hallucination prompts and dispersed fact placement cause some long-context LLMs to over-refuse or collapse in accuracy, while others remain robust.
Reference graph
Works this paper leans on
-
[1]
funOCLUST: Clustering Functional Data with Outliers Katharine M. Clark 1 and Paul D. McNicholas 2 1Department of Mathematics & Statistics, Trent University, Ontario, Canada. 2Department of Mathematics & Statistics, McMaster University, Ontario, Canada. Abstract Functional data present unique challenges for clustering due to their infinite-dimensional natu...
work page 2005
-
[3]
fbeta 2nh (nh 1)2 (dj c) p 2 ; nh p 1 2 (2) for c < d j < (nh1)2 2nh + c; nh > p + 1, where c = log ^ h + p 2 log(2 ) + 1 2 logjShj, nh is the number of points in cluster h, ^ h = nh=n, Sh = 1 nh 1 nX i=1 zih(xi xh)(xi xh)> is the sample covariance matrix of cluster h, and xh = 1 nh Pn i=1 zihxi: The trimmed model chosen minimizes the Kullback-Leibler (KL...
work page 2005
-
[4]
is a collection of isolation trees which randomly split the functional space into partitions, iteratively, until each function is isolated. Functions isolated in fewer iterations tend to be more outlying. An outlier score is calculated using a forest of these isolation trees. Finally, Hubert et al. (2015) determine functional outlyingness by first calcula...
work page 2015
-
[5]
As in Clark and McNicholas (2024), these assumptions can be relaxed in practice
f#2i 2nh (nh 1)2 (dj c) K + 4 2 ; nh K 5 2 (7) 7Q` c < d j < (nh1)2 2nh + c; nh > K + 5- r?2`2 c = log ^ h + K+4 2 log(2 ) + 1 2 logjShj- nh Bb i?2 MmK#2` Q7 TQBMib BM +Hmbi2` h- ^ h = nh=n- Sh = 1 nh 1 nX i=1 zih(bi bh)(bi bh)> Bb i?2 b KTH2 +Qp `B M+2 K i`Bt Q7 +Hmbi2` h- bh = 1 nh Pn i=1 zihbi; M/ K Bb i?2 MmK#2` Q7 BMi2`BQ` FMQib BM i?2 +m#B+ bTHBM2 #...
work page 2024
-
[6]
Thus, any nontrivial transformation of increases the MSD. This transformation carries over to the functional representation: B(t)b = B(t)(a + c) = a (t) + B(t)c; 7 where (t) = B(t) denotes the mean function. Hence, deviations of b from correspond directly to deviations of the associated function from the mean function. Because lXnb lX / DM (b ); observati...
work page 2015
-
[7]
Additionally, these outliers often exhibit unusual combinations of traffic patterns
Many of the outlying observations have higher-than-average traffic in the afternoon (noon to 5pm) or early morning (midnight to 4am). Additionally, these outliers often exhibit unusual combinations of traffic patterns. For example, traffic volume around noon and 7pm may resemble a typical workday, while the volume around 4pm is more similar to what is observ...
work page 2022
-
[11]
Data captured for the year 2017 from the Chinatown–Swanston St. (North) sensor are plotted in Figure 6, with observations coloured according to whether the data correspond to a workday or non-workday (i.e., a weekend or public holiday). There are 365 observations, each with values collected at 24 time points. Of the 365 days captured, 249 are workdays and...
work page 2017
-
[1982]
as the distance option. 9Xk *Hmbi2`BM; _2bmHib Each clustering algorithm is evaluated using the adjusted Rand index (ARI; Hubert and Arabie, 1985). The ARI compares two partitions — in this case, real and predicted classes. 11 The ARI equals one under perfect classification and has expected value zero under random class assignment. For funOCLUST and tkmea...
work page 1985
Show all 14 references
-
[2005]
Solid black curves indicate the group-wise hourly mean
Curves are coloured according to whether each day is a workday (W) or non-workday (NW). Solid black curves indicate the group-wise hourly mean. 0.51 to 0.86 places its best-performing models on par with existing approaches while demon- strating the importance of the choice of ...
2022
-
[2019]
have been pro- posed, but these methods remain focused on subspace-specific clustering. While subspace methods can be highly effective, there are settings in which the clustering structure requires the entire functional domain, making dimension reduction less desirable. Non-pa...
2024
-
[2021]
The BIC selects the optimal model
and h6mM>..* (Anton and Smith , 2023), respectively, each with 20 random k-means starts, all models, and with the Catell-scree threshold in {0.05, 0.1, 0.2, 0.6}. The BIC selects the optimal model. The mocca algorithm is run with the 7/ JQ++ package, setting the number of knot...
2023
-
[2022]
for all sensors between the years 2009 and
2009
-
[2023]
(2019), providing a natural basis for comparison with existing methods
and Rivera-García et al. (2019), providing a natural basis for comparison with existing methods. First used in Febrero et al. (2008), the data reflect hourly measurements of nitric oxide (NO) and nitrogen dioxide (NO
2019
-
[2025]
for _ (R Core Team , 2025), with X = fb1; : : : ;bng, G clusters, and F maximum outliers. 5: 2M/ T`Q+2/m`2 8 _2K `F kX h?2 +?QB+2 Q7 F Bb mb2`@bT2+B}2/ M/ /2MQi2b i?2 K tBKmK MmK#2` Q7 TQBMib `2KQp2/ /m`BM; i?2 QmiHB2` b2 `+?X AM i?2 #b2M+2 Q7 /QK BM FMQrH2/;2- F=2 Bb +QMb2`p ...
2003
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.