REVIEW 4 major objections 5 minor 2 references
ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read People can distinguish AI-written from expert-written archive catalogue descriptions at above-chance rates, and they rate the AI versions as less accurate and useful.
desk verdict A preregistered, well-run study showing people can detect curated VLM catalogue text above chance; the 'out-of-the-box' claim overshoots the data because the evaluated text went through a human selection and verification loop. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a controlled discrimination experiment built around a ten-field catalogue template derived from the Minimum Record Recommendation for museums and collections, so that human and machine outputs are structurally comparable. Each participant saw 80 photo-description pairs and made a binary 'AI or expert?' decision; the binary responses are scored with signal detection theory, using sensitivity d' to measure discriminability and response bias c to measure the tendency to answer 'expert-written', while generalized linear mixed models test trial-level effects of description type, expertise, and ratings. The signal detection framework is what converts the binary judgments into the paper's core quantitative claim of above-chance human detection.
What would settle it
Run the same 80-item discrimination experiment using the model's raw outputs before best-of-three selection, typo correction, translation, and expert verification; if d' is no longer above chance or the accuracy and usefulness ratings converge with expert texts, the paper's characterization of the out-of-the-box model would collapse.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that an unfine-tuned, out-of-the-box vision language model fails a Turing-style test for archival cataloguing: mean sensitivity was d' = 0.67, t(138) = 12.67, p < .001, far above chance, and this held for non-experts as well as experts, with no reliable expert advantage. AI-generated descriptions received lower perceived accuracy (67.61% vs. 76.11%) and lower perceived usefulness (58.01% vs. 70.01%) than expert-written ones. Exploratory analyses showed that higher-rated AI descriptions were harder to classify, and that direct experience with the outputs reduced general willingness to use AI and trust in AI; experts were less willing to adopt AI tools than non-experts. The authors conclude that human review remains necessary and recommend a collaborative workflow in which AI drafts initial descriptions and experts verify, with an explainable pipeline to build trust.
Load-bearing premise
The conclusion that an out-of-the-box model cannot meet archival standards assumes that the curated descriptions shown to participants—best of three generations, proofread, translated into German, and expert-verified—still represent the model's unaided capability rather than the human editing loop.
Editorial extensions
If this is right
- Above-chance detection at d' = 0.67 means an out-of-the-box vision language model cannot yet produce catalogue entries indistinguishable from expert writing.
- Because higher-rated AI descriptions were harder to classify, improvements in accuracy and usefulness may push detection toward chance, so the current detectability gap is not a fixed ceiling.
- Since exposure to outputs lowered willingness to use AI and experts reported lower willingness, adoption in archives depends on trust and workflow integration, not just output quality.
- Expert-written descriptions were rated as more accurate and useful, so a human-in-the-loop draft-then-verify workflow is the paper's recommended division of labour.
Reading between the lines
- Because participants saw the best of three generations that a human then cleaned and verified, the paper's 'out-of-the-box model' framing likely underestimates the raw model's error rate; a raw-output replication is the natural test.
- The inverse relation between perceived quality and detectability suggests the Turing-test gap is not a stable property of current VLMs: any technical improvement that raises rated accuracy and usefulness should push d' toward chance, so the shelf life of this result is short.
- The study's quality ratings capture perceived accuracy, not error counts; a findability-focused evaluation that counts transcription errors and hallucinations per record could show whether AI drafts still save archivist time even when detectable.
- The observed decline in trust after exposure parallels algorithm aversion; a direct follow-up comparing trust under an explainable pipeline versus a black-box pipeline would test whether transparency can offset that decline.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a preregistered online experiment (N = 139, including 29 experts) in which participants classified and rated 80 catalogue descriptions: 40 generated by InternVL2-Llama3-76B and 40 written by human experts, matched to 40 labelled photo cards from the LEIZA archaeological archive. The main results are above-chance discrimination of AI versus human texts (d' = 0.67, t(138) = 12.67, p < .001), no evidence of an expert advantage, self-assessed detection performance lower than actual performance, expert-written descriptions rated as more accurate and useful than AI-generated ones, and post-task declines in willingness to use and trust in AI tools. The authors conclude that an out-of-the-box VLM fails a Turing-test-like benchmark and that human review remains necessary for archival cataloguing.
Significance. The study has clear strengths: a realistic archival stimulus set, preregistered hypotheses, signal detection measures, Bayesian complements, and a stated commitment to open data and scripts. If the claims are scoped to the materials actually evaluated, the detection and rating effects provide useful human-centered evidence for AI-assisted cataloguing in GLAM institutions. However, the headline 'out-of-the-box' conclusion is not supported by the materials, because a human selected, cleaned, translated, and expert-verified every AI description before it was shown to participants. This scope-of-evidence gap changes what the experiment can say and requires either re-scoping the conclusions or adding a direct test of raw model outputs. The paper also uses trial-level GLMMs without item random effects, which weakens several secondary and exploratory analyses.
major comments (4)
- [Materials, 'Generating catalogue descriptions'; Discussion, Research Question 1] The descriptions evaluated as 'AI-generated' were not raw model output. The Methods state that InternVL2 generated three descriptions per photo card with temperature values 0.1, 0.5, and 0.75, from which a non-expert selected the best-fitting description, checked it for typos and non-expert detectable errors, and that the selected description was then translated into German and verified by an expert; the prompt was also iteratively engineered. Therefore the d' = 0.67 effect, the accuracy/usefulness ratings, and the classification learning effects characterize a human-in-the-loop pipeline, not 'the out-of-the-box VLM' named in the abstract and in the Research Question 1 discussion ('failed to pass the Turing test'). The conclusion that the out-of-the-box model is unsuitable for cataloguing is not derivable from these data; please re-scope the abstract, hypotheses, and discussion to 'curated VLM-assisted descriptions' or add a direct supplementary test of raw model outputs.
- [Results, 'Hypothesis 1' GLMM; 'Exploratory analyses'] The trial-level GLMM includes random slopes for description type by participant but no random intercepts for items (the 40 photo cards or 80 descriptions). Because every participant sees all items and each photo card contributes two descriptions, the conditional-independence assumption is violated, and omitting item random effects can understate standard errors and inflate chi-square statistics for description type and for the exploratory predictors (trial number, accuracy, usefulness). Please re-fit the models with item random intercepts or justify treating items as fixed, and report whether the main d'-based conclusion and the exploratory effects are robust; the participant-level d' analysis itself is not affected by this concern.
- [Results, 'Hypothesis 3'; Figure 6] The H3 comparison, r(27) = .87 versus r(108) = .69, is computed on participant-level aggregates, so it tests whether participants who give high average accuracy also give high average usefulness, not whether accuracy and usefulness are more tightly coupled across individual AI descriptions within raters. Aggregation can inflate correlations through between-participant differences in scale use. A multilevel model with random intercepts and slopes, or within-participant correlations with appropriate error terms, is needed before concluding that experts treat factual accuracy and practical usefulness as more intertwined than non-experts do.
- [Results, 'Hypothesis 1' and 'Hypotheses 2' Bayes factors] The Bayes factors are mislabeled. For the expert versus non-expert d' contrast, BF = 0.68 corresponds to roughly 1.5:1 in favor of the null, which is anecdotal, not 'moderate' as stated in the Results. For H2b, BF = 0.31 corresponds to roughly 3.2:1 for the null, which is moderate. The H1 statement that 'this ability did not depend on archival or archaeological expertise' is therefore stronger than the evidence warrants; please rephrase the conclusion and report BF01 values or equivalence tests where null claims are made.
minor comments (5)
- [References] The reference Hashmati et al. (2024) contains a placeholder DOI ('10.1145/xxxxxxx.xxxxxxx') and must be completed before publication.
- [Data availability] The Data availability section and footnotes provide a Zenodo preview URL with an access token rather than a stable DOI; please use the permanent record identifier and fix the 'zenodoi' typo in the Results text.
- [Supplementary Table B1.3] Supplementary Table B1.3 reports the Fisher z p-value for AI-generated descriptions as p = 0.14, while the main text reports p = .014; please make these consistent.
- [Abstract and Discussion] The abstract's phrase 'OCR errors and hallucinations limited perceived quality' is presented as a finding, but the paper does not report a systematic coding or analysis of error types; please either add such an analysis or soften the causal wording to 'were associated with lower perceived quality'.
- [Figure 6 caption] Figure 6 shows correlations between accuracy and usefulness but does not state that the plotted correlations are participant-level aggregates; please clarify this in the caption so readers do not interpret the lines as item-level relationships.
Circularity Check
Empirical study with independent participant data; no load-bearing circularity; the human-curation vs. out-of-the-box gap is a scope limitation, not a derivation-by-construction.
full rationale
This paper reports a preregistered behavioral experiment rather than a derivation chain. The headline result (d' = 0.67, t(138) = 12.67, p < .001) is computed from participants' binary classifications of 80 fixed photo-description items, and the accuracy and usefulness ratings are separate dependent variables, so the central claim is not defined in terms of its own conclusion and no fitted parameter is relabeled as a prediction. The exploratory GLMM showing that higher perceived accuracy and usefulness predict lower detectability is post hoc construct validation on the same dataset; it is methodologically soft but does not make the classification result equal to its inputs, because the two measures come from distinct responses. The only notable gap is one of scope, not circularity: the Materials section and Supplement A2 describe a human loop that selected the best of three temperature-varied generations, checked typos, translated the text into German, and had an expert verify it, while the Discussion attributes the failure to the out-of-the-box VLM failing the Turing test. That mismatch threatens construct or external validity, but it does not reduce the conclusion to its evidence by construction. Self-citations such as Schwesig et al. (2023) are used as background for risk perception and are not load-bearing for the detection result. No circular step meeting the requirement of a quoted equation or fitted-parameter reduction was found.
Assumptions & free parameters
assumptions (4)
- domain assumption Classification performance (d') is a valid proxy for human-likeness and quality of AI-generated descriptions.
- domain assumption Expert-written descriptions represent the gold standard for catalogue description quality.
- standard math Standard statistical assumptions of t-tests and GLMMs (normality, logit link, random-effects structure) hold.
- domain assumption The 40 selected photo cards are representative of the LEIZA archaeological archive.
Cite this review
Pith. "Pith review of ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing." pith.science (2026). https://pith.science/paper/6CI6GAQK
@misc{pith2026250707551,
author = {Pith},
title = {Pith review of: ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing},
year = {2026},
howpublished = {\url{https://pith.science/paper/6CI6GAQK}},
note = {Machine review of arXiv:2507.07551}
}
read the original abstract
The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether Al-generated catalogue descriptions can approximate human-written quality and how generative Al might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content, evaluated by archive and archaeology experts and non-experts in a human-centered, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect Al-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify, suggesting that human review is necessary to ensure the accuracy and quality of catalogue descriptions generated by the out-of-the-box model, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends not only on technical advancements, such as domain-specific fine-tuning, but even more on establishing trust among professionals, which could both be fostered through a transparent and explainable AI pipeline.
Figures
Reference graph
Works this paper leans on
-
[127]
ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing
https://doi.org/10.3390/bs12050127 ArchiveGPT 38 Data availability All materials, code, and datasets generated and analyzed during the study are available in a zenodo recordii. Supplementary information Supplementary materials contain information on material creation and additional results in tables. Acknowledgements We want to thank all of Dominik Kimmel...
-
[333]
https://doi.org/10.69554/IAGK5522 Colavizza, G., Blanke, T., Jeurgens, C., & Noordegraaf, J. (2021). Archives and AI: An Overview of Current Debates and Future Perspectives. ACM Journal on Computing and Cultural Heritage, 15(1), 4:1- 4:15. https://doi.org/10.1145/3479010 ArchiveGPT 34 Davet, J., Hamidzadeh, B., & Franks, P. (2023). Archivist in the machin...
arXiv 2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.