Pith. sign in

REVIEW 4 major objections 5 minor 2 references

ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read People can distinguish AI-written from expert-written archive catalogue descriptions at above-chance rates, and they rate the AI versions as less accurate and useful.

desk verdict A preregistered, well-run study showing people can detect curated VLM catalogue text above chance; the 'out-of-the-box' claim overshoots the data because the evaluated text went through a human selection and verification loop. read the letter →

arxiv 2507.07551 v1 pith:6CI6GAQK submitted 2025-07-10 cs.HC cs.AIcs.DL

classification cs.HCcs.AIcs.DL
keywords visionlanguagemodelscataloguedescriptionsarchivalcataloguingTuringtestsignaldetectiontheoryhuman-AIcollaborationtrustinAImuseumcollections
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether a vision language model (a model that reads both images and text) can produce archive catalogue descriptions that are indistinguishable from those written by human experts, and whether professionals would trust and adopt such output. Using forty labelled archaeological photo cards and a purpose-built ten-field catalogue template, the authors had InternVL2 generate descriptions and asked 139 participants (archive and archaeology experts plus non-experts) to classify each description as AI-written or expert-written, rate its accuracy and usefulness, and report trust and willingness to use AI. The central finding is that people classified the descriptions above chance (d' = 0.67), even without domain expertise, while expert-written descriptions were rated as more accurate and useful, and exposure to the AI outputs lowered participants' willingness to use AI tools. If this holds, an out-of-the-box VLM cannot yet take over cataloguing without human review, and successful integration depends on trust and transparent workflows as much as on technical quality.

What carries the argument

The central machinery is a controlled discrimination experiment built around a ten-field catalogue template derived from the Minimum Record Recommendation for museums and collections, so that human and machine outputs are structurally comparable. Each participant saw 80 photo-description pairs and made a binary 'AI or expert?' decision; the binary responses are scored with signal detection theory, using sensitivity d' to measure discriminability and response bias c to measure the tendency to answer 'expert-written', while generalized linear mixed models test trial-level effects of description type, expertise, and ratings. The signal detection framework is what converts the binary judgments into the paper's core quantitative claim of above-chance human detection.

What would settle it

Run the same 80-item discrimination experiment using the model's raw outputs before best-of-three selection, typo correction, translation, and expert verification; if d' is no longer above chance or the accuracy and usefulness ratings converge with expert texts, the paper's characterization of the out-of-the-box model would collapse.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that an unfine-tuned, out-of-the-box vision language model fails a Turing-style test for archival cataloguing: mean sensitivity was d' = 0.67, t(138) = 12.67, p < .001, far above chance, and this held for non-experts as well as experts, with no reliable expert advantage. AI-generated descriptions received lower perceived accuracy (67.61% vs. 76.11%) and lower perceived usefulness (58.01% vs. 70.01%) than expert-written ones. Exploratory analyses showed that higher-rated AI descriptions were harder to classify, and that direct experience with the outputs reduced general willingness to use AI and trust in AI; experts were less willing to adopt AI tools than non-experts. The authors conclude that human review remains necessary and recommend a collaborative workflow in which AI drafts initial descriptions and experts verify, with an explainable pipeline to build trust.

Load-bearing premise

The conclusion that an out-of-the-box model cannot meet archival standards assumes that the curated descriptions shown to participants—best of three generations, proofread, translated into German, and expert-verified—still represent the model's unaided capability rather than the human editing loop.

Editorial extensions

If this is right

  • Above-chance detection at d' = 0.67 means an out-of-the-box vision language model cannot yet produce catalogue entries indistinguishable from expert writing.
  • Because higher-rated AI descriptions were harder to classify, improvements in accuracy and usefulness may push detection toward chance, so the current detectability gap is not a fixed ceiling.
  • Since exposure to outputs lowered willingness to use AI and experts reported lower willingness, adoption in archives depends on trust and workflow integration, not just output quality.
  • Expert-written descriptions were rated as more accurate and useful, so a human-in-the-loop draft-then-verify workflow is the paper's recommended division of labour.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because participants saw the best of three generations that a human then cleaned and verified, the paper's 'out-of-the-box model' framing likely underestimates the raw model's error rate; a raw-output replication is the natural test.
  • The inverse relation between perceived quality and detectability suggests the Turing-test gap is not a stable property of current VLMs: any technical improvement that raises rated accuracy and usefulness should push d' toward chance, so the shelf life of this result is short.
  • The study's quality ratings capture perceived accuracy, not error counts; a findability-focused evaluation that counts transcription errors and hallucinations per record could show whether AI drafts still save archivist time even when detectable.
  • The observed decline in trust after exposure parallels algorithm aversion; a direct follow-up comparing trust under an explainable pipeline versus a black-box pipeline would test whether transparency can offset that decline.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports a preregistered online experiment (N = 139, including 29 experts) in which participants classified and rated 80 catalogue descriptions: 40 generated by InternVL2-Llama3-76B and 40 written by human experts, matched to 40 labelled photo cards from the LEIZA archaeological archive. The main results are above-chance discrimination of AI versus human texts (d' = 0.67, t(138) = 12.67, p < .001), no evidence of an expert advantage, self-assessed detection performance lower than actual performance, expert-written descriptions rated as more accurate and useful than AI-generated ones, and post-task declines in willingness to use and trust in AI tools. The authors conclude that an out-of-the-box VLM fails a Turing-test-like benchmark and that human review remains necessary for archival cataloguing.

Significance. The study has clear strengths: a realistic archival stimulus set, preregistered hypotheses, signal detection measures, Bayesian complements, and a stated commitment to open data and scripts. If the claims are scoped to the materials actually evaluated, the detection and rating effects provide useful human-centered evidence for AI-assisted cataloguing in GLAM institutions. However, the headline 'out-of-the-box' conclusion is not supported by the materials, because a human selected, cleaned, translated, and expert-verified every AI description before it was shown to participants. This scope-of-evidence gap changes what the experiment can say and requires either re-scoping the conclusions or adding a direct test of raw model outputs. The paper also uses trial-level GLMMs without item random effects, which weakens several secondary and exploratory analyses.

major comments (4)
  1. [Materials, 'Generating catalogue descriptions'; Discussion, Research Question 1] The descriptions evaluated as 'AI-generated' were not raw model output. The Methods state that InternVL2 generated three descriptions per photo card with temperature values 0.1, 0.5, and 0.75, from which a non-expert selected the best-fitting description, checked it for typos and non-expert detectable errors, and that the selected description was then translated into German and verified by an expert; the prompt was also iteratively engineered. Therefore the d' = 0.67 effect, the accuracy/usefulness ratings, and the classification learning effects characterize a human-in-the-loop pipeline, not 'the out-of-the-box VLM' named in the abstract and in the Research Question 1 discussion ('failed to pass the Turing test'). The conclusion that the out-of-the-box model is unsuitable for cataloguing is not derivable from these data; please re-scope the abstract, hypotheses, and discussion to 'curated VLM-assisted descriptions' or add a direct supplementary test of raw model outputs.
  2. [Results, 'Hypothesis 1' GLMM; 'Exploratory analyses'] The trial-level GLMM includes random slopes for description type by participant but no random intercepts for items (the 40 photo cards or 80 descriptions). Because every participant sees all items and each photo card contributes two descriptions, the conditional-independence assumption is violated, and omitting item random effects can understate standard errors and inflate chi-square statistics for description type and for the exploratory predictors (trial number, accuracy, usefulness). Please re-fit the models with item random intercepts or justify treating items as fixed, and report whether the main d'-based conclusion and the exploratory effects are robust; the participant-level d' analysis itself is not affected by this concern.
  3. [Results, 'Hypothesis 3'; Figure 6] The H3 comparison, r(27) = .87 versus r(108) = .69, is computed on participant-level aggregates, so it tests whether participants who give high average accuracy also give high average usefulness, not whether accuracy and usefulness are more tightly coupled across individual AI descriptions within raters. Aggregation can inflate correlations through between-participant differences in scale use. A multilevel model with random intercepts and slopes, or within-participant correlations with appropriate error terms, is needed before concluding that experts treat factual accuracy and practical usefulness as more intertwined than non-experts do.
  4. [Results, 'Hypothesis 1' and 'Hypotheses 2' Bayes factors] The Bayes factors are mislabeled. For the expert versus non-expert d' contrast, BF = 0.68 corresponds to roughly 1.5:1 in favor of the null, which is anecdotal, not 'moderate' as stated in the Results. For H2b, BF = 0.31 corresponds to roughly 3.2:1 for the null, which is moderate. The H1 statement that 'this ability did not depend on archival or archaeological expertise' is therefore stronger than the evidence warrants; please rephrase the conclusion and report BF01 values or equivalence tests where null claims are made.
minor comments (5)
  1. [References] The reference Hashmati et al. (2024) contains a placeholder DOI ('10.1145/xxxxxxx.xxxxxxx') and must be completed before publication.
  2. [Data availability] The Data availability section and footnotes provide a Zenodo preview URL with an access token rather than a stable DOI; please use the permanent record identifier and fix the 'zenodoi' typo in the Results text.
  3. [Supplementary Table B1.3] Supplementary Table B1.3 reports the Fisher z p-value for AI-generated descriptions as p = 0.14, while the main text reports p = .014; please make these consistent.
  4. [Abstract and Discussion] The abstract's phrase 'OCR errors and hallucinations limited perceived quality' is presented as a finding, but the paper does not report a systematic coding or analysis of error types; please either add such an analysis or soften the causal wording to 'were associated with lower perceived quality'.
  5. [Figure 6 caption] Figure 6 shows correlations between accuracy and usefulness but does not state that the plotted correlations are participant-level aggregates; please clarify this in the caption so readers do not interpret the lines as item-level relationships.

Circularity Check

0 steps flagged · score 0.0 of 10

Empirical study with independent participant data; no load-bearing circularity; the human-curation vs. out-of-the-box gap is a scope limitation, not a derivation-by-construction.

full rationale

This paper reports a preregistered behavioral experiment rather than a derivation chain. The headline result (d' = 0.67, t(138) = 12.67, p < .001) is computed from participants' binary classifications of 80 fixed photo-description items, and the accuracy and usefulness ratings are separate dependent variables, so the central claim is not defined in terms of its own conclusion and no fitted parameter is relabeled as a prediction. The exploratory GLMM showing that higher perceived accuracy and usefulness predict lower detectability is post hoc construct validation on the same dataset; it is methodologically soft but does not make the classification result equal to its inputs, because the two measures come from distinct responses. The only notable gap is one of scope, not circularity: the Materials section and Supplement A2 describe a human loop that selected the best of three temperature-varied generations, checked typos, translated the text into German, and had an expert verify it, while the Discussion attributes the failure to the out-of-the-box VLM failing the Turing test. That mismatch threatens construct or external validity, but it does not reduce the conclusion to its evidence by construction. Self-citations such as Schwesig et al. (2023) are used as background for risk perception and are not load-bearing for the detection result. No circular step meeting the requirement of a quoted equation or fitted-parameter reduction was found.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters were fitted to data; the statistical thresholds are standard and the temperature values are procedural hyperparameters, not tuned to maximize outcomes. The statistical and domain assumptions listed above are the main unstated premises the conclusions rest on.

assumptions (4)
  • domain assumption Classification performance (d') is a valid proxy for human-likeness and quality of AI-generated descriptions.
    The study operationalizes the Turing test as an indicator of quality (RQ1); the authors validate this with the exploratory finding that higher-rated descriptions are harder to classify, but this validation is post hoc.
  • domain assumption Expert-written descriptions represent the gold standard for catalogue description quality.
    Human expert descriptions are the baseline against which AI descriptions are compared on accuracy and usefulness; the study does not independently verify their factual accuracy.
  • standard math Standard statistical assumptions of t-tests and GLMMs (normality, logit link, random-effects structure) hold.
    The preregistered analyses rely on these assumptions; the GLMM does not include item-level random intercepts, which could affect inference.
  • domain assumption The 40 selected photo cards are representative of the LEIZA archaeological archive.
    The authors selected cards with varied dating, layouts, and motifs, but did not formally sample; generalization to the broader archive is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing." pith.science (2026). https://pith.science/paper/6CI6GAQK

@misc{pith2026250707551,
  author       = {Pith},
  title        = {Pith review of: ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6CI6GAQK}},
  note         = {Machine review of arXiv:2507.07551}
}
read the original abstract

The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether Al-generated catalogue descriptions can approximate human-written quality and how generative Al might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content, evaluated by archive and archaeology experts and non-experts in a human-centered, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect Al-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify, suggesting that human review is necessary to ensure the accuracy and quality of catalogue descriptions generated by the out-of-the-box model, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends not only on technical advancements, such as domain-specific fine-tuning, but even more on establishing trust among professionals, which could both be fostered through a transparent and explainable AI pipeline.

Figures

Figures reproduced from arXiv: 2507.07551 by the authors.

Figure 1
Figure 1. Materials and experiment Note. Creation procedure of the experimental material and how it was integrated in the experiment. Materials This subsection describes the creation of experimental materials for which photographic prints mounted on labelled cardboards from LEIZA’s image archive functioned as a database. Using a custom￾designed catalogue template, human experts and a VLM generated catalogue descriptions from … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 canonical work pages

  1. [127]

    ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing

    https://doi.org/10.3390/bs12050127 ArchiveGPT 38 Data availability All materials, code, and datasets generated and analyzed during the study are available in a zenodo recordii. Supplementary information Supplementary materials contain information on material creation and additional results in tables. Acknowledgements We want to thank all of Dominik Kimmel...

  2. [333]

    https://doi.org/10.69554/IAGK5522 Colavizza, G., Blanke, T., Jeurgens, C., & Noordegraaf, J. (2021). Archives and AI: An Overview of Current Debates and Future Perspectives. ACM Journal on Computing and Cultural Heritage, 15(1), 4:1- 4:15. https://doi.org/10.1145/3479010 ArchiveGPT 34 Davet, J., Hamidzadeh, B., & Franks, P. (2023). Archivist in the machin...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.