{"id":"01cb82cc-57fe-45a4-af22-86d70ef96e06","arxiv_id":"2507.07551","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An out-of-the-box vision language model's catalogue descriptions were detectable by humans and rated less accurate and useful than expert-written texts.","lead":"Researchers tested whether an image-captioning AI (InternVL2) can write museum catalogue descriptions that fool people. People could tell AI and human texts apart, rated human texts as better, and trust in AI dropped after seeing the outputs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The d' = 0.67 result pertains to human-curated model output, not to the out-of-the-box VLM; the human selection and verification loop is the load-bearing gap between the data and the discussion's central claim.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the evaluated AI descriptions passed through a human selection, proofreading, translation, and expert-verification loop, so the 'out-of-the-box model' framing in the abstract and Discussion overstates what was tested. My read therefore does not move the verdict; it remains CONDITIONAL with the condition that the 'out-of-the-box' claim be either retested with raw outputs or explicitly qualified to refer to the human-curated pipeline. I add two points of emphasis. First, the manuscript contains the relevant limitation itself in the Materials section, so the issue is not a hidden flaw but an interpretive gap between method and conclusion. Second, the core within-pipeline results are credible: the preregistration, the signal-detection analysis, the Bayesian checks, and the large sample support the claim that participants could detect curated AI descriptions above chance and rated them lower on accuracy and usefulness. The concern is specifically that those results cannot license the statement that 'the out-of-the-box model' failed the Turing test, because the model never produced the final evaluated texts by itself. A concrete test using raw outputs would settle whether the curation loop is doing the work; absent that test, the manuscript should be accepted only with the scoping caveat made explicit.","tokens_in":21935,"tokens_out":3118,"duration_ms":37407,"concrete_test":"Run the same classification and rating experiment on a new set of 40 items (or a random half of the current 40 cards) using raw InternVL2 outputs: a single generation at temperature 0.5, no selection among three candidates, no typo correction, no translation, and no expert verification. Keep prompts and the 2x2 procedure otherwise identical. If the raw-output d' is materially lower than 0.67, or if accuracy/usefulness ratings fall below the reported 67.6% and 58.0%, the human-curation loop explains part of the apparent human-likeness and the 'out-of-the-box' conclusion is unsupported. If d' and ratings are statistically equivalent to the curated values, the curation loop is not the driver and the central claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline conclusion is that 'the out-of-the-box VLM... failed to pass the Turing test' and is not yet suitable for cataloguing without human review. But the descriptions actually evaluated were not raw model output. In the Materials section, the authors state: 'a human loop was integrated... the model generated three descriptions per labelled photo card with varying parameter values from which the non-expert selected the best-fitting description per labelled photo card, ensuring it was free of typos or non-expert detectable errors. It was then translated into German, verified by an expert, and integrated into the experiment.' In addition, the prompt was iteratively engineered and the temperature parameter was tuned. Thus d' = 0.67, t(138) = 12.67, and the accuracy/usefulness ratings characterize a human-in-the-loop pipeline, not the unaided model. The central claim about 'out-of-the-box' capability is therefore not derivable from these data. The empirical finding within the curated pipeline is solid and interesting, and the conclusion may still be true (raw outputs could be even easier to detect), but the paper's framing overstates what was tested. This is not an internal inconsistency, but it is a scope-of-evidence problem because the abstract and discussion attribute the result to the model alone. The fix is either to qualify the claim to 'curated VLM-assisted descriptions' or to test raw outputs directly.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a preregistered online experiment (N = 139, including 29 experts) in which participants classified and rated 80 catalogue descriptions: 40 generated by InternVL2-Llama3-76B and 40 written by human experts, matched to 40 labelled photo cards from the LEIZA archaeological archive. The main results are above-chance discrimination of AI versus human texts (d' = 0.67, t(138) = 12.67, p < .001), no evidence of an expert advantage, self-assessed detection performance lower than actual performance, expert-written descriptions rated as more accurate and useful than AI-generated ones, and post-task declines in willingness to use and trust in AI tools. The authors conclude that an out-of-the-box VLM fails a Turing-test-like benchmark and that human review remains necessary for archival cataloguing.","tokens_in":22150,"tokens_out":10145,"duration_ms":117756,"significance":"The study has clear strengths: a realistic archival stimulus set, preregistered hypotheses, signal detection measures, Bayesian complements, and a stated commitment to open data and scripts. If the claims are scoped to the materials actually evaluated, the detection and rating effects provide useful human-centered evidence for AI-assisted cataloguing in GLAM institutions. However, the headline 'out-of-the-box' conclusion is not supported by the materials, because a human selected, cleaned, translated, and expert-verified every AI description before it was shown to participants. This scope-of-evidence gap changes what the experiment can say and requires either re-scoping the conclusions or adding a direct test of raw model outputs. The paper also uses trial-level GLMMs without item random effects, which weakens several secondary and exploratory analyses.","major_comments":[{"comment":"The descriptions evaluated as 'AI-generated' were not raw model output. The Methods state that InternVL2 generated three descriptions per photo card with temperature values 0.1, 0.5, and 0.75, from which a non-expert selected the best-fitting description, checked it for typos and non-expert detectable errors, and that the selected description was then translated into German and verified by an expert; the prompt was also iteratively engineered. Therefore the d' = 0.67 effect, the accuracy/usefulness ratings, and the classification learning effects characterize a human-in-the-loop pipeline, not 'the out-of-the-box VLM' named in the abstract and in the Research Question 1 discussion ('failed to pass the Turing test'). The conclusion that the out-of-the-box model is unsuitable for cataloguing is not derivable from these data; please re-scope the abstract, hypotheses, and discussion to 'curated VLM-assisted descriptions' or add a direct supplementary test of raw model outputs.","section":"Materials, 'Generating catalogue descriptions'; Discussion, Research Question 1"},{"comment":"The trial-level GLMM includes random slopes for description type by participant but no random intercepts for items (the 40 photo cards or 80 descriptions). Because every participant sees all items and each photo card contributes two descriptions, the conditional-independence assumption is violated, and omitting item random effects can understate standard errors and inflate chi-square statistics for description type and for the exploratory predictors (trial number, accuracy, usefulness). Please re-fit the models with item random intercepts or justify treating items as fixed, and report whether the main d'-based conclusion and the exploratory effects are robust; the participant-level d' analysis itself is not affected by this concern.","section":"Results, 'Hypothesis 1' GLMM; 'Exploratory analyses'"},{"comment":"The H3 comparison, r(27) = .87 versus r(108) = .69, is computed on participant-level aggregates, so it tests whether participants who give high average accuracy also give high average usefulness, not whether accuracy and usefulness are more tightly coupled across individual AI descriptions within raters. Aggregation can inflate correlations through between-participant differences in scale use. A multilevel model with random intercepts and slopes, or within-participant correlations with appropriate error terms, is needed before concluding that experts treat factual accuracy and practical usefulness as more intertwined than non-experts do.","section":"Results, 'Hypothesis 3'; Figure 6"},{"comment":"The Bayes factors are mislabeled. For the expert versus non-expert d' contrast, BF = 0.68 corresponds to roughly 1.5:1 in favor of the null, which is anecdotal, not 'moderate' as stated in the Results. For H2b, BF = 0.31 corresponds to roughly 3.2:1 for the null, which is moderate. The H1 statement that 'this ability did not depend on archival or archaeological expertise' is therefore stronger than the evidence warrants; please rephrase the conclusion and report BF01 values or equivalence tests where null claims are made.","section":"Results, 'Hypothesis 1' and 'Hypotheses 2' Bayes factors"}],"minor_comments":[{"comment":"The reference Hashmati et al. (2024) contains a placeholder DOI ('10.1145/xxxxxxx.xxxxxxx') and must be completed before publication.","section":"References"},{"comment":"The Data availability section and footnotes provide a Zenodo preview URL with an access token rather than a stable DOI; please use the permanent record identifier and fix the 'zenodoi' typo in the Results text.","section":"Data availability"},{"comment":"Supplementary Table B1.3 reports the Fisher z p-value for AI-generated descriptions as p = 0.14, while the main text reports p = .014; please make these consistent.","section":"Supplementary Table B1.3"},{"comment":"The abstract's phrase 'OCR errors and hallucinations limited perceived quality' is presented as a finding, but the paper does not report a systematic coding or analysis of error types; please either add such an analysis or soften the causal wording to 'were associated with lower perceived quality'.","section":"Abstract and Discussion"},{"comment":"Figure 6 shows correlations between accuracy and usefulness but does not state that the plotted correlations are participant-level aggregates; please clarify this in the caption so readers do not interpret the lines as item-level relationships.","section":"Figure 6 caption"}],"recommendation":"major_revision","confidential_remarks":"The empirical core of the paper is solid once the claims are scoped to the curated pipeline, and the authors have provided preregistration and open materials. The main risk is the 'out-of-the-box' framing in the abstract and discussion, which is contradicted by the Materials section's description of human selection, translation, and expert verification. I would urge the editor to require the re-scoping or a raw-output comparison before acceptance; the remaining statistical issues, while real, are secondary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is real: d' = 0.67, t(138) = 12.67, BF = 5.6 x 10^21, and both experts and non-experts classify AI-generated descriptions above chance. Human-written texts are rated more accurate and useful, and the response-bias and learning effects are sensible. If you work on human-AI collaboration in GLAM settings, this is a useful data point.\n\nWhat the paper does well: it takes a real archival corpus, builds a plausible catalogue template, and runs a preregistered experiment with both domain experts and laypeople. The use of signal detection theory is appropriate, and the Bayesian complements are a plus. The exploratory link between perceived quality and detectability—higher-rated AI texts are harder to classify—is a nice validation of the Turing-test proxy. The materials documentation is also thorough.\n\nThe soft spot is the one the stress-test flags, and it is real: the evaluated AI descriptions were not raw model output. The model generated three descriptions per photo card, a non-expert selected the best, checked for typos, translated it into German, and an expert verified it. So the d' and ratings characterize a curated, human-in-the-loop pipeline, not the 'out-of-the-box VLM' the abstract and discussion cite. The conclusion may still be true—raw output could be even easier to detect—but the data as they stand do not support the unqualified claim. That is a scope-of-evidence problem, not a fatal one. The fix is straightforward: re-scope the central claim to 'curated VLM-assisted descriptions' or add a raw-output condition.\n\nOther issues are minor. The pre-post trust and willingness measures lack a control group, so the causal reading is weaker than the text implies. The GLMM omits item-level random effects, which could inflate significance, though the t-test on d' is robust to that concern. The Zenodo link is a tokenized preview URL rather than a stable DOI, which is annoying for reproducibility. The H3 correlation is computed on participant-level aggregates, fine but worth noting.\n\nThis is a solid, honest empirical study with one load-bearing framing issue. It deserves a serious referee, with the request that the authors fix the scope-of-evidence gap. I would send it to peer review rather than desk reject.","headline":"A preregistered, well-run study showing people can detect curated VLM catalogue text above chance; the 'out-of-the-box' claim overshoots the data because the evaluated text went through a human selection and verification loop.","tokens_in":22724,"tokens_out":1868,"would_cite":true,"duration_ms":21023,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"People can distinguish AI-written from expert-written archive catalogue descriptions at above-chance rates, and they rate the AI versions as less accurate and useful.","keywords":["vision language models","catalogue descriptions","archival cataloguing","Turing test","signal detection theory","human-AI collaboration","trust in AI","museum collections"],"falsifier":"Run the same 80-item discrimination experiment using the model's raw outputs before best-of-three selection, typo correction, translation, and expert verification; if d' is no longer above chance or the accuracy and usefulness ratings converge with expert texts, the paper's characterization of the out-of-the-box model would collapse.","tokens_in":21706,"feed_emoji":"🖼️","tokens_out":7518,"duration_ms":75008,"temperature":0.7,"pith_summary":"The paper asks whether a vision language model (a model that reads both images and text) can produce archive catalogue descriptions that are indistinguishable from those written by human experts, and whether professionals would trust and adopt such output. Using forty labelled archaeological photo cards and a purpose-built ten-field catalogue template, the authors had InternVL2 generate descriptions and asked 139 participants (archive and archaeology experts plus non-experts) to classify each description as AI-written or expert-written, rate its accuracy and usefulness, and report trust and willingness to use AI. The central finding is that people classified the descriptions above chance (d' = 0.67), even without domain expertise, while expert-written descriptions were rated as more accurate and useful, and exposure to the AI outputs lowered participants' willingness to use AI tools. If this holds, an out-of-the-box VLM cannot yet take over cataloguing without human review, and successful integration depends on trust and transparent workflows as much as on technical quality.","feed_headline":"People detect AI-written archive entries above chance","feed_subtitle":"Even non-experts could tell machine from expert text; AI entries scored lower on accuracy and usefulness.","key_machinery":"The central machinery is a controlled discrimination experiment built around a ten-field catalogue template derived from the Minimum Record Recommendation for museums and collections, so that human and machine outputs are structurally comparable. Each participant saw 80 photo-description pairs and made a binary 'AI or expert?' decision; the binary responses are scored with signal detection theory, using sensitivity d' to measure discriminability and response bias c to measure the tendency to answer 'expert-written', while generalized linear mixed models test trial-level effects of description type, expertise, and ratings. The signal detection framework is what converts the binary judgments into the paper's core quantitative claim of above-chance human detection.","core_discovery":"On the paper's own terms, the discovery is that an unfine-tuned, out-of-the-box vision language model fails a Turing-style test for archival cataloguing: mean sensitivity was d' = 0.67, t(138) = 12.67, p < .001, far above chance, and this held for non-experts as well as experts, with no reliable expert advantage. AI-generated descriptions received lower perceived accuracy (67.61% vs. 76.11%) and lower perceived usefulness (58.01% vs. 70.01%) than expert-written ones. Exploratory analyses showed that higher-rated AI descriptions were harder to classify, and that direct experience with the outputs reduced general willingness to use AI and trust in AI; experts were less willing to adopt AI tools than non-experts. The authors conclude that human review remains necessary and recommend a collaborative workflow in which AI drafts initial descriptions and experts verify, with an explainable pipeline to build trust.","pith_inferences":["Because participants saw the best of three generations that a human then cleaned and verified, the paper's 'out-of-the-box model' framing likely underestimates the raw model's error rate; a raw-output replication is the natural test.","The inverse relation between perceived quality and detectability suggests the Turing-test gap is not a stable property of current VLMs: any technical improvement that raises rated accuracy and usefulness should push d' toward chance, so the shelf life of this result is short.","The study's quality ratings capture perceived accuracy, not error counts; a findability-focused evaluation that counts transcription errors and hallucinations per record could show whether AI drafts still save archivist time even when detectable.","The observed decline in trust after exposure parallels algorithm aversion; a direct follow-up comparing trust under an explainable pipeline versus a black-box pipeline would test whether transparency can offset that decline."],"forward_implications":["Above-chance detection at d' = 0.67 means an out-of-the-box vision language model cannot yet produce catalogue entries indistinguishable from expert writing.","Because higher-rated AI descriptions were harder to classify, improvements in accuracy and usefulness may push detection toward chance, so the current detectability gap is not a fixed ceiling.","Since exposure to outputs lowered willingness to use AI and experts reported lower willingness, adoption in archives depends on trust and workflow integration, not just output quality.","Expert-written descriptions were rated as more accurate and useful, so a human-in-the-loop draft-then-verify workflow is the paper's recommended division of labour."],"supporting_citations":[{"why":"Provides the signal detection framework (d' and c) used to quantify classification performance.","marker":"Green & Swets, 1966"},{"why":"Defines the imitation-game framing that motivates Research Question 1.","marker":"Turing, 1950"},{"why":"Describes InternVL2, the vision language model whose descriptions are evaluated.","marker":"Wang et al., 2024"},{"why":"Supplies the prompt-engineering method used to instruct the model.","marker":"Sahoo et al, 2024"},{"why":"Basis of the ten-field catalogue template that makes human and AI outputs comparable.","marker":"Minimum Record Working Group et al., 2024"},{"why":"Prior evidence that risk and opportunity perception predict willingness to use AI, used to interpret adoption findings.","marker":"Schwesig et al., 2023"},{"why":"Supports the claim that trust and professional ethics are decisive for AI uptake in archives.","marker":"Jaillant & Rees, 2023"},{"why":"Algorithm aversion research used to explain why exposure to errors lowers willingness to use AI.","marker":"Dietvorst et al., 2015"},{"why":"Provides the lme4 package and mixed-model analysis used in the trial-level analyses.","marker":"Bates et al., 2015"}],"fun_headline_variants":["AI archive entries fail Turing test, humans spot them","Even non-experts outsmart AI in archive cataloguing","AI-written archive labels rated lower, humans detect them","Humans beat out-of-the-box AI at archive descriptions","Expert reviewers still required for AI archive cataloguing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that an out-of-the-box model cannot meet archival standards assumes that the curated descriptions shown to participants—best of three generations, proofread, translated into German, and expert-verified—still represent the model's unaided capability rather than the human editing loop.","fun_headline_variants_meta":{"raw":{"variants":["AI archive entries fail Turing test, humans spot them","Even non-experts outsmart AI in archive cataloguing","AI-written archive labels rated lower, humans detect them","Humans beat out-of-the-box AI at archive descriptions","Expert reviewers still required for AI archive cataloguing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000609,"raw_usage":{"total_tokens":2861,"prompt_tokens":999,"completion_tokens":1862,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":1785}},"tokens_in":615,"tokens_out":1862,"duration_ms":14874,"temperature":1.0,"reasoning_tokens":1785,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:37:25.303818+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 80-item discrimination experiment using the model's raw outputs before best-of-three selection, typo correction, translation, and expert verification; if d' is no longer above chance or the accuracy and usefulness ratings converge with expert texts, the paper's characterization of the out-of-the-box model would collapse.","supporting_citations":[],"review_version":1}