{"id":"988a83d6-3ebb-4094-aafd-d3ac96a8b5b0","arxiv_id":"2607.09324","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"On the Their Finest Hour archive, open extractive NER and statistical keyword extraction outperform generative AI for scalable, accountable keyword assignment in crowdsourced collections.","lead":"A team tested statistical, open neural, and generative AI methods for keywording a crowdsourced WWII archive of living contributors' stories. Open extractive models beat GenAI on accuracy and accountability, so stewards should prefer them when metadata must stay faithful to contributors.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Quantitative rankings rest on n=10 team-annotated records; this underpins the claim that open extractive models outperform GenAI for responsible keywording.","rationale":"The Reader correctly isolates the n=10 gold-standard limitation as the weakest load-bearing premise. The paper is transparent about it, supplies open code/data, and frames results as a case study rather than a definitive sector ranking. The ethics argument about living contributors and answerability for GenAI outputs can stand independently of the precise numerical order. Because the authors already flag the sample as indicative, the concern does not require a harsher verdict; CONDITIONAL remains appropriate provided readers treat the model rankings as provisional. No stronger internal inconsistency or fabrication risk is present. The concrete test above would settle whether the ranking generalises or is sample-specific.","tokens_in":32861,"tokens_out":542,"duration_ms":6186,"concrete_test":"Independently re-annotate a stratified sample of ≥50 TFH Description fields (or a public subset) by at least two annotators external to the project team, using the same KE@5 and NER precision/recall protocol; recompute Tables 5–6. If the relative ordering of the top statistical KE methods and Flair/SpaCy vs GPT-4o reverses or the gaps shrink below practical significance, the generalisation claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest technical claim—that open-weight extractive models (Flair/SpaCy NER, statistical KE, GLiNER) are preferable to GenAI for accountable keyword extraction—depends on the quantitative rankings in Tables 5–6. Those rankings are computed against a gold standard of only ten records, double-annotated by two project-team members with consensus resolution (§3.5). The paper itself labels the sample 'indicative rather than definitive.' Because the same team also built the original TFH keyword workflow and now evaluates the models, the gold labels may systematically favour extractive, text-grounded terms that match their prior practice and ethical stance, while undervaluing abstractive or generative outputs. If the ranking order (MultipartiteRank/TopicRank/PositionRank and Flair/SpaCy above GPT-4o) is an artefact of this small, non-independent sample, the recommendation that open extractive models 'emerge … as best placed to support responsible deployment' loses its empirical footing and becomes primarily a normative argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper evaluates three NLP approaches—Named Entity Recognition, Keyword Extraction, and Topic Modelling—for automated keyword assignment on the Their Finest Hour crowdsourced Second World War archive (2,003 description fields). It compares statistical methods, open-weight specialised models (Flair, SpaCy, GLiNER, KeyBERT, BERTopic, etc.), and closed generative models (GPT-4o/4o-mini) via quantitative precision/recall on a 10-record gold standard (Tables 5–6) and qualitative analysis of extractive vs abstractive outputs, relational terms, and stewardship. The central claim is that NLP can support keyword extraction at scale in crowdsourced collections, but no single method is complete; open-weight extractive models best support responsible deployment because their outputs are text-grounded and answerable, whereas generative AI introduces accountability risks that stewards of living-contributor collections must weigh carefully. Methods and code are released openly.","tokens_in":33129,"tokens_out":812,"duration_ms":7531,"significance":"The work is a useful, practice-oriented contribution at the intersection of digital humanities, GLAM metadata, and applied NLP. Its distinctive value is the integrated treatment of technical performance with stewardship and answerability for living contributors, grounded in a real archive the same team crowdsourced and published. Strengths include the breadth of methods compared, the open-source reusable toolkit (Arana-Catania et al. 2025), explicit discussion of extractive vs abstractive keywords and relational terms, and a clear normative argument that model choice is an ethical as well as technical decision. If the rankings and recommendations hold under larger evaluation, the paper would give GLAM practitioners concrete, actionable guidance on when open extractive models are preferable to GenAI for keywording.","major_comments":[{"comment":"§3.5 and Tables 5–6: The quantitative rankings that underwrite the claim that open extractive models (MultipartiteRank/TopicRank/PositionRank for KE; Flair/SpaCy/GLiNER for NER) outperform GPT-4o and are ‘best placed to support responsible deployment’ rest on a gold standard of only ten records, double-annotated by two project-team members with consensus resolution. The paper itself labels this sample ‘indicative rather than definitive.’ With n=10 the ranking order is under-powered and may be sensitive to the team’s prior extractive, text-grounded practice (the original TFH controlled-vocabulary workflow). Either enlarge and independently re-annotate the gold standard, report confidence intervals / bootstrap stability of the rankings, or reframe the quantitative results as exploratory and rest the stewardship recommendation more explicitly on the qualitative and ethical argument rather t","section":null},{"comment":"§4.1 / Table 5 vs §5 (Extraction, abstraction, and generative AI): KE evaluation scores only extractive, in-text keywords against a gold standard that contains no absent keywords (§3.5). Abstractive or generative outputs that introduce useful umbrella terms (e.g. ‘D-Day’, ‘agriculture’) are therefore systematically disadvantaged on precision/recall@5 even when the discussion later treats such terms as potentially valuable for discovery. This mismatch means the quantitative claim that statistical KE and open NER ‘outperform’ GenAI does not fairly test the abstractive use-case the paper itself raises. Clarify the evaluation protocol (e.g. separate extractive vs abstractive scoring, or a second gold set that allows absent keywords) or qualify the ranking claim so it applies only to extractive keywording.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: this is a careful applied case study on the Their Finest Hour WWII crowdsourced archive (2,003 description fields), not a new NLP method. Same team that crowdsourced and published the collection now ranks KE, NER, and TM from statistical tools through Flair/SpaCy/GLiNER to GPT-4o, ships open code and data, and argues that open extractive models better support stewardship of living-contributor stories than abstractive GenAI.\n\nWhat is new is the integrated setting and the ethics frame, not the stack. Tables 5–6 are clear: MultipartiteRank/TopicRank/PositionRank lead KE precision@5; Flair ner-large and SpaCy trf lead NER precision/recall; GPT-4o sits mid-pack; KeyBERT is weak. Qualitative points land—relational terms (\"mother\"), GLiNER’s custom categories, extractive vs abstractive discovery, and the responsibility-gap argument for stewards. Code and archive DOIs are real assets.\n\nSoft spot, in proportion: the quantitative ranking that underwrites “open extractive models emerge as best placed” is scored on ten records double-annotated by two project members with consensus. They call it indicative; that is correct. Same-team gold can tilt toward text-grounded extractive terms that match their prior TFH workflow and ethical stance, so the order vs GPT-4o is not a large independent benchmark. TM is qualitative only. None of that invents a circularity problem or kills the paper; it just means the recommendation is practice-informed guidance, not definitive model ranking for the sector.\n\nMath and citations look ordinary and honest—standard libraries, no fabricated entities, self-cites are the archive and tools repo. For digital-humanities, library, and cultural-heritage AI readers who need a reusable workflow and a clear stewardship argument, this is worth the time. I would send it to peer review with the sample-size limit stated up front; I would not treat the ranking as settled science.","headline":"Useful GLAM case study with open code; the open-extractive-over-GenAI ranking is real on their sample but rests on n=10 team gold labels, so treat it as indicative practice guidance, not a sector benchmark.","tokens_in":33731,"tokens_out":532,"would_cite":true,"duration_ms":8031,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Open extractive AI can scale keywords for living contributors' archives without surrendering answerability; generative AI cannot.","keywords":["Keywords","Crowdsourced","NLP","NER","Keyword Extraction","Topic Modelling","Responsible AI","Digital archives"],"falsifier":"Re-run the same KE and NER evaluation protocol on a substantially larger, independently annotated sample (hundreds of records) drawn from this archive or a comparable crowdsourced collection; if open specialised NER and statistical KE no longer dominate generative models on precision, recall, and answerability, the central ranking collapses.","tokens_in":33748,"feed_emoji":"🗂️","tokens_out":588,"duration_ms":6293,"temperature":0.7,"pith_summary":"Crowdsourced collections force a hard trade-off: platforms demand keywords at scale, but imposed terms risk misrepresenting living contributors. Using a Second World War archive of more than two thousand personal stories, this paper tests Named Entity Recognition, keyword extraction, and topic modelling across statistical methods, open specialised neural models, and closed generative AI. Statistical keyword extractors deliver high precision; specialised open NER models deliver both high precision and high recall; generative models lag and introduce outputs that cannot be fully traced. The authors conclude that no method is complete, that model choice shapes what can be discovered, and that extractive open-weight models best preserve stewardship duties because their keywords stay grounded in the contributor's own words and remain checkable. Generative abstraction may improve findability, but at the cost of answerability for people whose family histories are being described.","feed_headline":"Open extractive AI scales archive keywords without losing answerability","feed_subtitle":"Statistical KE and specialised NER beat generative models on a wartime crowdsourced collection","key_machinery":"Comparative evaluation of three NLP approaches (Named Entity Recognition, Keyword Extraction, Topic Modelling) across statistical, open specialised neural, and closed generative implementations, scored against a human gold standard and judged for stewardship fitness.","core_discovery":"Natural Language Processing can automate keyword extraction for crowdsourced collections, yet no single approach solves the problem. Open-weight extractive models—especially specialised NER and statistical keyword extractors—best support responsible deployment because their outputs remain text-grounded and answerable, while generative AI's abstractive power introduces accountability risks that stewards of living contributors' material must weigh carefully.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Open extractive models ground keywords in wartime crowdsourced archives","NER and statistical KE beat GenAI for answerable archive keywords","Extractive AI keeps keyword scale answerable for living collections","No single NLP approach fully solves crowdsourced keyword extraction","Open-weight extractive models best balance scale and archive stewardship"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The ranking of models rests on a gold standard of only ten records double-annotated by the project team, which the paper itself treats as indicative rather than definitive for the full archive.","fun_headline_variants_meta":{"raw":{"variants":["Open extractive models ground keywords in wartime crowdsourced archives","NER and statistical KE beat GenAI for answerable archive keywords","Extractive AI keeps keyword scale answerable for living collections","No single NLP approach fully solves crowdsourced keyword extraction","Open-weight extractive models best balance scale and archive stewardship"]},"model":"grok-4.5","effort":"low","cost_usd":0.003608,"raw_usage":{"total_tokens":1172,"prompt_tokens":760,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":36080000,"prompt_tokens_details":{"text_tokens":760,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":345,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":760,"tokens_out":67,"duration_ms":4908,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T03:51:51.457985+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same KE and NER evaluation protocol on a substantially larger, independently annotated sample (hundreds of records) drawn from this archive or a comparable crowdsourced collection; if open specialised NER and statistical KE no longer dominate generative models on precision, recall, and answerability, the central ranking collapses.","supporting_citations":[],"review_version":1}