{"id":"a9d1e451-9484-431b-b720-ed931d761c72","arxiv_id":"2506.20918","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A hybrid manual and LLM workflow enriches HathiTrust ETD metadata with generated keywords and abstracts, yielding a 5,760-record dataset.","lead":"This paper builds a dataset of 5,760 digitized theses and dissertations with metadata enriched by manual curation and large language models. It shows how combining human review with off-the-shelf LLM tools can add keywords and abstracts that repositories like HathiTrust did not originally provide.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed 'high degree of semantic accuracy' for LLM-generated keywords and abstracts rests on an unreported author self-review; no quantitative evidence supports it.","rationale":"I agree with the reader that the weakest assumption is the unquantified 10% manual check. The paper's main positive contributions are a public dataset and a clear field mapping; these are real and reproducible. However, the abstract's benefit claim is the load-bearing conclusion, and it depends on an accuracy assertion that has no measured outcome. A secondary internal inconsistency also deserves attention: starting from 6,445 ETDs, removing 1,131 non-ETD, 216 OCR-poor, and 56 missing-full-text documents yields 5,042, not 5,760, unless the final count counts split author/degree rows rather than documents; this should be reconciled in a revision. Neither issue is fatal given the artifact's availability, but both require the authors to either add a quantitative evaluation or weaken the central claim. Thus the conditional verdict is appropriate and my read does not change it.","tokens_in":5767,"tokens_out":6523,"duration_ms":74343,"concrete_test":"Draw a stratified random sample of 100 records from the Zenodo dataset (balanced by publication decade and degree level). Two independent annotators, using the original full text, rate the generated dc:subject keywords and dc:description:abstract on a three-point scale (accurate / partially accurate / inaccurate), with a rubric requiring abstracts to contain only source-grounded content and keywords to be present in or directly inferable from the document. Compute per-field accuracy and Cohen's kappa. Pre-register thresholds (e.g., acceptable rate >= 90% and kappa >= 0.6). If thresholds are not met, the paper should weaken 'high degree of semantic accuracy' or report the observed error rates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the hybrid workflow 'makes it possible to generate and integrate missing metadata fields with a high degree of semantic accuracy.' The only evaluation cited is the sentence in the Semantic Enrichment section: 'we randomly selected and manually reviewed 10% of the metadata records enriched through keyword extraction and text summarization techniques against the original source documents, confirming their accuracy.' This cannot support the claim as written. The paper reports no error counts, no per-field accuracy rates for keywords versus abstracts, no rubric for what counts as an accurate abstract, no inter-rater agreement, and no baseline such as original HathiTrust fields, title-only keyword inference, or a non-LLM extractor. Because the review was done by the authors themselves, 'confirming their accuracy' is a self-assessment, not a checkable result. The released Zenodo dataset is genuine independent evidence of the artifact, but it does not by itself establish the semantic-accuracy claim; an accuracy estimate requires scored comparisons against source documents. Without these numbers the central claim is not verifiable from the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This short paper reports a hybrid manual-and-LLM workflow for enriching metadata of English-language theses and dissertations (ETDs) from the HathiTrust Digital Library. The authors retrieved 6,445 ETD records, mapped 20 fields to Dublin Core / ETD-MS, manually cleaned and completed fields from title and preliminary pages, used KeyLLM to generate keywords and PaperQA with GPT-4o-mini to generate abstracts, and released a claimed 5,760-record enriched dataset on Zenodo under CC BY-NC-SA 4.0. The central claim is that this hybrid approach makes it possible to 'generate and integrate missing metadata fields with a high degree of semantic accuracy' and that LLM enrichment is 'particularly beneficial' for repositories with missing metadata.","tokens_in":5935,"tokens_out":4286,"duration_ms":40666,"significance":"If the quality of the enriched fields were properly quantified, the released dataset and the field-mapping exercise would be a useful contribution to digital-library and ETD metadata practice. The authors should be credited for providing a concrete, reusable artifact with a persistent DOI, for documenting a nontrivial manual-cleaning process, and for mapping HathiTrust fields to Dublin Core/ETD-MS. The main scientific value, however, depends on evidence that the LLM-generated keywords and abstracts are actually accurate; that evidence is currently only a self-reported, underspecified 10% review. The paper's secondary claims about improved search and accessibility are not evaluated at all.","major_comments":[{"comment":"The only evidence cited for the central claim of 'high degree of semantic accuracy' is the sentence: 'we randomly selected and manually reviewed 10% of the metadata records enriched through keyword extraction and text summarization techniques against the original source documents, confirming their accuracy.' No sample size, per-field error rates, rubric, inter-rater agreement, or comparison baseline is reported, and the review was conducted by the authors themselves. This is insufficient to support the strong accuracy claim made in the Conclusion; please add quantitative quality statistics or substantially weaken the claim.","section":"Semantic Enrichment"},{"comment":"The record counts are internally inconsistent. The paper reports 6,445 initial ETDs, removal of 1,131 non-ETD/duplicate titles, removal of 216 ETDs with OCR errors, and removal of 56 titles lacking full-text files, with a final dataset of 5,760 records. Those numbers sum to 6,445 - (1,131 + 216 + 56) = 5,042, not 5,760. Because the released dataset is the paper's primary artifact, this arithmetic discrepancy must be resolved and explained in the text.","section":"Findings"},{"comment":"The abstract and conclusion claim that the approach enhances 'search results' and improves 'the accessibility of the digital repository,' but the paper reports no search/retrieval evaluation, no user study, and no accessibility analysis. These claims should either be supported with an appropriate evaluation or removed/restricted to a statement about the availability of additional metadata access points.","section":"Abstract and Conclusion"},{"comment":"The paper states: 'In our study, we used the ‘title’ field to extract relevant keywords using KeyLLM and populated the results in the ‘dc:subject’ metadata field.' If keywords were generated from the title alone, then dc:subject does not describe the full document content as Table 1's description ('Keywords or subject terms describing the thesis') implies, and the semantic enrichment claim is weakened. Please clarify whether the full text or only the title was used; if only the title was used, this limitation must be stated explicitly.","section":"Semantic Enrichment - Keyword Extraction"}],"minor_comments":[{"comment":"The phrase 'has can be replicated across other digital libraries' should read 'can be replicated'.","section":"Conclusion"},{"comment":"The text says 'we added 10 more fields' but Table 1 lists 11 fields marked n/a (advisor, committeeChair, committeeMember, department, discipline, grantor, degree name, degree level, spatial, subject, abstract). Please reconcile this count.","section":"Findings"},{"comment":"The search query description ('keywords such as “dissertation”, or “academic” in the subject field') is too vague for replication; please provide the exact query, the search date, and the inclusion/exclusion criteria.","section":"Methodology - Data Collection"},{"comment":"For reproducibility, please report the prompt templates, model parameters (e.g., temperature, max tokens), and the date of OpenAI API access for both KeyLLM and PaperQA, or state that default settings were used.","section":"Semantic Enrichment"},{"comment":"The Zenodo DOI is provided, but no dataset version or file checksum is reported; adding these would improve reproducibility of the released artifact.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short conference paper, so I do not expect a full retrieval evaluation, but the current omission of any quantitative quality assessment is too large for the paper's accuracy claim. The arithmetic inconsistency in the record counts also needs correction before the dataset can be trusted. I would be willing to see a revised version that adds per-field accuracy numbers (even on the 10% sample), reconciles the counts, and softens the search/accessibility claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a dataset-and-workflow paper, and the dataset is the part that matters. The 5,760 enriched records with 20 mapped Dublin Core fields—including manually added advisor/committee/degree data and LLM-generated keywords and abstracts—are a genuinely reusable resource for digital humanities and repository work. The authors did the careful, boring parts well: they document de-duplication, OCR failures, and missing fields in plain language, and they ship the data on Zenodo with a license. That is real evidence of care.\n\nThe soft spots are in the evaluation of the LLM output. The paper says a 10% random sample was manually reviewed against source documents \"confirming their accuracy,\" but gives no counts, error rates, per-field breakdown, rubric, or inter-rater agreement. That cannot support the abstract's \"particularly beneficial\" and \"high degree of semantic accuracy\" claims. Also, the record-count arithmetic does not reconcile: 6,445 minus 1,131 misclassified minus 216 poor OCR minus 56 missing full text leaves 5,042, yet the paper reports 5,760. There is a 718-record gap, probably due to later reintegration, but the paper never walks through it.\n\nNone of this kills the contribution. For a short conference paper, the dataset and the field mapping are the contribution; the evaluation claims can be either scaled back or substantiated by reporting actual numbers. If the authors revise, the central argument would hold if they present the accuracy check as a quality-assurance step, not as a measured result. I would send this to peer review as a short paper, but ask for the evaluation to be either tightened or removed, and the counts reconciled.","headline":"A useful, carefully documented ETD enrichment dataset whose central accuracy claim is under-evidenced and whose record counts don't reconcile.","tokens_in":6482,"tokens_out":2571,"would_cite":true,"duration_ms":24646,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid human-LLM pipeline can enrich long-text metadata with high semantic accuracy, producing a reusable dataset of 5,760 theses and dissertations.","keywords":["metadata enrichment","electronic theses and dissertations","large language models","semantic enrichment","keyword extraction","abstract generation","digital libraries","Dublin Core"],"falsifier":"Take any random sample of, say, 100 enriched records and have independent evaluators compare each generated keyword set and abstract against the original source document using a defined scoring rubric; if the measured error rate exceeds the level implied by 'high degree of semantic accuracy,' or if retrieval tests show enriched metadata does not outperform the original sparse metadata, the central claim is falsified.","tokens_in":5541,"feed_emoji":"🎓","tokens_out":9037,"duration_ms":80191,"temperature":0.7,"pith_summary":"This paper shows that a hybrid approach—manual curation for structured fields plus large language models for keywords and abstracts—can fill metadata gaps in long text documents such as theses and dissertations. Working from a collection of digitized dissertation records, the authors cleaned and mapped the metadata to the Dublin Core standard and then used an LLM keyword extractor and a retrieval-augmented summarizer to create two fields that the original records lacked. They report that a 10% manual review against source documents confirmed the semantic accuracy of the generated content. The result is a reusable dataset of 5,760 enriched records that provides new access points for search and discovery, and the method is proposed as a scalable solution for repositories with incomplete metadata.","feed_headline":"Hybrid human-LLM method fills metadata gaps for 5,760 dissertations","feed_subtitle":"A reusable dataset shows how digital repositories can add keywords and abstracts at scale.","key_machinery":"The load-bearing mechanism is the two-stage semantic enrichment pipeline. In the first stage, structured metadata fields such as advisor, committee, degree, and department are populated by hand from title pages and preliminary pages, and the original record is normalized and mapped to the Dublin Core standard for electronic theses and dissertations. In the second stage, an LLM-based keyword extractor takes the title as input and produces subject terms, while a retrieval-augmented generation package indexes the merged full text and answers the prompt 'What is the abstract of this paper?' to produce the abstract. A manual review pass over 10% of the outputs serves as the accuracy gate. The enriched fields are organized under a Dublin Core mapping designed for ETDs, and the published dataset carries the resulting 20-field metadata schema.","core_discovery":"The central claim is that the hybrid workflow allows missing metadata fields to be generated and integrated with a high degree of semantic accuracy. The paper demonstrates this on a collection of 6,445 English-language theses and dissertations, which after cleaning and de-duplication became 5,760 records. It adds a subject/keyword field by extracting key terms from each title with an LLM, and an abstract field by generating summaries from the full text with a retrieval-augmented question-answering package. The new fields are designed to capture document themes and produce coherent, comprehensive summaries, not just surface-level tokens. The authors state that random manual review of 10% of the enriched records against the original documents confirmed accuracy, and they release the enriched dataset as a reusable resource.","pith_inferences":["A sharper test of the method would measure retrieval performance, for instance whether topic-based queries return more relevant enriched records than unenriched ones; the paper's manual review does not quantify search improvement.","The approach could plausibly transfer to historical or OCR-degraded documents, but success will depend on the quality of the full text, since the abstract generator explicitly fails when OCR errors are severe.","The validation rests on a 10% sample without reported error metrics; an external evaluation with multiple annotators and an explicit error budget would make the 'high degree of semantic accuracy' claim testable.","The specific LLM and extractor choices are incidental; the paper's core demonstration is that a hybrid pipeline can integrate LLM outputs into a standards-compliant metadata record, not that any particular model is necessary."],"forward_implications":["Repositories with sparse metadata can add searchable keywords and abstracts to long text documents without re-cataloging each item manually, improving discoverability.","The Dublin Core mapping and cleaning rules offer a template for standardizing ETD metadata across institutions, improving interoperability.","The released dataset of 5,760 enriched records gives computational social science and digital humanities researchers a ready-made corpus with structured fields they can search and analyze.","Because the enrichment pipeline is automated after the manual curation of structured fields, it can scale to collections larger than the one tested here.","If the semantic accuracy holds, the method can be extended to other document types whose original metadata was designed for a different content type, such as books rather than theses."],"supporting_citations":[{"why":"Prior work using LLMs to improve ETD metadata fields, providing the baseline the paper extends toward a hybrid approach.","marker":"Choudhury, et al., 2023"},{"why":"Supplies the keyword extraction algorithm used to populate the dc:subject field from titles.","marker":"KeyLLM, 2024"},{"why":"Provides the PaperQA retrieval-augmented package used to generate abstract field content from full text.","marker":"Lála et al., 2023"},{"why":"Defines the ETD metadata standard and Dublin Core mapping that structures the enriched record.","marker":"NDLTD, 2023"},{"why":"Establishes the metadata quality criteria of accuracy, completeness, and consistency that motivate the enrichment goals.","marker":"Park, 2009"},{"why":"Supports the claim that metadata quality is crucial for interoperability and research use.","marker":"Liu et al., 2024"}],"fun_headline_variants":["LLM adds missing keywords and abstracts to 5,760 dissertations","Hybrid workflow enriches 5,760 theses with new metadata fields","Dataset offers LLM-enriched metadata for 5,760 dissertations","Human-LLM pipeline creates searchable abstracts and keywords","5,760 dissertations gain LLM-generated metadata access points"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim of high semantic accuracy rests on the authors' manual review of a randomly selected 10% of the enriched records, with no reported error rates, sample-size justification, or inter-rater agreement, so that 10% is assumed to stand for the whole dataset.","fun_headline_variants_meta":{"raw":{"variants":["LLM adds missing keywords and abstracts to 5,760 dissertations","Hybrid workflow enriches 5,760 theses with new metadata fields","Dataset offers LLM-enriched metadata for 5,760 dissertations","Human-LLM pipeline creates searchable abstracts and keywords","5,760 dissertations gain LLM-generated metadata access points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000285,"raw_usage":{"total_tokens":1611,"prompt_tokens":808,"completion_tokens":803,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":708}},"tokens_in":424,"tokens_out":803,"duration_ms":8183,"temperature":1.0,"reasoning_tokens":708,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:38:22.107070+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any random sample of, say, 100 enriched records and have independent evaluators compare each generated keyword set and abstract against the original source document using a defined scoring rubric; if the measured error rate exceeds the level implied by 'high degree of semantic accuracy,' or if retrieval tests show enriched metadata does not outperform the original sparse metadata, the central claim is falsified.","supporting_citations":[],"review_version":1}