{"id":"d0f65310-c163-4862-b5e2-4ec6619a75ae","arxiv_id":"2508.06729","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"The abstract claims LLMs can annotate large oral history collections, but the full text is a different paper on MIP optimization, leaving the claim unsupported.","lead":"The abstract reports an LLM pipeline that tags 92,191 sentences from Japanese American incarceration oral histories for sentiment and meaning. The enclosed full text is an unrelated paper on parallel mixed-integer programming, so the abstract's results have no in-scope evidence.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's central claim is unsupported: the submitted full text is a different paper on parallel mixed-integer programming, containing none of the oral-history experiments or results cited.","rationale":"The reader's verdict of REJECT is appropriate because the abstract and full text are about entirely different studies. The reader's weakest_assumption focused on the representativeness of the 558 labeled sentences and transferability to 92,191 sentences—a relevant secondary concern if the oral-history study were actually present. But the more load-bearing problem is more fundamental: the full text provides no experimental content whatsoever to support the abstract's empirical claims. The F1 scores, dataset construction, prompt strategies, and scaled annotation are asserted in the abstract but never described or derived in the document. Under the review rule that all provided text is in-scope, this internal inconsistency means the central claim cannot be validated. I agree with the reader's ultimate rejection, but the primary reason is the complete absence of the study rather than a subtle sampling limitation. Thus 'partial' agreement: the reader's rationale hints at the mismatch but their named weakest_assumption is a downstream issue. The concrete test I propose—checking the arXiv version and GitHub repo—would definitively determine whether the mismatch is a pipeline error or a genuine deficiency; if it is a mismatch, the correct paper may exist elsewhere, but as submitted, the claim is unsupported. I recommend UNCHANGED because the reader's REJECT verdict correctly reflects the state of the submission: an abstract with unsupported empirical claims and a full text that contradicts the abstract's topic.","tokens_in":3767,"tokens_out":2726,"duration_ms":33030,"concrete_test":"Verify whether the actual arXiv record 2508.06729 on the arXiv website contains the oral-history full text, and inspect the linked GitHub repository (github.com/kc6699c/LLM4OralHistoryAnalysis) for the described pipeline. If the full text is indeed the PARBALANS paper (or otherwise contains no oral-history experimental sections, data, or code), and the repository has no annotation code or evaluation results for the JAIOH collection, then the central claim has no supporting evidence in this submission. Alternatively, if the supplied abstract/full-text pair is a pipeline mispairing, recover the correct full text and reassess; otherwise the mismatch and lack of support stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim—that LLMs can effectively perform semantic and sentiment annotation across large oral history collections—rests entirely on empirical results reported in the abstract: 558 labeled sentences, F1 scores (ChatGPT 88.71%, Llama 84.99%, Qwen 83.72% for semantic; Llama 82.66%, Qwen 82.29%, ChatGPT 82.29% for sentiment), and annotation of 92,191 sentences from 1,002 interviews. However, the full-text document is 'PARBALANS', a paper by different authors on parallel multi-armed bandits-based adaptive large neighborhood search for mixed-integer programming. It contains no mention of oral histories, prompt engineering, sentiment/semantic annotation, or the JAIOH collection. No methodology, dataset description, evaluation protocol, or raw results supporting the abstract's claims appear anywhere in the submission. This is an internal inconsistency, not merely a weak assumption about sample representativeness: the document that would substantiate the central claim is absent. The abstract's conclusion is therefore untestable from the provided text, and the reported F1 scores are orphaned assertions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission as received is internally inconsistent. The abstract advertises an oral-history NLP paper: a scalable LLM framework for semantic and sentiment annotation of Japanese American Incarceration Oral History, reporting F1 scores for ChatGPT, Llama, and Qwen, plus annotation of 92,191 sentences from 1,002 interviews. The full text, however, is a completely different paper titled 'PARBALANS: Parallel Multi-Armed Bandits-based Adaptive Large Neighborhood Search,' which concerns mixed-integer programming solvers and contains no mention of oral histories, LLMs, sentiment analysis, or the JAIOH collection. The central empirical claims of the abstract therefore have no supporting methodology, dataset description, evaluation protocol, or results anywhere in the submission.","tokens_in":4023,"tokens_out":2068,"duration_ms":24206,"significance":"If the advertised oral-history results were real and properly documented, the contribution could be significant for digital humanities: a reusable, ethics-aware annotation pipeline for culturally sensitive archives, with practical guidance on prompt design. However, as submitted, the contribution cannot be assessed because the body of the paper does not contain the study. The MIP solver content in the full text is a separate contribution that is not the subject of the abstract, and it is not evaluated here. The abstract's GitHub link and numerical findings are orphaned assertions with no in-text support.","major_comments":[{"comment":"The central claim, 'Our findings show that LLMs can effectively perform semantic and sentiment annotation across large oral history collections when guided by well-designed prompts,' is unsupported by the submitted body. The full text is the PARBALANS paper by Yilmaz, Cai, Kadioglu, and Dilkina, about parallel multi-armed bandits-based large neighborhood search for MIP. It contains no section, equation, table, or experiment related to oral history, sentence annotation, sentiment, semantic classification, or prompt engineering. The abstract and the manuscript are two different papers.","section":"Abstract vs. Full Text"},{"comment":"The quantitative evidence base is entirely missing. The abstract reports 558 expert-labeled sentences from 15 narrators, F1 scores (ChatGPT 88.71%, Llama 84.99%, Qwen 83.72% for semantic; Llama 82.66%, Qwen 82.29%, ChatGPT 82.29% for sentiment), and annotation of 92,191 sentences from 1,002 interviews. None of these numbers appear in the full text, and there is no experimental setup, no label taxonomy, no prompt templates, no train/test split, no standard deviation or error analysis, and no table or figure to support them. The numbers cannot be verified or reproduced from the submitted manuscript.","section":"Abstract, F1 scores and dataset counts"},{"comment":"Even taking the abstract at face value, the reported F1 scores are computed on a 558-sentence labeled set, and the best prompt configurations are then used to annotate 92,191 sentences from 1,002 interviews. The abstract provides no justification that the 15 narrators are representative of the full JAIOH collection, nor any independent validation of the annotations on the large corpus. As submitted, there is no way to assess whether the reported F1 transfers to the full corpus. This is a load-bearing gap in the empirical argument, though it is secondary to the complete absence of the study body.","section":"Methodology and transferability"}],"minor_comments":[{"comment":"The arXiv abstract, title, author list, and GitHub link refer to an oral-history study, while the full-text title, authors, abstract, and contributions refer to PARBALANS. The submission package appears to contain the wrong paper or a corrupted upload; this must be corrected before any further review.","section":"Metadata and submission integrity"},{"comment":"The abstract says 'This study provides a reusable annotation pipeline,' but no pipeline or code is described in the body. The GitHub URL is the only artifact, and it is not referenced or described anywhere in the full text.","section":"Abstract"},{"comment":"Numerous presentation issues in the full text (e.g., spacing, table layout, repeated 'GUROBI' formatting) may be artifacts of the unrelated PARBALANS paper. They are not relevant to the advertised oral-history study but would need attention if the correct manuscript is submitted.","section":"General presentation"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission-integrity problem: the uploaded full text is an entirely different paper, so the advertised oral-history study cannot be evaluated. I am not speculating about intent, but the editor should verify whether the wrong PDF was uploaded. If the intended oral-history manuscript exists, it should be submitted afresh with the full experimental details; the current version cannot be accepted or revised into a coherent paper within the scope of a normal revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the abstract and full text are not the same paper. The abstract describes an LLM annotation pipeline for Japanese American oral histories, with F1 scores and a 92k-sentence dataset. The full text is PARBALANS, a parallel mixed-integer programming solver paper by Yilmaz, Cai, Kadıoğlu, and Dilkina. It contains no mention of oral history, sentiment analysis, or prompt engineering. The reported results are orphaned assertions.\n\nTo give credit where it's due: the PARBALANS manuscript looks like a competent computational study. The benchmarks against Gurobi are extensive, the tables are clear, and the writing is straightforward. If that paper were submitted on its own, it would likely deserve a serious referee. But it is not the work described in the title or abstract of this submission.\n\nThe soft spots are not subtle. The abstract's central claim—that LLMs can effectively annotate large oral history collections—rests entirely on experiments that are absent from the body. There is no methodology, no dataset description, no annotation protocol, no evaluation details. Even taken on its own terms, the abstract raises a sample-size concern: tuning prompts on 558 sentences from 15 narrators and then applying to 92,191 sentences from 1,002 interviews assumes the labeled set is representative. That transferability is not justified. But the bigger problem is that the supporting document simply isn't here.\n\nWho is this for? A reader interested in oral history NLP would find nothing usable. A reader interested in MIP solvers might find the PARBALANS paper interesting, but it is not the paper under review. The mismatch is internal and load-bearing.\n\nRecommendation: desk reject or return to authors as a submission error. Do not send to peer review until the full text matches the abstract. If the correct full text exists, it may be worth a look, but this version is not reviewable.","headline":"The submission is an abstract for an oral history LLM study attached to an unrelated MIP solver paper, so the central claims are unsupported.","tokens_in":4455,"tokens_out":2132,"would_cite":false,"duration_ms":24060,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that well-designed prompts let LLMs annotate large oral history collections, reporting 88.71% F1 for semantic classification and 82.66% for sentiment, then scaling to 92,191 sentences.","keywords":["oral history","LLM annotation","sentiment analysis","semantic classification","prompt engineering","Japanese American incarceration","digital humanities","archival analysis"],"falsifier":"Take a stratified sample of sentences from the 92,191 that were not among the 558 used for prompt tuning, have expert annotators label them for the same semantic and sentiment categories, run the best prompts on the sample, and compare F1 scores. If the held-out F1 is materially lower than the development-set numbers, or varies sharply across narrators or topics, the transfer claim fails.","tokens_in":3711,"feed_emoji":"🗣️","tokens_out":5718,"duration_ms":58044,"temperature":0.7,"pith_summary":"This paper tries to show that large language models, steered by the right prompts, can carry out accurate semantic and sentiment annotation on large oral history archives, not just on small curated samples. Working with the Japanese American Incarceration Oral History collection, the authors built a multiphase pipeline: expert labeling of 558 sentences from 15 narrators, prompt design and model evaluation (zero-shot, few-shot, retrieval-augmented), and final annotation of 92,191 sentences from 1,002 interviews. On the labeled benchmark, ChatGPT reached 88.71% F1 for semantic classification and Llama reached 82.66% for sentiment, with all three model families within a few points of each other. If the transfer to the full corpus holds, the pipeline would let archives with sensitive, emotionally complex material become searchable and analyzable at scale without manual labeling of every sentence.","feed_headline":"One prompt recipe annotates 92,191 oral-history sentences","feed_subtitle":"ChatGPT hits 88.7% F1 on semantics; Llama leads sentiment in a Japanese American incarceration archive.","key_machinery":"The load-bearing machinery is the multiphase annotation workflow: a small expert-labeled set of 558 sentences, prompt strategies (zero-shot, few-shot, and retrieval-augmented generation), three LLM families (ChatGPT, Llama, Qwen), and a final large-scale annotation pass. The essential mechanism is prompt conditioning—the idea that the LLM's output can be aligned with an expert label schema by carefully specifying task, context, and examples, making the model act as a reliable annotator rather than a free-form responder.","core_discovery":"On its own terms, the paper reports that a well-designed prompt can make an off-the-shelf LLM agree with expert annotations on oral history sentences at F1 levels around 83–89%, and that this agreement is enough to justify annotating the entire collection. The discovery is not a new model or a new linguistic theory; it is a practical result about controllability: semantic classification accuracy depends strongly on prompt configuration, while sentiment analysis is more evenly matched across ChatGPT, Llama, and Qwen. The paper claims that the best prompt configurations transfer from the 558-sentence development set to the 92,191-sentence corpus, producing a reusable annotation pipeline that r","pith_inferences":["The development-set F1 numbers are measured on 558 sentences; the paper does not report a held-out evaluation on a fresh sample of the 92,191 annotated sentences, so the transfer claim is an assumption to test rather than a fully verified result.","Retrieval-augmented generation was among the prompt strategies tested; for culturally dense archives, grounding prompts in narrator-specific or event-specific context could plausibly improve accuracy beyond the zero-shot/few-shot results reported, but that is an extension the paper does not demonstrate.","The full text supplied with this record describes a different study, so the specific prompt templates, model versions, and annotation procedures behind the reported figures are not verifiable from the text given here; a reader would need to consult the authors' released materials to reproduce the numbers."],"forward_implications":["A 558-sentence expert-labeled set is enough to tune prompts that the paper then applies to the full 92,191-sentence corpus, dramatically lowering annotation cost.","Semantic classification is more prompt-sensitive than sentiment analysis: ChatGPT's F1 (88.71%) leads, but the gap among ChatGPT, Llama (84.99%), and Qwen (83.72%) shows no single model dominates.","Sentiment performance is comparable across models (Llama 82.66%, Qwen 82.29%, ChatGPT 82.29%), so model choice for sentiment can be driven by cost, access, or cultural-appropriateness considerations.","The pipeline is reusable: the same expert-annotation-plus-prompt-engineering workflow can be applied to other oral history collections with sensitive content."],"supporting_citations":[],"fun_headline_variants":["LLM prompt recipe labels 92,191 oral-history sentences","ChatGPT tops semantic scores, Llama leads sentiment on oral histories","AI annotates 92K oral history sentences with expert-level accuracy","Prompt design unlocks LLM annotation for 1,002 oral history interviews","LLMs classify 92K sentences: meaning and emotion from oral history"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The 558 expert-labeled sentences from 15 narrators are representative enough to tune prompts that perform equally well across all 92,191 sentences from 1,002 interviews.","fun_headline_variants_meta":{"raw":{"variants":["LLM prompt recipe labels 92,191 oral-history sentences","ChatGPT tops semantic scores, Llama leads sentiment on oral histories","AI annotates 92K oral history sentences with expert-level accuracy","Prompt design unlocks LLM annotation for 1,002 oral history interviews","LLMs classify 92K sentences: meaning and emotion from oral history"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1325,"prompt_tokens":845,"completion_tokens":480,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":388}},"tokens_in":589,"tokens_out":480,"duration_ms":5155,"temperature":1.0,"reasoning_tokens":388,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:33:13.612876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified sample of sentences from the 92,191 that were not among the 558 used for prompt tuning, have expert annotators label them for the same semantic and sentiment categories, run the best prompts on the sample, and compare F1 scores. If the held-out F1 is materially lower than the development-set numbers, or varies sharply across narrators or topics, the transfer claim fails.","supporting_citations":[],"review_version":1}