{"id":"205d60e9-eca1-4ad7-953e-d8788aae9867","arxiv_id":"2603.17723","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A supervised LLM framework for systematic literature reviews improves paper-classification accuracy over unconstrained models and maps option-pricing research.","lead":"LR-Robot is a framework that combines human supervision with large language models to sort and summarize large bodies of research papers. The authors test it on option-pricing literature, reporting better classification accuracy when human-written rules guide the AI.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported classification gains rely on the same 1,000-paper sample used for prompt refinement; without a held-out set, the central claim of improved accuracy and the validity of Section 4 analyses are not yet established.","rationale":"The Reader's weakest_assumption is exactly the concern I identify: the evaluation is conducted on the same sample used for prompt refinement, with no held-out set. This is the load-bearing issue because the paper's main empirical claims—that human-in-the-loop instructions improve classification and that the resulting labels enable accurate literature synthesis—depend on the reported metrics reflecting true generalization performance. The in-sample evaluation makes the accuracy improvements in Tables 1 and 2 difficult to interpret, and the downstream analyses in Section 4 inherit any label errors. The internal inconsistency in the reported Lenient Accuracy for model types (0.6739 in text vs 0.8657 in Table 4) reinforces the concern that the empirical record is not yet reliable. I do not see a deeper flaw in the framework's architecture or motivation; the idea of combining expert supervision with LLM classification is reasonable and the workflow is clearly described. However, the empirical support is not sufficient for the claims as stated. Since the Reader already recommended a CONDITIONAL verdict (essentially requiring held-out evaluation and resolution of inconsistencies), my assessment does not change that verdict. I therefore mark verdict_should_be as UNCHANGED: the paper should remain conditional pending the concrete validation test.","tokens_in":12981,"tokens_out":3157,"duration_ms":35958,"concrete_test":"Before any further prompt refinement, split the 1,000 labeled papers into a development set (e.g., 700) and a held-out test set (e.g., 300). Refine constraints and select the model using only the development set. Freeze the final prompt, then run Gemini Flash 2.0 three times on the held-out test set and compute accuracy, F1, self-consistency, Jaccard, and Lenient Accuracy for all four dimensions, comparing against the unconstrained baseline. In addition, recompute Lenient Accuracy for the model-type dimension on the 417-paper sample using the definition in Section 3.2.2 (at least one human-labeled category correctly predicted) to resolve whether the true value is 0.6739 or 0.8657. If the held-out accuracy/F1 for the model-type dimension is substantially lower than the in-sample values, or if Lenient Accuracy is 0.6739 rather than 0.8657, the Section 4 analyses should be treated as illust","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that human-in-the-loop prompt constraints make LLM classification more accurate and consistent than unconstrained LLMs (Tables 1 vs 2), and that the resulting labels support a reliable systematic literature review. The load-bearing assumption is that the reported metrics are unbiased estimates of how the final prompts perform on the full corpus. This assumption fails because the same 1,000-paper sample is used both to iteratively refine prompts and to evaluate them. Sections 2.2.1 and 3.2.1 describe selecting 1,000 papers, manually labeling them, evaluating multiple LLMs, and using those results to improve prompts; Tables 3 and 4 then report accuracy/F1/consistency on the same sample (or the 417-paper subset derived from it). No held-out set is mentioned anywhere in Sections 2–3 or Appendix A. This creates an in-sample optimism problem: the prompts and model were selected to fit these particular labels, so the reported improvements over the unconstrained baseline are likely inflated. The problem is especially acute for the model-type dimension, where Table 4's text reports Lenient Accuracy as 0.6739 while the table itself reports 0.8657, and micro-F1 is only 0.6586 with Jaccard 0.5545. If the true held-out performance is closer to the lower values, the Section 4 model-type proportions, chord diagram, and category-specific PageRank networks—all built from these labels—are not reliable. Because the entire empirical contribution depends on generalization of the refined prompts, the missing held-out evaluation is the single most load-bearing weakness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LR-Robot, a supervised human-in-the-loop framework for systematic literature reviews using large language models. The framework defines four classification dimensions (whether a paper develops/compares pricing models; underlying asset type; option type; pricing-model type), uses a 1,000-paper sample to compare five LLMs and refine prompts, and then applies the selected model to 11,916 option-pricing papers. Tables 1–4 report that human-supplied constraints improve accuracy, F1, and self-consistency relative to unconstrained LLMs. Section 4 uses the resulting labels to present frequency tables, a chord diagram, citation networks, and topic-evolution analyses. The central claim is that expert-constrained LLM classification is accurate and consistent enough to support valid field-level literature review.","tokens_in":13304,"tokens_out":4944,"duration_ms":53530,"significance":"The framework is concrete and addresses a real need. The paper's strengths include explicit prompt design in the appendices, evaluation of multiple LLMs, a with/without-constraint comparison, and reporting of precision, recall, F1, and self-consistency. If the reported performance generalized to new papers, LR-Robot would be a useful practical contribution to AI-assisted systematic reviews. However, the empirical support is not yet convincing: the evaluation protocol uses the same 1,000-paper sample both to refine prompts and to report accuracy, with no held-out set; the model-type dimension has low F1/precision and contains a direct numerical inconsistency; and the full-corpus analyses in Section 4 inherit these label-quality issues. The contribution is potentially publishable, but the validation protocol and reporting need substantial revision.","major_comments":[{"comment":"The evaluation protocol is in-sample. Section 2.2.1 states that the human-in-the-loop loop evaluates models on a representative sample and 'informs refinements to the prompts'; Section 3.2.1 then reports accuracy/F1/consistency on the same 1,000-paper sample used to select the best model and refine the constraints. Tables 1–4 are therefore optimistic estimates of final-prompt performance, not independent estimates. This is load-bearing for the central claim that human constraints improve classification accuracy and that the Section 4 analyses are reliable. Please provide a held-out evaluation set (or cross-validation) that is not used in any prompt/model selection, and report metrics on that set.","section":"§2.2.1 and §3.2.1, Tables 1–4"},{"comment":"There is a direct numerical inconsistency in the model-type evaluation. The text reports 'Lenient Accuracy is 0.6739' while Table 4 reports 'AI vs. Human (Lenient Accuracy) 0.8657'. These values lead to opposite interpretations of the model-type classification quality. A similar smaller discrepancy appears in §3.2.2 (text: self-consistency 0.9517; Table 3: 0.9594). Please reconcile both, define Lenient Accuracy explicitly, and ensure the text and tables report identical numbers.","section":"§3.2.4, Table 4"},{"comment":"The full-corpus statistical claims rest on labels whose model-type dimension has micro-F1 0.6586, precision 0.5505, and mean Jaccard similarity 0.5545 (Table 4). Even if these estimates are unbiased, a large fraction of model-type labels is wrong, so the category proportions in Table 7, the chord diagram in Fig. 3, and the category-specific PageRank analyses in Table 9 may change materially with label noise. The paper should include a sensitivity analysis, confidence intervals for proportions, or a comparison using only high-confidence labels before presenting these results as field-level facts.","section":"§4.1–§4.2"},{"comment":"The first-stage classification selects 417 positive papers from the 1,000-paper sample, and dimensions 2–4 are evaluated only on those 417 papers. Section 4 then applies the same prompts to all 5,942 papers that pass the first-stage filter. The accuracy estimates for dimensions 2–4 may not transfer from the 417-paper subsample to the full 5,942-paper set if the positive set is heterogeneous. Additionally, the text calls the full-corpus first-stage proportion of 49.86% 'close' to the sample estimate of 41.7%, an 8.2-percentage-point gap; a confidence interval or hypothesis test is needed. Please report dimension 2–4 evaluation on an independent sample drawn from the full positive set.","section":"§4.1"}],"minor_comments":[{"comment":"The column header 'A verage Accuracy' should read 'Average Accuracy'; the same typo appears in both tables.","section":"Tables 1–2"},{"comment":"'similiar' should be 'similar' in the sentence about the pattern of model type (5).","section":"§4.3"},{"comment":"'toxonomy' should be 'taxonomy' in the prompt text.","section":"Appendix A.3"},{"comment":"The metrics 'Lenient Accuracy' and 'Mean Jaccard similarity' are used without definitions, while the Micro-F1 and Sample-F1 formulas are given in a footnote. Please define all evaluation metrics in one place.","section":"§3.2.2–§3.2.4"},{"comment":"The statement 'Data and code will be available upon request' is not sufficient for reproducibility, especially because prompt versions and model outputs drive the results. Please provide a repository or DOI with the code, prompts, model outputs, and evaluation labels.","section":"Data Availability"},{"comment":"Figures 1–5 are captioned but the images are not present in the manuscript text. Ensure the final submission includes the actual figures.","section":"Figures"},{"comment":"The term 'Real-Time' is used in the title and abstract, but the pipeline performs daily updates rather than true real-time processing. Consider clarifying that user-mode querying is real-time while corpus updates are periodic.","section":"Title/Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's main weakness is methodological: the reported evaluation does not separate prompt refinement from assessment, and the model-type numbers are internally inconsistent. These are fixable within the scope of the paper if the authors can supply a held-out evaluation, reconcile Table 4, and temper the Section 4 claims accordingly. I see no fundamental flaw in the proposed framework, but the current evidence does not justify acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The framework itself is a reasonable integration of known components: human-in-the-loop prompt refinement, LLM classification, RAG, citation networks, and topic-evolution analysis. The option-pricing case study shows a plausible workflow, and the appendix gives enough detail to see exactly what the prompts do. That is real work and worth acknowledging.\n\nThe soft spot is the evaluation. The same 1,000-paper sample is used both to iteratively refine the prompts and then to report accuracy. That makes the headline comparison (Tables 1 vs. 2) an in-sample result: the prompts were tuned to fit those particular labels. There is no held-out set anywhere in Sections 2–3. The claim that human-in-the-loop guidance improves classification accuracy is plausible, but the numbers as reported are not credible estimates of out-of-sample performance. This is the load-bearing issue, and it affects the Section 4 statistics built on these labels.\n\nThere is also a concrete reporting inconsistency: for model types, the text says Lenient Accuracy is 0.6739 while Table 4 reports 0.8657. That is the kind of discrepancy a referee would catch, and it matters because the model-type dimension has the weakest agreement with human labels (micro-F1 0.66, Jaccard 0.55). The full-corpus model-type proportions and network analyses are only as good as these labels, and the evidence suggests they are moderately noisy.\n\nMissing code and data make it harder to check the numbers, though the authors do say materials will be available upon request.\n\nAll of this is fixable. The paper deserves peer review because the framework is useful and the idea is worth testing properly. A referee should ask for a held-out evaluation, clarification of the metric inconsistency, and ideally code/data release. For a reader working on LLM-based review tools, this is a worthwhile workflow reference, but not yet a source of reliable accuracy evidence.","headline":"A sensible framework for LLM-assisted literature review, but the reported accuracy gains rest on an in-sample evaluation with no held-out set and an internal metric inconsistency.","tokens_in":13804,"tokens_out":1474,"would_cite":true,"duration_ms":16562,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human-guided prompts lift LLM accuracy and consistency in systematic literature reviews, claims a new supervised framework.","keywords":["systematic literature review","large language models","human-in-the-loop","retrieval-augmented generation","option pricing","classification","citation network","topic evolution"],"falsifier":"Take a fresh random sample of 1,000 abstracts from the same option-pricing corpus, label them by human experts, run the final constrained prompts, and compare accuracy, F1, and Jaccard similarity to Tables 1, 3, and 4. A substantial drop would indicate the tuning sample did not generalise. Separately, reconcile the Model Types Lenient Accuracy values stated in the text (0.6739) and in Table 4 (0.8657); only one can be right.","tokens_in":12837,"feed_emoji":"🤖","tokens_out":4246,"duration_ms":37129,"temperature":0.7,"pith_summary":"The paper introduces LR-Robot, a supervised framework for systematic literature reviews, arguing that letting a human expert refine and constrain LLM prompts yields more accurate and more consistent classification of research papers than unconstrained LLMs. It evaluates five LLMs on a labeled sample of 1,000 option-pricing papers and reports that guided models outperform unguided ones on accuracy, F1, and self-consistency. Using the best model, the framework classifies the full 11,916-paper corpus along four dimensions, builds citation networks, and traces topic evolution, producing a field-level picture of option pricing research. The point of the framework is to speed up the labor-intensive stages of a review while keeping human interpretive control.","feed_headline":"Supervised LLM prompts beat unguided ones","feed_subtitle":"New framework claims faster, accurate systematic reviews; case study covers all 11,916 option-pricing papers.","key_machinery":"The load-bearing component is the human-in-the-loop prompt-refinement cycle: an LLM drafts sub-review tasks, a human researcher approves and constrains the instructions, the constrained prompts are evaluated on a labeled sample, and the feedback is used to improve the prompts before deployment on the full corpus. Its companion parts are the RAG database that stores all outputs for downstream tasks, and the four-layer architecture that separates developer-mode evaluation from user-mode queries.","core_discovery":"The central claim is that human-in-the-loop supervision—expert-written constraints appended to LLM prompts—measurably improves the reliability of AI-based classification of academic abstracts. On a 1,000-paper sample from the option pricing literature, guided Gemini Flash 2.0 reports accuracy 0.8327 and F1 0.8152 versus 0.7281 and 0.7419 without constraints, with self-consistency rising from 0.905 to 0.947. The same model, applied to 11,916 papers, labels 49.86% as model-development or comparison studies and produces the field statistics, network rankings, and evolution curves reported in Section 4. The paper presents LR-Robot as a practical route to real-time systematic reviews that preserv","pith_inferences":["The claim that expert-guided prompts 'preserve interpretive accuracy' would be stronger with a held-out test set; since the same 1,000-paper sample drives both prompt tuning and reported metrics, the published accuracy likely overstates generalisation.","The model-type dimension is the weakest link: the text and Table 4 give different values for lenient accuracy (0.6739 vs 0.8657) and the Jaccard similarity is 0.5545; if model-type labels are noisy, the co-occurrence and topic-evolution analyses built on those labels inherit the noise.","The framework's portability to other fields depends on the cost and quality of expert-labeled samples; for a new domain, a human must still build the constraint set and validate it.","A testable extension: use the same supervised prompt protocol on a different financial corpus, or on a set of papers with a known gold-standard taxonomy, and compare the accuracy gain from supervision."],"forward_implications":["If the reported numbers hold, supervised instruction gives a practical protocol for making LLM-based literature classification both faster and more trustworthy than unconstrained prompting.","The framework can produce multidimensional categorizations (e.g., underlying asset, option type, model type) that feed citation networks and temporal trend analysis.","The option-pricing case study yields concrete field-level claims: analytical and numerical models dominate the literature; machine-learning approaches are emergent from the 2010s onward.","The framework's daily update protocol means the review can stay current, addressing a known weakness of traditional systematic reviews."],"fun_headline_variants":["Supervised LLM prompts lift review accuracy to 0.83","Human-in-the-loop constraints boost AI literature reviews","Guided LLMs beat unguided ones in paper classification","LR-Robot: expert supervision improves AI systematic reviews","Supervised prompts raise LLM review F1 by 7 points"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework's reported accuracy is measured on the same 1,000-paper sample used to refine the prompts, with no held-out set, so the full-corpus statistics assume those tuning-phase numbers generalise to all 11,916 papers.","fun_headline_variants_meta":{"raw":{"variants":["Supervised LLM prompts lift review accuracy to 0.83","Human-in-the-loop constraints boost AI literature reviews","Guided LLMs beat unguided ones in paper classification","LR-Robot: expert supervision improves AI systematic reviews","Supervised prompts raise LLM review F1 by 7 points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1224,"prompt_tokens":779,"completion_tokens":445,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":363}},"tokens_in":523,"tokens_out":445,"duration_ms":4164,"temperature":1.0,"reasoning_tokens":363,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:55:26.013744+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh random sample of 1,000 abstracts from the same option-pricing corpus, label them by human experts, run the final constrained prompts, and compare accuracy, F1, and Jaccard similarity to Tables 1, 3, and 4. A substantial drop would indicate the tuning sample did not generalise. Separately, reconcile the Model Types Lenient Accuracy values stated in the text (0.6739) and in Table 4 (0.8657); only one can be right.","supporting_citations":[],"review_version":1}