{"id":"149f877f-2de5-4c44-bc7f-1af715f368b1","arxiv_id":"2412.02549","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Patent-CR provides the first English patent claim revision dataset, and benchmark results show current LLMs, including GPT-4, cannot yet revise claims to examination standard.","lead":"The paper introduces Patent-CR, a dataset of 22,606 pairs of European patent applications and their granted claims, and evaluates LLMs on the new task of revising patent claims. It finds most models produce ineffective edits, GPT-4 scores highest but still below professional standard, and GPT-4-based evaluation aligns best with human judgments.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"A1/B1 claim pairs may not be a well-posed revision task; the granted target often depends on prosecution history and disclosure content absent from the source.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing point: A1/A2-to-B1 pairing is not a clean revision relation because prosecution is an interactive, externally informed process. I agree, and I add a concrete operationalization: the paper's own five-way edit taxonomy includes content amendment that draws on information outside the claim set, which means the target is not a function of the provided source. This is more fundamental than the small-sample human evaluation issue because it affects the validity of the dataset itself and therefore every downstream ranking and metric. I still would not reject the paper: the corpus could be repurposed as a parallel claims corpus, and the empirical findings are described as preliminary. The conditional verdict should stand, but the release should include or be accompanied by an amendment-history validation study. The proposed concrete test would settle whether the concern is real: if most pairs are directly derivable from A1 by renumbering, merging, and deleting existing claims, then the task is well-posed and the current benchmark is acceptable; if a large share require external disclosure, the benchmark design needs revision before its central claims can be relied upon.","tokens_in":16855,"tokens_out":4004,"duration_ms":46718,"concrete_test":"Sample 100 random Patent-CR pairs. For each pair, retrieve the EPO register file wrapper via the OPS register API or EP Register and reconstruct the amendment history between A1/A2 and B1. Classify each pair as 'directly derivable' if B1 can be obtained from A1 by claim-internal edits, deletion, renumbering, or merging of existing A1 dependent claims; classify it as 'externally grounded' if the granted claims import features from the description/drawings or are shaped by a specific examiner objection recorded in the file wrapper. If the externally grounded fraction is substantial (e.g., >20%), the source-target pairs are not a well-posed revision task from the source alone, and the benchmark should either be re-scoped as parallel claim rewriting or be augmented with the intervening office actions and description text.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's central premise is that an EPO A1/A2 application claim set and its later B1 granted claim set form a source-target revision pair, with B1 usable as gold output. This premise is weakest because EPO prosecution is interactive: applicant amendments respond to examiner objections, and final claims frequently import features from the description, drawings, or newly constructed dependent claims rather than editing the A1 text. The paper's own taxonomy in Section 1 acknowledges this under 'Content amendment', where 'essential information missing in the draft is included' — information that cannot be inferred from the source claim set alone. If a substantial fraction of B1 claim sets are not recoverable from A1 by any deterministic edit sequence, then supervised fine-tuning and lexical metrics (SARI, BLEU, ROUGE-L) are trained and evaluated against an underdetermined target, while human evaluators judge models against a reference those models were never given enough information to produce. The result would not invalidate the parallel-corpus value of Patent-CR, but it would undercut the claim that the benchmark measures claim-revision ability, and low LLM scores could reflect task ill-posedness rather than model failure. The Limitations section does not address this pairing validity issue, so it remains the most load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Patent-CR, a dataset of 22,606 pairs of EPO A1/A2 application claim sets and corresponding B1 granted claim sets, framed as the first English resource for a new 'patent claim revision' task. The authors define a five-type taxonomy of revisions, report descriptive statistics, and evaluate ten systems (including Copy, general-purpose LLMs, a legal-domain model, fine-tuned variants, and GPT-4) using both professional human ratings on a 60-example subset and automated metrics (SARI, BLEU, ROUGE-L, BERTScore, and GPT-4-based G-Eval). They report that GPT-4 receives the highest human quality score (6.40/10), that fine-tuning improves base models, and that G-Eval correlates best with human judgments among automated metrics, while also concluding that all model outputs remain below the examination standard.","tokens_in":17084,"tokens_out":5744,"duration_ms":54042,"significance":"If the A1/B1 pairing is accepted as a valid revision task, Patent-CR is a valuable, large-scale resource in an underexplored domain, and the paper's reproducible OPS-based construction pipeline, clear taxonomy, explicit model versions, prompt disclosure, and released code are genuine strengths. The finding that patent claim revision differs from generic text revision (targets become more complex and less readable) is interesting and useful for the community. However, the empirical evaluation is severely underpowered for the comparative claims it makes, and the well-posedness of the A1-to-B1 revision task is not established; both issues are load-bearing for the paper's central claims.","major_comments":[{"comment":"The dataset's core assumption is that an EPO A1/A2 application claim set and the later B1 granted claim set of the same patent form a valid source–target revision pair, with B1 usable as gold output. This is not validated. The paper's own taxonomy in §1 includes 'content amendment', where 'essential information missing in the draft is included'; in real EPO prosecution that information typically comes from the description, drawings, or the applicant's response to examiner objections, none of which is present in the A1 claim text. Consequently, a substantial fraction of B1 claim sets may not be recoverable from A1 by any deterministic editing process, making the task underdetermined for supervised fine-tuning and for lexical metrics such as SARI, BLEU, and ROUGE-L. The paper should provide quantitative evidence on recoverability (e.g., the fraction of B1 claims or n-grams that have no source in A1, or alignment statistics) or explicitly reframe the resource as a parallel corpus of pre- and post-prosecution claim sets rather than as a well-posed revision benchmark. The Limitations section does not address this pairing-validity issue.","section":"§3.1 (Steps 1–2) and §1 (taxonomy)"},{"comment":"The human evaluation is too small to support the paper's ranking claims: only 60 examples in total, 6 per model, two raters, no reported inter-annotator agreement, no confidence intervals, and no significance tests. Differences such as GPT-4 at 6.40 versus SaulLM-7B-FT at 6.38 versus Llama-3.1-8B-FT at 6.03 are within plausible noise for n=6. The manuscript should either report per-item scores with bootstrap confidence intervals and inter-rater reliability, or present the human evaluation as qualitative and avoid comparative formulations such as 'GPT-4 outperforms other tested LLMs' (Abstract and §5.5). Without this, the central empirical claim is not supported.","section":"§4.2 and Table 3"},{"comment":"The correlation analysis between automated metrics and human judgments uses only 9 data points for G-Eval because GPT-4 is deliberately excluded from G-Eval evaluation. With n=9, a Spearman rho of 0.600 is not statistically significant at the conventional 0.05 level, and the paper reports no p-values, confidence intervals, or permutation-based intervals. The claim that 'GPT-4-based automated evaluation has the highest correlation with human judgment' (Abstract) is therefore unsupported. The authors should report significance tests or uncertainty estimates, or temper the claim accordingly.","section":"Table 5"}],"minor_comments":[{"comment":"The text states that GPT-4's feature-linkage score rises 'from 5.67 to 6.67', but Table 3 lists the Copy baseline linkage as 5.33 and GPT-4 as 6.33. The numbers should be corrected.","section":"§5.5"},{"comment":"The model descriptions refer to 'Llama-3-8B-Instruct' and 'Llama-3-70B-Instruct', but the rest of the paper consistently uses 'Llama-3.1-8B' and 'Llama-3.1-70B'. Please make the naming consistent.","section":"Appendix D"},{"comment":"The labels 'Claim before' and 'Claim after' are ambiguous; the B1 text is a granted patent, not merely a 'published' version. Suggest renaming to 'Application claims' and 'Granted claims' to avoid confusion, especially since A1/A2 documents are also published.","section":"Figure 2 and Table 2"},{"comment":"The G-Eval prompt instructs the model to rate 'draft claims' against 'referenced claims', but in the experiment G-Eval is used to score model-generated claims against the gold B1 text. The role of the input text should be clarified so that readers do not think the prompt matches the actual setup.","section":"Appendix E.3"},{"comment":"The phrase 'initial patent applications rejected by patent examiners' is inaccurate: A1/A2 documents are published applications that may have received objections but are not necessarily 'rejected'. Suggest rewording to 'applications as initially published'.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The dataset resource itself is potentially useful and the construction pipeline is transparent, but the empirical sections need substantial strengthening before the comparative claims are publishable. The manuscript also draws heavily on the authors' own related work (Jiang et al., 2025a,b,c; Jiang and Goetz, 2025); this is not inappropriate, but the novelty relative to those concurrent papers should be carefully checked by the editor. If the authors can supply a convincing recoverability analysis for the A1-to-B1 pairing and rework the human-evaluation and correlation statistics, a revised version could make a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jiang et al. bring us the first English dataset for patent claim revision, 22.6K paired EPO A1/A2 application claims and B1 granted claims. That resource is real and likely to be used; the collection pipeline is transparent and the release is a service to the patent-NLP community. The paper also makes a useful negative point: current LLMs do not produce claims that pass examination, and standard lexical metrics correlate poorly with professional judgments. Credit where due: the expert evaluation is a genuine step up from automated-only benchmarking.\n\nThe soft spots are the usual suspects but they are real. First, the stress-test on pairing validity concerns me. An EPO B1 claim set is the output of an interactive prosecution, not a single revision pass. Features get imported from the description or added in response to examiner objections, so the target is often underdetermined by the A1 text alone. The paper's own taxonomy (content amendment) admits essential information missing in the draft gets included. That means a model could be perfectly competent and still fail to reproduce B1, and the lexical metrics (SARI, BLEU, ROUGE-L) bake that underdetermination into the benchmark. The dataset keeps its value as a parallel corpus, but the framing as a revision task with B1 as gold needs a caveat the paper does not give.\n\nSecond, the human evaluation is too small to support the ranking claims. Sixty examples total, six per model, with two raters, no inter-annotator agreement reported, and no confidence intervals. GPT-4's 6.40 versus fine-tuned SaulLM's 6.38 is a tie for all practical purposes, yet the text says GPT-4 'stands out.' The correlation analysis in Table 5 has only ten model-level points; those coefficients are noisy. The test split being a single month is a minor quibble, but it should at least be justified.\n\nAll of that said, the central resource holds up. The task is new, the dataset is useful, and the paper is clearly written. I would send it to peer review, but the authors should be asked to add uncertainty to the empirical claims and to discuss the prosecution-history problem directly.","headline":"A genuinely useful patent-claim parallel corpus whose empirical ranking claims rest on a tiny expert-rated sample and an unexamined pairing assumption.","tokens_in":17568,"tokens_out":3001,"would_cite":true,"duration_ms":31626,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Patent claim revision gets its first English dataset, with 22,606 draft-to-grant claim pairs, and GPT-4 scoring 6.40 out of 10 in professional human evaluation, still below the examination bar.","keywords":["patent claim revision","dataset","large language models","text revision","European Patent Office","human evaluation","G-Eval","fine-tuning"],"falsifier":"Take a random sample of 100 Patent-CR pairs, give patent attorneys only the A1/A2 claims, and ask them to reproduce the B1 claims; if the added limitations and merged claims cannot be anticipated, then the target is not recoverable from the source and the supervised pairing cannot support the claimed task.","tokens_in":16685,"feed_emoji":"⚖️","tokens_out":5296,"duration_ms":48375,"temperature":0.7,"pith_summary":"This paper proposes a new task, patent claim revision, and introduces Patent-CR, the first English dataset built for it: 22,606 pairs in which pre-grant application claims (EPO A1/A2 publications) are the source and granted claims (B1) are the target. The motivation is that revising claims to withstand legal scrutiny, not just improve readability, is expensive and currently done by specialists. The empirical study evaluates ten model configurations against ratings from patent professionals and finds that most LLMs make ineffective edits, while fine-tuned domain models approach GPT-4's quality; GPT-4 scores 6.40 out of 10, still below the examination standard. The paper also argues that standard automated metrics mislead on this task, and that GPT-4-based G-Eval correlates most closely with human judgment.","feed_headline":"GPT-4 leads patent claim revision but still fails the legal bar","feed_subtitle":"First English dataset with 22,606 draft-to-grant claim pairs; top score 6.40/10 is below examination standard.","key_machinery":"The load-bearing object is the Patent-CR dataset itself, built by pairing the EPO A1/A2 application claims with the later B1 granted claims of the same patent, then filtering to 22,606 pairs. A taxonomy of five revision types (content amendment, term consistency, language precision, concision, renumbering) and a weighted quality formula, Quality = (Completeness*4 + Clarity*2 + Consistency*2 + Linkage*3)/11, carry the evaluation; the professional human evaluation on 60 selected examples and the G-Eval prompt that mirrors it are what the comparisons rest on.","core_discovery":"On the paper's own terms, the central claim is that patent claim revision is a distinct and harder text-revision task, and that a parallel corpus of EPO application and grant claim sets can support it. The dataset construction assumes the granted B1 claims are the gold revision: average claims drop from 13.85 to 10.66 per document while claim length and structural complexity rise, showing revision is about densifying and legally hardening claims rather than simplifying them. With professional human evaluation as the standard, GPT-4 produces the highest quality revision (6.40), fine-tuned SaulLM-7B nearly matches it (6.38), and most other models fail to beat a copy baseline; none reach examination standard. The paper further claims that lexical overlap metrics like BLEU and ROUGE partially track human judgments, SARI and BERTScore do not, and GPT-4-based G-Eval correlates most strongly (Spearman 0.600) with human quality scores.","pith_inferences":["If the B1-as-gold assumption is imperfect, then models trained on Patent-CR may learn examiner-driven narrowing and claim consolidation rather than general drafting skill, so downstream evaluation should test on real prosecution documents.","The five revision types suggest an auxiliary supervision signal: predicting the edit type per claim could sharpen revision models and make their errors more interpretable.","The finding that G-Eval aligns with human judgment could be extended to a cheaper screening pipeline: use G-Eval to rank candidate revisions and send only the top candidates to human experts.","One testable extension is to build matched pairs from other jurisdictions or from intermediate prosecution documents to see whether the gap between GPT-4 and examination standard is intrinsic or an artifact of the EPO pairing."],"forward_implications":["A shared benchmark now exists for training and comparing models on legal-grade claim revision, with the B1 claims as reference targets.","Fine-tuning a reasonably sized language model on patent claim pairs appears to be the most reliable path to improvement, since both fine-tuned models beat their base versions on every human criterion.","Automated metrics should be used with caution on this task; GPT-4-based G-Eval is the closest automatic proxy for professional judgment.","Because GPT-4 still scores 6.40 out of 10, no current LLM output can be trusted to pass examination without human expert post-editing.","The targets' increasing structural complexity and decreasing readability imply that revision models should not optimize for simplification."],"supporting_citations":[{"why":"Prior work on generating patent claims with GPT-2, the gap this task extends.","marker":"Lee and Hsiang (2020)"},{"why":"Provides the five evaluation criteria for claim quality and the finding that LLM claim generation is unsatisfactory.","marker":"Jiang et al. (2025c)"},{"why":"CoEdIT, the state-of-the-art text revision model used as baseline.","marker":"Raheja et al. (2023)"},{"why":"SaulLM-7B, the law-specific LLM evaluated here.","marker":"Colombo et al. (2024)"},{"why":"ITERATER dataset, the main text-revision baseline compared in Table 1.","marker":"Du et al. (2022)"},{"why":"CASIMIR, the scientific-article revision dataset contrasted with Patent-CR.","marker":"Jourdan et al. (2024)"},{"why":"GPT-4, the best-performing model and the basis of G-Eval.","marker":"(OpenAI, 2023)"},{"why":"G-Eval, the GPT-4-based automated evaluation with the highest human correlation.","marker":"(Liu et al., 2023)"}],"fun_headline_variants":["GPT-4 best at patent claim revision, still fails exam bar","New patent dataset: GPT-4 leads revision, no one passes exam","Patent claim revision dataset: LLMs can't meet legal standard","Domain-specific LLMs nearly match GPT-4 on patent claims","First patent claim revision dataset shows GPT-4's limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The granted B1 claims are treated as the gold revision of the A1/A2 application claims, even though real prosecution involves examiner objections and strategic choices whose traces are not in the application text alone.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 best at patent claim revision, still fails exam bar","New patent dataset: GPT-4 leads revision, no one passes exam","Patent claim revision dataset: LLMs can't meet legal standard","Domain-specific LLMs nearly match GPT-4 on patent claims","First patent claim revision dataset shows GPT-4's limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001135,"raw_usage":{"total_tokens":4718,"prompt_tokens":952,"completion_tokens":3766,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":3678}},"tokens_in":568,"tokens_out":3766,"duration_ms":28243,"temperature":1.0,"reasoning_tokens":3678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:19:20.745793+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 100 Patent-CR pairs, give patent attorneys only the A1/A2 claims, and ask them to reproduce the B1 claims; if the added limitations and merged claims cannot be anticipated, then the target is not recoverable from the source and the supervised pairing cannot support the claimed task.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CoEdIT, the state-of-the-art text revision model used as baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CASIMIR, the scientific-article revision dataset contrasted with Patent-CR."}],"review_version":1}