{"id":"524c0c6e-f5d2-482f-9d87-2436b8fe6e0a","arxiv_id":"2606.09334","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Gemini 3.1-Pro with Ukrainian minimal-edits + few-shot prompting reaches F0.5=69.22 on Ukrainian GEC, closing over 90% of the gap to fine-tuned SOTA at 73.14.","lead":"This paper evaluates prompting strategies with 11 commercial LLMs and one open-source model on the UNLP 2023 Ukrainian minimal-edit GEC benchmark. A smart generalist might read it to see how far general-purpose LLMs can go on low-resource language tasks without fine-tuning.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"UNLP 2023 benchmark + minimal-edit protocol may not reflect real-world Ukrainian GEC preferences or error distributions","rationale":"The reader's weakest_assumption is precisely the load-bearing one for the performance claim. No other internal inconsistency (e.g., prompt construction details or statistical reporting) is visible from the provided abstract that would supersede this. Full-text access would be needed only to check whether additional robustness experiments already address it; absent that, the assumption remains the primary risk.","tokens_in":1711,"tokens_out":353,"duration_ms":15092,"concrete_test":"Take the 100 highest-frequency error types from the UNLP 2023 test set plus 100 new Ukrainian sentences drawn from contemporary web/news text; obtain independent human minimal-edit annotations and human preference judgments (minimal-edit vs. fluent); recompute F0.5 for the published Gemini minimal-edit + few-shot prompt on both sets and measure rank correlation with human preference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Gemini 3.1-Pro F0.5=69.22 closes >90% of gap to fine-tuned SOTA 73.14) rests entirely on scores under the UNLP 2023 GEC-only benchmark and the authors' minimal-edit prompting/evaluation rules. The abstract itself notes that detailed minimal-edit instructions cause models to abandon low-frequency error categories and produce five recurring over-correction patterns tied to Ukrainian-specific phenomena (case, punctuation). If the benchmark's error mix or the strict minimal-edit preference diverges from what users actually want (more fluent rewrites, handling of low-frequency phenomena), the numerical gap closure does not support the implied practical conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper evaluates 11 commercial LLMs and one open-source Ukrainian model on the UNLP 2023 GEC-only benchmark using zero-shot, few-shot, minimal-edits, and LLM-assisted prompt optimization strategies. It reports that Gemini 3.1-Pro with Ukrainian minimal-edits + few-shot prompting reaches F0.5=69.22 (closing >90% of the gap to fine-tuned SOTA at 73.14), notes that Ukrainian instructions help only for Claude in zero-shot, identifies five recurring over-correction patterns, and observes that detailed minimal-edit rules improve punctuation/case but cause abandonment of low-frequency error categories. Code, prompts, and outputs are released.","tokens_in":1851,"tokens_out":577,"duration_ms":19502,"significance":"If the UNLP 2023 benchmark and minimal-edit protocol are accepted as representative, the result shows that carefully engineered prompting can nearly match fine-tuned performance for Ukrainian GEC without task-specific training, which is relevant for other low-resource languages. The public release of prompts and outputs supports reproducibility. The work also surfaces Ukrainian-specific linguistic phenomena in over-corrections.","major_comments":[{"comment":"Abstract: the headline claim that Gemini 3.1-Pro 'closes over 90% of the gap' to fine-tuned SOTA is presented as a primary result, yet the error analysis shows that the minimal-edit protocol causes models to abandon several low-frequency error categories. This makes the numerical gap-closure claim load-bearing only under the specific benchmark protocol and weakens the implied practical conclusion unless qualified.","section":"Abstract"},{"comment":"Results section (performance table): single-point F0.5 scores are reported without standard deviations, multiple runs, or statistical significance tests against the SOTA baseline. Given that the central claim rests on the 69.22 vs. 73.14 comparison, the absence of uncertainty estimates makes it impossible to judge whether the gap closure is reliable.","section":"Results"}],"minor_comments":[{"comment":"The abstract states 'Gemini 3.1-Pro'; confirm the exact model name and version against the experimental setup section for consistency.","section":"Abstract"},{"comment":"Error analysis identifies five over-correction patterns but does not quantify their frequency or contribution to the overall F0.5 drop; adding counts or a breakdown table would strengthen the section.","section":"Error Analysis"},{"comment":"The paper would benefit from an explicit limitations paragraph discussing how the minimal-edit preference may diverge from user expectations for fluency in real-world Ukrainian GEC.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and the recommendation of minor revision. We address each major comment below, agreeing that the abstract claim benefits from additional qualification and that the results reporting can be clarified.","responses":[{"response":"We agree that the gap-closure figure is protocol-specific and that the observed abandonment of low-frequency error categories (already detailed in the error analysis) represents an important trade-off. We will revise the abstract to qualify the primary claim by explicitly noting that the reported performance is achieved under the minimal-edit prompting protocol, which involves such category-specific trade-offs. This will better contextualize the headline result without altering the numerical findings.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the headline claim that Gemini 3.1-Pro 'closes over 90% of the gap' to fine-tuned SOTA is presented as a primary result, yet the error analysis shows that the minimal-edit protocol causes models to abandon several low-frequency error categories. This makes the numerical gap-closure claim load-bearing only under the specific benchmark protocol and weakens the implied practical conclusion unless qualified."},{"response":"We acknowledge that single-point estimates without uncertainty quantification limit assessment of the comparison's reliability. Our experiments used single runs primarily due to the prohibitive cost of repeated API calls across 12 models and multiple prompting configurations. We will add a clarifying sentence in the results section noting this practical constraint and the largely deterministic nature of the evaluated prompts. We also observe that the fine-tuned SOTA baseline is itself reported as a single point in the UNLP 2023 literature.","revision_made":"partial","referee_comment":"[Results] Results section (performance table): single-point F0.5 scores are reported without standard deviations, multiple runs, or statistical significance tests against the SOTA baseline. Given that the central claim rests on the 69.22 vs. 73.14 comparison, the absence of uncertainty estimates makes it impossible to judge whether the gap closure is reliable."}],"tokens_in":1436,"tokens_out":437,"duration_ms":14422,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper shows that commercial LLMs prompted with Ukrainian minimal-edit instructions can reach F0.5 of 69.22 on the UNLP 2023 benchmark, closing most of the gap to the fine-tuned SOTA at 73.14. Gemini 3.1-Pro leads, and Ukrainian-specific prompts outperform English ones for most models.\n\nWhat stands out is the scale: 11 commercial models plus one open-source Ukrainian model, tested across zero-shot, few-shot, minimal-edit, and optimization setups. They document five recurring overcorrection patterns tied to case and punctuation, and they release the prompts, code, and outputs. That release is the most immediately useful part.\n\nThe soft spot is the evaluation setup itself. The abstract notes that detailed minimal-edit rules make models drop low-frequency error categories, and the stress-test concern holds: if real users prefer fluent rewrites over strict minimal changes, these scores do not directly show practical value. No statistical significance or variance numbers are mentioned in the abstract, and commercial model versions change, so the exact ranking may not stick. The work stays empirical and avoids circular claims.\n\nThis is for researchers doing GEC on low-resource languages or testing prompting for structured output tasks. It adds concrete numbers where few existed and deserves peer review to check the methods and discuss how well the benchmark matches user needs.","headline":"Prompting gets close to fine-tuned results on Ukrainian GEC but the minimal-edit benchmark and protocol limit how far the numbers generalize.","tokens_in":2313,"tokens_out":345,"would_cite":false,"duration_ms":8780,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Ukrainian minimal-edit prompting with commercial LLMs closes over 90 percent of the gap to fine-tuned grammatical error correction systems.","keywords":["grammatical error correction","Ukrainian language","prompt engineering","large language models","minimal edits","few-shot learning","zero-shot learning","UNLP benchmark"],"falsifier":"Running the same prompted models on a freshly collected set of Ukrainian errors from authentic sources like forums or documents and measuring the F0.5 score under minimal-edit rules would confirm or refute the performance claims.","tokens_in":2606,"feed_emoji":"📝","tokens_out":657,"duration_ms":21213,"temperature":0.7,"pith_summary":"The paper tests whether prompting alone can handle minimal-edit Ukrainian grammatical error correction on commercial LLMs. Using the UNLP 2023 benchmark, it finds that Gemini 3.1-Pro with Ukrainian minimal-edits prompts and optimization reaches an F0.5 of 69.22, compared to 73.14 for the fine-tuned best system. Ukrainian instructions are necessary to express precise rules, and they work best when combined with few-shot examples and LLM help for prompt tuning. This shows prompting can substitute for fine-tuning in many cases but still misses some error categories and introduces specific overcorrections.","feed_headline":"Prompting closes 90% gap to fine-tuned Ukrainian GEC","feed_subtitle":"Gemini 3.1-Pro hits F0.5=69.22 with Ukrainian minimal-edit prompts versus 73.14 for trained models on UNLP 2023.","key_machinery":"Ukrainian minimal-edits prompts that specify language-specific correction rules, combined with few-shot examples and LLM-assisted optimization.","core_discovery":"Our best configuration using Gemini 3.1-Pro with LLM-assisted prompt optimization on minimal-edits and few-shot prompts achieves F0.5=69.22 on the UNLP 2023 GEC-only benchmark. This closes over 90% of the gap to the fine-tuned SOTA of F0.5=73.14. Zero-shot Ukrainian instructions help only Claude models, while all models perform best with Ukrainian minimal-edits prompts. Detailed instructions improve results on punctuation and case errors but lead models to ignore several low-frequency categories. Five recurring overcorrection patterns related to Ukrainian linguistic phenomena are identified in the error analysis.","pith_inferences":["Similar prompting strategies might allow rapid deployment of GEC for other low-resource languages without large training sets.","Hybrid systems could combine these prompts with targeted rules to address the observed overcorrection patterns.","The abandonment of low-frequency error categories suggests that prompting may require supplementary mechanisms for complete coverage.","Evaluating the same prompts on out-of-domain Ukrainian text would test whether the benchmark scores generalize."],"forward_implications":["Zero-shot prompts in Ukrainian improve performance only for Claude models among the tested LLMs.","Minimal-edits prompts in Ukrainian outperform those in English for every model tested.","LLM-assisted prompt optimization yields the single highest score when added to minimal-edits plus few-shot.","Detailed minimal-edits instructions produce the largest gains on punctuation and case errors.","Five recurring overcorrection patterns appear that are linked to Ukrainian-specific features."],"fun_headline_variants":["Gemini 3.1-Pro reaches F0.5 69.22 with Ukrainian prompts","Ukrainian minimal-edits prompts close 90% gap to GEC SOTA","Prompt optimization with minimal edits reaches UNLP GEC SOTA","Five overcorrection patterns identified in Ukrainian LLM GEC"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The UNLP 2023 GEC-only benchmark and its minimal-edit evaluation protocol match the distribution and desired behavior of real-world Ukrainian grammatical errors.","fun_headline_variants_meta":{"raw":{"variants":["Gemini 3.1-Pro reaches F0.5 69.22 with Ukrainian prompts","Ukrainian minimal-edits prompts close 90% gap to GEC SOTA","Prompt optimization with minimal edits reaches UNLP GEC SOTA","Five overcorrection patterns identified in Ukrainian LLM GEC"]},"model":"grok-4.3","cost_usd":0.01314,"raw_usage":{"total_tokens":5728,"prompt_tokens":729,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":131399500,"prompt_tokens_details":{"text_tokens":729,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4921,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":729,"tokens_out":78,"duration_ms":32601,"temperature":1.0,"reasoning_tokens":4921,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T16:32:26.462057+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same prompted models on a freshly collected set of Ukrainian errors from authentic sources like forums or documents and measuring the F0.5 score under minimal-edit rules would confirm or refute the performance claims.","supporting_citations":[],"review_version":1}