{"id":"f95b2a42-4408-4a0d-a4a8-ee797c152eaa","arxiv_id":"2504.18180","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"DPO and RLHF preference training give inconsistent gains for Icelandic legal summarization, and ROUGE scores clash with expert human evaluation.","lead":"This paper tests whether two preference-based training methods, DPO and RLHF, produce better Icelandic legal summaries than standard supervised fine-tuning. It finds that these methods help some models and hurt others, and that automated ROUGE scores often disagree with legal expert judgments.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 8 contradicts the abstract's general claim: legal accuracy improves for only one of three preference-trained variants, and the pooled mean is flat; the central claim should be weakened to 'inconsistent'.","rationale":"The reader's conditional verdict and recommended revision are sound. My analysis supports the same bottom line, but the load-bearing concern is slightly different from the reader's stated weakest assumption. The reader emphasizes the reliability of a single legal expert scoring 25 rulings; that is a real limitation. However, the paper's central claim fails even before questioning the expert's reliability: the reported scores in Table 8 improve for only one of three preference-trained variants (Llama2-DPO versus Llama2-SFT), while GPT-SW3-RLHF and GPT-SW3-DPO both score below GPT-SW3-SFT on legal accuracy. The pooled mean is flat. Thus the abstract's general statement that preference training improves legal accuracy is an overgeneralization of the paper's own data. This is an internal consistency issue, not a question of whether the expert raters are representative. The ROUGE-based construction of the preference dataset is a related weakness: it means the 'preference' signal in training was not human feedback, further limiting the generalizability of any claim about RLHF or DPO improving legal quality. Still, the primary fix is to revise the abstract and conclusions to say preference training helps inconsistently, which is exactly what the reader recommended. I therefore keep the verdict at CONDITIONAL rather than moving it, and I partially agree with the reader's weakest assumption because the expert-sample concern is secondary to the direct contradiction in the reported results.","tokens_in":12380,"tokens_out":6831,"duration_ms":72360,"concrete_test":"Recompute the primary expert's per-ruling legal-accuracy scores for the 25 rulings behind Table 8 and compare SFT versus preference-trained variants with a paired nonparametric test, such as Wilcoxon signed-rank, separately for each model pair and pooled across all preference variants. Also report the aggregate mean legal accuracy before and after preference training. If neither the pooled comparison nor a majority of per-model comparisons shows a significant positive difference, the abstract must be revised to state that preference training improves legal accuracy inconsistently and only for some models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is that the paper's headline conclusion is not supported by its own reported evaluation. Table 8 gives legal accuracy scores: GPT-SW3-SFT 2.68, GPT-SW3-RLHF 2.56, GPT-SW3-DPO 1.96, Llama2-SFT 2.04, and Llama2-DPO 2.52. Preference training therefore improves legal accuracy in only one of three comparisons, and the mean legal accuracy across preference-trained variants (2.35) is essentially equal to the SFT mean (2.36). The abstract states that 'preference training improves the legal accuracy of generated summaries over standard fine-tuning' as a general result, but the empirical basis is a single favorable pairwise comparison. This concern does not depend on rater noise: even if the primary expert's scores are taken as ground truth, the data contradict the general claim. The paper's own conclusion in Section 6 is more cautious, saying the effect is 'not consistently' observed, so the abstract overstates the finding. A related issue is that the preference dataset was constructed using ROUGE-based top responses rather than human preferences, as acknowledged in Section 6, so the 'preference training' studied here partly aligns to an automatic metric rather than to human judgment.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether preference-based training (RLHF and DPO) improves the quality of Icelandic legal text summarization over standard supervised fine-tuning (SFT). The authors take two base models (GPT-SW3-1.3B and Llama2-7B), further pre-train them on domain-specific Icelandic legal text, instruction-tune them on court rulings and their human-written summaries, and then apply DPO or RLHF. Evaluation uses perplexity, ROUGE scores on a 300-item test set, and a human expert evaluation of 25 rulings (with two additional experts for five rulings) covering ranked preferences, Icelandic language quality, and legal accuracy. The abstract claims that preference training improves legal accuracy over SFT but does not improve overall Icelandic quality, and that automated metrics diverge from human judgment. The paper's own Section 6 more cautiously states that improvements are 'not consistently' observed. The core empirical contribution is the head-to-head comparison of DPO and RLHF in a low-resource language with expert human evaluation, but the reported evidence only partially supports the stated headline conclusion.","tokens_in":12720,"tokens_out":2958,"duration_ms":29689,"significance":"If the results are interpreted carefully, the paper makes a useful empirical contribution to legal NLP for low-resource languages. Strengths include the use of professional legal experts for evaluation, the direct comparison of DPO and RLHF on the same task, and the demonstration that ROUGE improvements need not correspond to expert-preferred legal accuracy (Tables 9 and 10). The study also provides a practical example of the instability of PPO training and the need for reward normalization. However, the significance is limited by the small human-evaluation sample, the lack of statistical testing, and a headline claim that outruns the data. The paper would be more valuable if the conclusions were explicitly restricted to the observed pairwise comparisons and presented as a case study rather than as a general finding about preference training for legal summarization.","major_comments":[{"comment":"The abstract states that 'preference training improves the legal accuracy of generated summaries over standard fine-tuning', but Table 8 does not support this as a general claim. In the three pairwise comparisons, legal accuracy improves only for Llama2-DPO over Llama2-SFT (2.52 vs. 2.04); it decreases for GPT-SW3-RLHF (2.56 vs. 2.68) and for GPT-SW3-DPO (1.96 vs. 2.68). The pooled mean legal accuracy across preference-trained variants (2.35) is essentially equal to the SFT mean (2.36). The paper's own Section 6 acknowledges that the effect is 'not consistently' observed. The abstract and Section 1 need to be revised to reflect the actual pattern, for instance by reporting the inconsistent direction and noting that the only positive result is for Llama2.","section":"Abstract and Table 8"},{"comment":"The preference dataset was constructed by selecting responses with top ROUGE scores after SFT, not by human preference judgments, as acknowledged in Section 6. This creates a self-reinforcing loop between the DPO training signal and the ROUGE evaluation metric, and it weakens the claim that the study evaluates alignment with 'user preferences.' The paper should either label this as 'metric-based preference training' or provide a human-preference benchmark for the preference dataset itself. This is a load-bearing issue because the central ROUGE-vs-expert discrepancy is partly explained by the fact that the training signal was already ROUGE-oriented.","section":"Sections 3.4 and 6"},{"comment":"The human evaluation is based on a single primary expert scoring 25 rulings, with only five rulings scored by two additional experts. No inter-annotator agreement statistic (e.g., Cohen's kappa) is reported, and Table 7 shows nontrivial disagreement (e.g., GPT-SW3-DPO average rank 4.2 for the primary expert vs. 3.2 for the other experts). With this small, partly inconsistent sample, the paper's claims about 'legal accuracy' and 'Icelandic language usage' are fragile. The authors should report per-ruling variance, agreement measures, and ideally confidence intervals or at least acknowledge the limited reliability of the ground truth. This concern applies directly to the abstract's second claim that preference training 'does not significantly enhance' Icelandic quality, as an absence of a significant difference in a 25-item sample cannot be interpreted as evidence of no effect.","section":"Section 3.5 and Tables 6-8"}],"minor_comments":[{"comment":"The text says the model was 'pre-trained on the ICG sub-corpus' but the corpus is the Icelandic Gigaword Corpus (IGC); this appears to be a typo.","section":"Section 4"},{"comment":"The sentence 'The same limitations can also be observed in in Table 10' contains a duplicated 'in'.","section":"Section 5.2"},{"comment":"The column heading 'Comparison' is unclear; it should specify whether it is the average of the two additional experts or some other pooled measure.","section":"Table 7"},{"comment":"Figure 1 is referenced in Section 4.2.2 but is not present in the provided manuscript text; please ensure the figure is included in the final submission.","section":"Figure 1"},{"comment":"No confidence intervals or significance tests are reported for the ROUGE differences or the expert score differences. Given the small evaluation sets, adding bootstrap confidence intervals or at least standard deviations would help readers judge the stability of the reported differences.","section":"Tables 3-5 and 8"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the empirical setup is honest, but the abstract overstates the findings and the human evaluation is too small to support the general claims. The revision should temper the conclusions, provide uncertainty measures, and explicitly acknowledge the ROUGE-based construction of the preference dataset. Assuming the authors make these changes, the paper could be a solid empirical contribution for low-resource legal summarization."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it, and the stress-test note is correct. The abstract says preference training improves legal accuracy, but Table 8 shows only Llama2-DPO beats its SFT baseline; both GPT-SW3 variants get worse. Pooled means are flat. Section 6 says \"not consistently,\" so the abstract overstates. That needs fixing.\n\nWhat the paper actually contributes: it is the first DPO/RLHF vs SFT comparison for Icelandic legal summarization I know of, with domain-expert evaluation. The qualitative failure modes in Section 5.3 are concrete and useful. The GPT-SW3-DPO inversion—highest ROUGE, worst expert preference—is a nice cautionary case. They also report RLHF instability honestly and link the dataset.\n\nSoft spots: the preference dataset was built from top-ROUGE responses, not human preferences (acknowledged in Section 6), so part of the ROUGE gain is circular. Human evaluation rests on one expert scoring 25 rulings, with two extra experts on only 5; no agreement metrics. No confidence intervals or significance tests anywhere. For a low-resource language with limited expert access, these are understandable but they weaken the central claim.\n\nThe central discrepancy between automated metrics and expert judgment survives even with rater noise. I would send this to peer review, but require the abstract to match the data and a clearer description of how the 25 rulings were selected. Adding uncertainty estimates would help.\n\nWho should read it: people working on low-resource legal NLP or applying DPO to domain-specific summarization. I would cite it for the ROUGE/expert divergence example, and it is a reasonable reading group pick.","headline":"The abstract overclaims: preference training improves legal accuracy in only one of three pairwise comparisons, but the paper is an honest, useful low-resource study.","tokens_in":724,"tokens_out":1213,"would_cite":true,"duration_ms":34196,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that preference training (DPO and RLHF) improves the legal accuracy of Icelandic legal summaries without improving the quality of the Icelandic language itself, and that ROUGE scores can rank models opposite to…","keywords":["Icelandic legal text summarization","Direct Preference Optimization","Reinforcement Learning from Human Feedback","legal language models","ROUGE evaluation","human evaluation","low-resource language NLP","domain-specific pre-training"],"falsifier":"Take the same trained models and have a panel of at least five legal experts rate a random sample of 100 or more summaries, blind to model identity, on separate legal-accuracy and Icelandic-quality scales; if DPO or RLHF models do not systematically beat the supervised-only baselines on legal accuracy, or if ROUGE rankings and expert rankings agree once more summaries are rated, the paper's central claim would fail.","tokens_in":12199,"feed_emoji":"⚖️","tokens_out":6103,"duration_ms":54490,"temperature":0.7,"pith_summary":"This paper asks whether preference-based training—RLHF and DPO—can make language models generate Icelandic legal summaries that legal professionals prefer over summaries from plain supervised fine-tuning. Using two open models further pre-trained on Icelandic court rulings and then instruction-tuned on lawyer-written summaries, the authors find that preference training improves the legal accuracy of generated summaries but does not consistently improve the quality of the Icelandic itself. The paper also reports a sharp split between automated ROUGE scores and expert rankings: the model with the highest ROUGE was ranked last by a legal expert, while the expert's top choices had lower ROUGE. The upshot is that for low-resource legal language, human qualitative evaluation is not merely a supplement to automated metrics but a necessary check, and language-specific pre-training remains the main lever for language quality.","feed_headline":"Preference training boosts legal accuracy, not Icelandic style","feed_subtitle":"An expert panel ranked the lowest-scoring ROUGE model first, so automated metrics alone mislead.","key_machinery":"The load-bearing mechanism is a three-stage training pipeline. Stage one further pre-trains a base model on Icelandic court rulings only (next-token prediction on packed 512-token blocks); stage two is supervised instruction fine-tuning on court rulings paired with lawyer-written summaries; stage three is preference training on a pairwise dataset built from the stage-two model's highest-ROUGE outputs. DPO converts reward maximization into a classification loss over preferred and rejected pairs, while RLHF trains a reward model on the same pairs and optimizes the policy with PPO, with reward normalization (zero mean, unit variance) needed to stop the policy from exploiting the reward model. Perplexity on legal text and ROUGE on the test set measure progress at each stage, and a legal expert's 1-to-5 scores for legal accuracy and Icelandic quality provide the final judgement.","core_discovery":"The central claim is that applying DPO or RLHF on top of domain-specific pre-training and supervised instruction fine-tuning improves the legal accuracy of Icelandic court-summary generation without a corresponding improvement in general Icelandic language quality. In the authors' own evaluation, the RLHF-tuned GPT-SW3 model and the supervised-only GPT-SW3 model were the expert's most-preferred outputs, both clearly ahead of the DPO variant; meanwhile GPT-SW3-DPO achieved the highest ROUGE score yet received the lowest expert legal-accuracy score, and Llama2-DPO improved legal accuracy over Llama2-SFT while leaving Icelandic quality essentially unchanged. The paper interprets this as evidence that preference training can add domain precision, but that ROUGE is not a trustworthy proxy for legal-expert preference and that language-specific pre-training, not preference optimization, is what determines written Icelandic quality.","pith_inferences":["If the preference pairs had been selected by expert preference instead of by ROUGE at stage three, the DPO and RLHF gains in legal accuracy might be larger; the authors themselves suggest this but did not test it.","The ROUGE-expert gap may be intrinsic to court summaries, which are concise outcome descriptions with low n-gram overlap with the ruling; under such conditions ROUGE measures lexical similarity rather than legal correctness.","The same pattern may hold in other low-resource languages with strict register requirements: preference training helps only after a strong language-model base, so language-specific pre-training investment should precede alignment work."],"forward_implications":["In Icelandic legal summarization, DPO or RLHF should be treated as targeted additions for legal accuracy, not as fixes for language quality.","ROUGE alone is insufficient for model selection in this setting; expert preference can invert ROUGE rankings.","Language-specific pre-training sets the ceiling on Icelandic quality: GPT-SW3's outputs outranked the larger Llama2 variants on expert language scores.","DPO is cheaper and easier to stabilize but prone to overfitting; RLHF is harder and compute-heavy yet produced the expert's top-ranked outputs.","All model outputs remained far below human-written summaries, particularly on legal accuracy."],"supporting_citations":[{"why":"Defines DPO, the classification-based preference optimization used in phase three.","marker":"Rafailov et al., 2024"},{"why":"Supplies the RLHF-with-human-feedback summarization approach and the PPO recipe adapted here.","marker":"Stiennon et al., 2020"},{"why":"Shows instruction-following models trained with RLHF are preferred over larger baselines, motivating the preference-training comparison.","marker":"Ouyang et al., 2022"},{"why":"Defines ROUGE, the automated metric used to select preference pairs and evaluate all model variants.","marker":"Lin, 2004"},{"why":"Provides GPT-SW3, the Nordic-pretrained 1.3B model that produced the expert-preferred summaries.","marker":"Ekgren et al., 2024"},{"why":"Provides Llama2-7B, the English-pretrained base model compared against GPT-SW3.","marker":"Touvron et al., 2023"},{"why":"Supplies LoRA, the low-rank adapter method used in all training phases.","marker":"Hu et al., 2022"},{"why":"Provides the DPO training recipe, including epoch counts and learning rate, adopted for the DPO runs.","marker":"Tunstall et al., 2024"},{"why":"Provides the Icelandic Gigaword Corpus sub-sample used to create Ice-Llama2.","marker":"Barkarson et al., 2022"},{"why":"Shows domain-specific pre-training improves legal NLP, the motivation for phase-one training.","marker":"Chalkidis et al., 2020"}],"fun_headline_variants":["Preference training sharpens legal accuracy, not Icelandic prose","For Icelandic legal summaries, ROUGE misleads human experts","DPO refines legal terms, but pre-training owns the language","Experts overrule ROUGE in Icelandic legal summary test","Legal precision gains from RLHF, but Icelandic style stays put"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's conclusion rests on one legal expert's ratings of 25 court-ruling summaries, with only five rulings independently checked by two other experts; if those sparse ratings are not representative of the 300-ruling test set, the reported ROUGE-versus-expert discrepancy may reflect rater idiosyncrasy rather than a real property of the models.","fun_headline_variants_meta":{"raw":{"variants":["Preference training sharpens legal accuracy, not Icelandic prose","For Icelandic legal summaries, ROUGE misleads human experts","DPO refines legal terms, but pre-training owns the language","Experts overrule ROUGE in Icelandic legal summary test","Legal precision gains from RLHF, but Icelandic style stays put"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1339,"prompt_tokens":840,"completion_tokens":499,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":456,"tokens_out":499,"duration_ms":5072,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:21:52.398709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same trained models and have a panel of at least five legal experts rate a random sample of 100 or more summaries, blind to model identity, on separate legal-accuracy and Icelandic-quality scales; if DPO or RLHF models do not systematically beat the supervised-only baselines on legal accuracy, or if ROUGE rankings and expert rankings agree once more summaries are rated, the paper's central claim would fail.","supporting_citations":[{"cited_title":"Manning, and Chelsea Finn","cited_arxiv_id":null,"evidence_quote":"Defines DPO, the classification-based preference optimization used in phase three."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides GPT-SW3, the Nordic-pretrained 1.3B model that produced the expert-preferred summaries."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Icelandic Gigaword Corpus sub-sample used to create Ice-Llama2."}],"review_version":1}