{"id":"b9629dcb-2c86-4898-a167-82231234f2a3","arxiv_id":"2606.19266","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Empirical comparison across model families finds SFT often sufficient for MCQA while CPT aids OEQA metrics in French medical LLM adaptation, with cross-lingual transfer observed.","lead":"This paper empirically compares continual pretraining, supervised fine-tuning, and their combination for adapting LLMs to French medical question-answering. A smart generalist might read it to learn cost-effective strategies when adapting models to specialized domains under limited compute.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's identification of the evaluation protocol as the weakest point aligns with the empirical nature of the work; the abstract and claim description show appropriate caution. No stronger, more specific technical vulnerability (e.g., in the CPT/SFT disentanglement or cross-lingual transfer) emerges from the provided summary.","tokens_in":1720,"tokens_out":267,"duration_ms":12149,"concrete_test":"Re-run the MCQA statistical significance tests (Table X, §4) after applying a multiple-comparison correction across all model-size pairs; if the fraction of 'not significant' CPT+SFT vs SFT comparisons changes by >20%, the cost-effectiveness recommendation requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical observation across model families, sizes, and initialization types, with explicit hedging that MCQA gains from CPT+SFT over SFT are small and often statistically insignificant. The OEQA findings similarly distinguish metric types and evaluation protocols. No internal inconsistency, hidden assumption in the disentanglement procedure, or unsupported leap from results to guidelines is apparent. The reader's weakest assumption (datasets/metrics/judge protocol) is a standard limitation of such studies rather than a load-bearing flaw in the reported comparisons.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents an empirical study of medical domain adaptation for French QA, comparing continual pretraining (CPT), supervised fine-tuning (SFT), and CPT+SFT across three model families, multiple sizes, and initialization types. It disentangles adaptation effects from base model choice and evaluates MCQA and OEQA using automatic metrics plus LLM-as-a-Judge under greedy and constrained decoding, plus cross-lingual transfer to English. Key findings are that CPT+SFT most often wins on MCQA but gains over SFT are small and frequently insignificant (making SFT a cost-effective default), while CPT improves overlap metrics on OEQA but SFT can degrade generation quality (with instruction tuning and CPT+SFT preferred by LLM judges).","tokens_in":1810,"tokens_out":531,"duration_ms":12561,"significance":"If the comparative results hold, the work supplies actionable guidelines for choosing adaptation strategies under compute constraints in specialized domains. Strengths include the explicit disentanglement of adaptation from base-model effects, the hedging on statistical significance of small gains, the distinction between metric types (overlap vs. LLM judge), and the cross-lingual transfer experiments; these elements make the ablation more informative than typical single-model studies.","major_comments":[{"comment":"Methods section: the abstract and results claim statistical significance tests and cross-lingual transfer, yet no details appear on dataset sizes, exclusion criteria, exact tests (e.g., paired t-test or bootstrap), multiple-comparison correction, or error bars; this information is load-bearing for the central claim that CPT+SFT gains over SFT are “frequently not statistically significant.”","section":"Methods / Results"},{"comment":"Evaluation protocol (OEQA subsection): reliance on LLM-as-a-Judge without a reported human validation subset or inter-annotator agreement on a medical sample risks systematic bias, especially since the paper notes imperfect correlation with human judgment; this directly affects the claim that instruction tuning and CPT+SFT are preferred.","section":"Evaluation / OEQA results"}],"minor_comments":[{"comment":"Table captions and axis labels should explicitly state the number of runs or seeds used for each reported score to allow readers to assess variability.","section":"Tables"},{"comment":"The cross-lingual transfer paragraph would benefit from a short statement on whether the English benchmarks were seen during any pretraining stage of the base models.","section":"Cross-lingual experiments"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment of the manuscript's significance and for the constructive feedback. We address each major comment below and will revise the manuscript accordingly.","responses":[{"response":"We agree that the statistical analysis details were insufficiently reported. The revised manuscript will add the missing information, including dataset sizes for each experiment and language, exclusion criteria for QA samples, the precise tests used (paired t-tests with bootstrap resampling), multiple-comparison correction (Bonferroni), and error bars or confidence intervals on all relevant figures and tables. These additions will directly support the claims about the (non-)significance of CPT+SFT gains over SFT alone.","revision_made":"yes","referee_comment":"[Methods / Results] Methods section: the abstract and results claim statistical significance tests and cross-lingual transfer, yet no details appear on dataset sizes, exclusion criteria, exact tests (e.g., paired t-test or bootstrap), multiple-comparison correction, or error bars; this information is load-bearing for the central claim that CPT+SFT gains over SFT are “frequently not statistically significant.”"},{"response":"We acknowledge the risk of systematic bias when relying solely on LLM-as-a-Judge for medical content. The original text already notes the imperfect correlation with human judgment. In revision we will add a human validation study on a representative subset of the OEQA outputs, reporting inter-annotator agreement, to strengthen the preference claims for instruction tuning and CPT+SFT. If the added study confirms the LLM-judge rankings, we will state this explicitly; otherwise we will qualify the claims accordingly.","revision_made":"yes","referee_comment":"[Evaluation / OEQA results] Evaluation protocol (OEQA subsection): reliance on LLM-as-a-Judge without a reported human validation subset or inter-annotator agreement on a medical sample risks systematic bias, especially since the paper notes imperfect correlation with human judgment; this directly affects the claim that instruction tuning and CPT+SFT are preferred."}],"tokens_in":1412,"tokens_out":437,"duration_ms":26799,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that for French medical multiple-choice QA, supervised fine-tuning gets most of the gains and adding continual pretraining on top rarely moves the needle enough to matter, while open-ended QA shows CPT helping overlap metrics but SFT sometimes hurting generation quality under the LLM judge.\n\nThe paper runs the same set of adaptations across three model families, multiple sizes, and different starting points. That disentangling is the useful part because it separates strategy effects from base model choice. They cover both MCQA and OEQA, test greedy and constrained decoding, report automatic metrics plus LLM-as-a-Judge, and include a cross-lingual check to English benchmarks. The hedging on small or insignificant gains for MCQA is straightforward and avoids overclaiming.\n\nWhat it does well is stay empirical and comparative at a decent scale. The preference ordering under the judge for OEQA and the note that SFT is cost-effective are the kind of practical observations that can save people compute when adapting to a new language and domain.\n\nThe soft spots are standard for this type of work. LLM-as-a-Judge is known to be imperfect for medical content, the tasks are limited to QA, and everything is French-specific so the guidelines may not travel far. The abstract flags statistical tests but leaves dataset sizes, exact methods, and error bars unclear, which makes it harder to assess robustness without the full text. No load-bearing flaws show up in the reported claims, just the usual limits on how far the results generalize.\n\nThis is for groups doing medical LLM work in non-English settings who need quick rules of thumb on adaptation order. A reader focused on compute-efficient domain adaptation would get some usable signals. It deserves a serious referee because the comparisons are broad enough to check and the tempered conclusions make the paper worth the time even if the core pattern is not surprising.","headline":"SFT alone is often enough for French medical MCQA while CPT adds limited value, with more mixed results for open-ended QA in this systematic but narrow empirical study.","tokens_in":2313,"tokens_out":451,"would_cite":false,"duration_ms":20000,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"For French medical multiple-choice QA, supervised fine-tuning alone matches or nearly matches the gains from adding continual pretraining, while open-ended QA shows different trade-offs.","keywords":["medical LLM adaptation","French QA","continual pretraining","supervised fine-tuning","MCQA","OEQA","domain adaptation","LLM evaluation"],"falsifier":"A human expert evaluation of the same model outputs on a held-out clinical French QA set that contradicts the ranking produced by the LLM judge or by the overlap metrics.","tokens_in":2632,"feed_emoji":"🩺","tokens_out":738,"duration_ms":10920,"temperature":0.7,"pith_summary":"The paper compares three adaptation strategies—continual pretraining, supervised fine-tuning, and their combination—on French medical question-answering tasks using multiple model families and sizes. It evaluates performance separately on multiple-choice questions and open-ended questions with both automatic metrics and LLM-based judging. The central finding is that for multiple-choice questions the combined approach wins most often but the extra benefit over fine-tuning alone is small and often statistically insignificant, while for open-ended questions continual pretraining helps overlap metrics yet fine-tuning tends to lower generation quality. Cross-lingual tests show that French-adapted models also improve on English medical benchmarks. These results are meant to guide practical choices when compute is limited.","feed_headline":"SFT nearly matches CPT+SFT for French medical MCQA","feed_subtitle":"Empirical comparison finds small, often insignificant gains from adding continual pretraining on multiple-choice tasks, with clearer differe","key_machinery":"Head-to-head comparison of continual pretraining (CPT), supervised fine-tuning (SFT), and CPT+SFT on French medical MCQA and OEQA datasets, disentangling effects from base model choice and initialization.","core_discovery":"Across three model families, CPT+SFT most often records the highest scores on multiple-choice medical QA, yet the margin over SFT alone is small and frequently fails to reach statistical significance; therefore SFT constitutes a strong, lower-cost default. On open-ended QA, CPT reliably raises overlap-based automatic metrics, whereas SFT tends to reduce generation quality; instruction tuning and the combined CPT+SFT route receive higher preference under LLM-as-a-Judge evaluation. Adaptation performed on French data transfers positively to English medical benchmarks.","pith_inferences":["The same trade-off pattern may appear in other low-resource languages when medical corpora are similarly scarce.","If clinical deployment values factual correctness over surface overlap, the preference for instruction-tuned or combined routes on OEQA would matter more than the automatic scores suggest.","Future work could test whether the small CPT+SFT margin persists when the base models are already instruction-tuned on general medical English data."],"forward_implications":["Under compute constraints, practitioners can default to SFT for multiple-choice medical QA without expecting large further gains from CPT.","For open-ended generation, CPT should be retained if overlap metrics matter, while SFT may need to be skipped or paired with instruction tuning.","French-domain adaptation produces measurable positive transfer to English medical QA benchmarks.","Model size and family interact with adaptation strategy, so results should be checked per architecture rather than assumed universal."],"fun_headline_variants":["SFT rivals CPT+SFT on French medical MCQA","Tiny CPT gains over SFT in medical multiple-choice QA","CPT raises OEQA metrics but SFT harms generation quality","French medical adaptation transfers to English benchmarks"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The selected French medical QA datasets together with the automatic metrics and LLM-as-a-Judge protocol give unbiased, representative measures of adaptation quality and clinical utility.","fun_headline_variants_meta":{"raw":{"variants":["SFT rivals CPT+SFT on French medical MCQA","Tiny CPT gains over SFT in medical multiple-choice QA","CPT raises OEQA metrics but SFT harms generation quality","French medical adaptation transfers to English benchmarks"]},"model":"grok-4.3","cost_usd":0.002605,"raw_usage":{"total_tokens":1491,"prompt_tokens":683,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":26049500,"prompt_tokens_details":{"text_tokens":683,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":747,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":683,"tokens_out":61,"duration_ms":6740,"temperature":1.0,"reasoning_tokens":747,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T20:58:31.742830+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A human expert evaluation of the same model outputs on a held-out clinical French QA set that contradicts the ranking produced by the LLM judge or by the overlap metrics.","supporting_citations":[],"review_version":1}