{"id":"559c0715-8d0d-4119-82f0-8dc83626cd60","arxiv_id":"2607.06127","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"General-domain OS-sLLMs moderately correlate with human OPTION12 SDM scores on Dutch melanoma transcripts; medical models fail via hallucination, and a Judge-LLM consensus is proposed.","lead":"Open-source smaller LLMs can partially score shared decision-making in Dutch melanoma consultations under OPTION12, with Gemma3:12b reaching moderate correlation (r=0.51) to humans while medical models hallucinate. The work shows privacy-preserving local models as a foundation for human-in-the-loop clinical coding rather than full automation.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The reported moderate correlations and Gemma ranking rest on only 7 cleaned development consultations after exclusions, rendering model selection and the foundation claim statistically fragile.","rationale":"The reader correctly isolates the single most load-bearing vulnerability: the entire quantitative scaffolding (model ranking, correlation magnitudes, Judge-LLM design) is derived from a post-exclusion N=7 pilot. That sample is too small and potentially biased for the claims that rest on it, while the qualitative error taxonomy and medical-model failure demonstrations remain solid. Because the paper already frames itself as a development-phase pilot and the test-set/Judge-LLM results are future work, the CONDITIONAL verdict with MODERATE confidence is appropriate; no stronger rejection is warranted, nor is an upgrade. The concrete bootstrap-plus-reprocessing check would settle whether the numbers survive even modest expansion of the development set.","tokens_in":14302,"tokens_out":569,"duration_ms":26741,"concrete_test":"Re-process the 3 timing-format and 2 multi-CG files with corrected extraction, recompute Pearson/Spearman on the full original development set of 11 (or the maximal clean subset), and obtain 95 % bootstrap CIs by resampling the 7–11 files; if Gemma’s r drops below 0.3 or the CI includes 0, the ranking and “moderate agreement” claim used for Judge-LLM selection no longer hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Gemma3:12b Pearson r=0.51 / Spearman ρ=0.59 as strongest agreement, supporting OS-sLLMs as a privacy-preserving foundation) is computed exclusively on the 7 successfully processed development consultations remaining after 4 of 11 were excluded for preprocessing failures or SDM unsuitability (Table 1, Section 4.3). Exact-match rates are already low (Gemma total consensus 20/84 item-file pairs ≈ 24 %, mean 2.86 items per file). With N=7 the item-level vectors and any file-level aggregation have high sampling variance; no confidence intervals, significance tests, or leave-one-out stability checks are reported. If the excluded files (different timing formats, multi-caregiver labels) systematically differ in discourse structure or SDM density, both the model ranking used to designate the Judge-sLLM and the claim of “moderate” alignment become non-generalizable to the held-out 15 and to real deployment. The medical-model hallucination examples are robust, but the positive general-domain numbers that underwrite the paper’s main recommendation are not.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces LLM4SDM, the first evaluation of open-source smaller LLMs (OS-sLLMs) for automated Observer OPTION12 coding of shared decision-making on Dutch melanoma consultation transcripts. It compares three general-domain models (Gemma3:12b, Llama3.1:8b, Mistral7b) and two medical-domain models (MedLlama2:7b, Meditron7b) on a development set of expert double-coded consultations, reports that general models outperform medical ones (which hallucinate and fail instruction following), finds Gemma3:12b strongest (Pearson r=0.51, Spearman ρ=0.59), provides item-level agreement tables and a qualitative error taxonomy (temporality, role attribution, evidence grounding), and proposes a Judge-LLM consensus framework to reconcile multi-model outputs for a planned human-in-the-loop pipeline.","tokens_in":14581,"tokens_out":1157,"duration_ms":24066,"significance":"If the pilot findings hold under larger evaluation, the work supplies a privacy-preserving, locally deployable alternative to commercial LLMs previously used only for the coarser OPTION5 instrument. The documented medical-model failure modes, the item-difficulty analysis, and the explicit error taxonomy are immediately useful to the clinical-NLP and SDM-measurement communities. The Judge-LLM design is a concrete, falsifiable proposal for multi-model consensus that mirrors human double-coding practice. These contributions are novel for OPTION12 and for open smaller models; the sustainability/privacy framing is a genuine advance over prior commercial-model studies.","major_comments":[{"comment":"Section 4.3 and Table 1: after excluding 4 of 11 development files for preprocessing failures or SDM unsuitability, all quantitative claims (Tables 2–4, including Gemma’s r=0.51 / ρ=0.59) rest on N=7 consultations (84 item-file pairs). Exact-match rates are already low (Gemma 20/84 ≈ 24 %). No confidence intervals, bootstrap, leave-one-out stability, or significance tests are supplied. With such a small cleaned sample the model ranking used to designate the Judge-sLLM and the claim of “moderate” alignment are statistically fragile and may not generalise to the held-out 15 or to deployment; this is load-bearing for the paper’s central recommendation.","section":"Section 4.3, Tables 1–4"},{"comment":"The methodology (Figure 1, Section 3) and conclusions (Section 6) describe a full testing-phase pipeline that deploys the selected Judge-sLLM on the 15 held-out consultations and evaluates consensus against human gold labels. No such results appear. Presenting the development-only pilot numbers in the abstract and as “findings” while leaving the planned evaluation unexecuted overstates the current evidence for the Judge-LLM framework and for the “promising foundation” claim.","section":"Section 3, Figure 1, Section 6, Abstract"},{"comment":"The “Accepted Abstract” block reports substantially higher correlations (0.83/0.80/0.64) and a different model ranking (Mistral best by consensus count) than the body (Tables 2 and 4: Gemma best, r≈0.51). These contradictory numbers cannot both be correct; the discrepancy must be resolved or the outdated abstract block removed, otherwise readers cannot trust the reported performance figures.","section":"Accepted Abstract vs. Tables 2 & 4"}],"minor_comments":[{"comment":"Human inter-rater reliability (before consensus) is never quantified, yet the paper repeatedly compares LLM–human agreement to an implicit human standard. Even a simple percentage agreement or weighted kappa on the original double codes would strengthen interpretation of the LLM numbers.","section":"Section 4.1 / future-work paragraph"},{"comment":"Table 6 (speaker statistics) is presented without any statistical test or clear link to the OPTION12 scores; either analyse the relationship formally or move the table to an appendix.","section":"Appendix Table 6"},{"comment":"Several figures (2–4) illustrate medical-model hallucinations well, but the captions and surrounding text could more explicitly state that these are representative failure modes rather than isolated runs.","section":"Figures 2–4"},{"comment":"Minor typographical inconsistencies appear (e.g., “OPION12”, “Medtron” vs “Meditron”, “sLLM” vs “OS-sLLM”). A careful proof-read would remove them.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The work is clearly a pilot write-up of an ongoing project (ISDM 2026 abstract extension). The core idea and error analysis are publishable, but the current manuscript over-claims relative to the executed experiments. I would accept a revised version that either (a) reports the held-out test-set + Judge-LLM results or (b) re-frames the entire paper strictly as a development-set pilot with explicit statistical caveats and removes the unfinished pipeline claims from the abstract and conclusions. Scope is appropriate for a clinical-NLP or digital-health venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a methods pilot, not a finished evaluation. What is new is the first look at open-source smaller models on the full OPTION12 instrument, using Dutch melanoma transcripts and a privacy-first local-deployment framing. Prior work used commercial giants on OPTION5 English breast-cancer data; this paper correctly flags that gap and fills it with a concrete comparison.\n\nWhat they do well: the medical-model failure is documented with clear, reproducible hallucination examples (repeated evidence, score-justification mismatches, invented mean scores). The qualitative error taxonomy is the most useful part of the paper—temporality failures, role confusion, evidence hallucination, score-evidence mismatch. That taxonomy will travel. The Judge-LLM consensus idea is a sensible adaptation of how human coders actually resolve disagreements, even if it is only proposed here.\n\nThe soft spot is real and load-bearing for the positive claim. After excluding 4 of 11 development files for preprocessing or unsuitability, the reported Gemma correlations (r=0.51, ρ=0.59) and the model ranking that designates the Judge rest on N=7 consultations. Exact-match rates are already low (~18–24 %). No CIs, no stability checks, no human inter-rater numbers yet. The medical-model failures do not depend on that N; the “moderate agreement / foundation for human-in-the-loop” claim does. The authors are transparent that the test set and Judge evaluation are future work, so the paper should be read as a scoped pilot, not as a completed validation.\n\nMath and citation pattern look clean enough for a pilot; no circular scoring. Data and code are not released, which is a practical limit for anyone who wants to build on it.\n\nThis is for medical-NLP and SDM-measurement people who care about local models and clinical dialogue. It deserves a serious referee once the small-N and missing-test-set limitations are stated up front. I would engage with the taxonomy and the privacy framing; I would not yet treat the correlation numbers as settled.","headline":"Useful first pilot of local OS-sLLMs on OPTION12 Dutch melanoma data, but the headline correlations rest on only seven cleaned consultations and the test-set/Judge-LLM results are still missing.","tokens_in":15206,"tokens_out":524,"would_cite":true,"duration_ms":5906,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Open-source smaller language models reach moderate agreement with human OPTION12 shared-decision scores and can support privacy-preserving coding with humans in the loop, while medical-domain models fail.","keywords":["shared decision making","OPTION12","open-source LLMs","privacy-preserving AI","clinical consultation coding","human-in-the-loop","melanoma transcripts","discourse analysis"],"falsifier":"Apply the same prompts, models, and Judge-LLM procedure to the held-out fifteen-consultation test set and check whether correlations with human consensus collapse toward zero, or whether medical-domain models suddenly outperform the general-domain ones that led on the development set.","tokens_in":15190,"feed_emoji":"🩺","tokens_out":1020,"duration_ms":25317,"temperature":0.7,"pith_summary":"This paper asks whether open-source smaller language models, run locally, can score how well clinicians involve patients in treatment decisions using the twelve-item OPTION12 observer scale on Dutch melanoma consultation transcripts. It finds that general-domain models reach moderate correlation with expert human coders, while the tested medical-domain models hallucinate and fail to follow instructions. Gemma3:12b is strongest, with Pearson correlation 0.51 and Spearman 0.59 against human consensus. Because models still miss many items and struggle with timing of decision framing, speaker roles, and grounding evidence in the transcript, they cannot replace human annotators. They can, however, draft scores and justifications that humans check, and a Judge-LLM can reconcile disagreements among models the way human coders discuss until they agree. A sympathetic reader cares because manual OPTION12 coding is slow and often inconsistent, and commercial cloud models raise privacy and cost barriers for clinical data.","feed_headline":"Small open LLMs score shared decisions at r=0.51","feed_subtitle":"General models beat medical ones; local drafts can cut coding load while keeping patient data private","key_machinery":"The Judge-LLM consensus framework: several OS-sLLMs independently produce OPTION12 scores plus evidence and justification for each item; a selected best-performing model then acts as judge to reconcile those candidates into one consensus label, mirroring independent human coding followed by discussion to agreement.","core_discovery":"On expert-annotated Dutch melanoma consultations, general-domain open-source smaller language models achieve moderate agreement with human OPTION12 shared-decision scores, with Gemma3:12b strongest at Pearson r=0.51 and Spearman ρ=0.59, while medical-domain models fail through hallucination and instruction-following failures. Exact item matches remain low, and errors cluster around temporal discourse reasoning, role attribution, and evidence grounding. Current OS-sLLMs therefore cannot replace human annotators but form a usable foundation for privacy-preserving human-in-the-loop assessment, aided by a Judge-LLM that resolves multi-model disagreements into consensus scores.","pith_inferences":["The consistent superiority of general over medical models implies that OPTION12 coding is mainly a discourse and role-reasoning problem, not a medical-knowledge problem.","If the Judge-LLM procedure holds on the test set, the same multi-model-plus-judge pattern could transfer to other observer instruments used in clinical communication research.","Caregiver label noise and timing-format failures in preprocessing may have understated true model capability; cleaner speaker-normalized transcripts could lift exact-match rates without changing the models.","Privacy-preserving local scoring may enable multi-site SDM quality benchmarking that commercial cloud models cannot legally support under strict health-data rules."],"forward_implications":["Locally run OS-sLLMs can draft OPTION12 scores and evidence so human coders spend time verifying rather than coding from scratch, while patient transcripts never leave institutional systems.","Item difficulty is systematic: temporal decision framing and clinician understanding-checks are hardest, so multi-model ensembles can cover complementary strengths across the twelve items.","Few-shot examples and later fine-tuning or parameter-efficient adaptation are expected to raise agreement on the remaining hard items toward usable human-in-the-loop levels.","Medical-domain smaller models need further work on instruction following and hallucination control before they can be trusted for discourse-level clinical coding tasks.","A Judge-LLM step can stand in for a third human when independent model scores disagree, reducing the cost of consensus meetings."],"fun_headline_variants":["Gemma3:12b hits r=0.51 on open LLM OPTION12 SDM scoring","General OS-sLLMs beat medical models for private SDM coding","Local small LLMs reach ρ=0.59 on Dutch melanoma shared decisions","OS-sLLMs draft OPTION12 scores but need human-in-loop fixes","Judge-LLM consensus aids multi-model SDM agreement resolution"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The reported correlations and model ranking rest on only seven successfully processed development consultations after four of eleven were dropped for formatting or suitability problems, so those results may not represent the full set of consultations.","fun_headline_variants_meta":{"raw":{"variants":["Gemma3:12b hits r=0.51 on open LLM OPTION12 SDM scoring","General OS-sLLMs beat medical models for private SDM coding","Local small LLMs reach ρ=0.59 on Dutch melanoma shared decisions","OS-sLLMs draft OPTION12 scores but need human-in-loop fixes","Judge-LLM consensus aids multi-model SDM agreement resolution"]},"model":"grok-4.5","effort":"low","cost_usd":0.006462,"raw_usage":{"total_tokens":1633,"prompt_tokens":834,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":64620000,"prompt_tokens_details":{"text_tokens":834,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":715,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":834,"tokens_out":84,"duration_ms":7500,"temperature":1.0,"reasoning_tokens":715,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T01:16:50.586126+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Apply the same prompts, models, and Judge-LLM procedure to the held-out fifteen-consultation test set and check whether correlations with human consensus collapse toward zero, or whether medical-domain models suddenly outperform the general-domain ones that led on the development set.","supporting_citations":[],"review_version":2}