{"id":"05f9e94c-e057-47c9-bf38-971e2be38c2c","arxiv_id":"2607.21570","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MedGame converts static clinical case reports into structured interactive storytelling games and shows that fine-tuned open-source LLMs approach commercial performance on a new 5,000-case benchmark.","lead":"MedGame is a two-engine LLM framework that turns static medical case reports into interactive, decision-based storytelling games for students. The paper ships a 5,000-case benchmark and finds that fine-tuning open-source models on commercial-model outputs improves structural and content quality, while an eight-student pilot reports higher perceived engagement than text-only cases.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MedGame Bench content-quality gains (Tables 2–3) rely on GPT-5.2 judge with human r=0.61 on Story Direction and r=0.48 on API Type Selection (n=50); without larger human validation, the fine-tuning/gap-narrowing claim may be a judge artifact.","rationale":"The reader and I identify the same soft spot. The framework's structural claims—JSON/schema/business-logic validation, DAG executability—are largely independent of the LLM judge and are credible. The leakage analysis (Appendix E.4) is a genuine strength, and the release of the platform supports reproducibility. The central quantitative contribution, however, is the benchmark comparison showing fine-tuning improves content quality and narrows the gap. Those comparisons are made with a judge whose human correlation is moderate for Story Direction and, at the indicator level, low. Since the fine-tuned models are trained on Gemini-3-Pro references, a judge with stylistic biases could produce exactly the observed pattern without true quality gains. The expert-revision data reinforce that raw outputs are imperfect, so an unreliable judge cannot be treated as a minor issue. The existing CONDITIONAL verdict is appropriate: the work is promising, but the headline quantitative claims should be re-verified with larger human scoring before being taken at face value. Therefore I do not change the verdict.","tokens_in":36396,"tokens_out":5118,"duration_ms":52954,"concrete_test":"Select 200 Story Direction outputs stratified by model (e.g., 50 each from Gemini-3-Pro, Claude-Sonnet-4.5, Qwen3.5-27B*, and Qwen3-32B*) and by score decile. Have three experienced game developers and two clinicians, blinded to model identity, score RA/ATS/PC with the same rubrics (Appendix D.3). Compute per-indicator Pearson r vs GPT-5.2 and, crucially, compare human mean scores across models to GPT-5.2 model ranking. If fine-tuned models no longer significantly outperform their prompted baselines in human scores, the 'substantially improves and narrows the gap' conclusion for Story Direction fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's core quantitative claim—task-specific fine-tuning substantially improves open-source LLMs and narrows the gap with commercial models on MedGame Bench—is supported for structure by rule-based checks, but the content-oriented indicators (CCI, CSU, NQ, CDA, ODA, MEA, QDQ, FQ, RA, ATS, PC) are all scored by GPT-5.2 as an LLM judge (Section 5, Appendix D.2–D.3). Human validation (Section 6.4, Figure 18) uses only 50 cases per task: average r=0.81 for Medical Narrative Generation but r=0.61 for Story Direction, with the lowest indicator r=0.48 for API Type Selection and r=0.67 for Parameter Content. With n=50, the 95% CI for r=0.61 spans roughly 0.39–0.77; for r=0.48 it spans 0.23–0.67, so the 'moderate' agreement is too imprecise to validate 1–10 scores used for model ranking. Because fine-tuned models were trained on Gemini-3-Pro reference trajectories, their outputs may match the conventions of the reference corpus; a judge that rewards those conventions could show gains that human experts would not confirm. The medical-accuracy and educational-quality claims in Table 2 and the task-reasonability claims in Table 3 therefore rest on a measurement whose validity is weakest exactly in the range used to establish the headline result. Appendix C.2 acknowledges the reference trajectories lack independent medical verification, and Appendix E.3 shows experts mark 8.3–8.9 revision regions per draft, confirming raw LLM outputs contain substantial medical and pedagogical flaws that an imperfect judge may miss.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces MedGame, a dual-engine framework that turns static clinical case summaries into structured, executable storytelling games. A Medical Narrative Designer generates a hierarchical clinical storyline (Acts, Scenes, Decision Nodes), and a Story Director converts that storyline into a directed acyclic graph of multimodal generation tasks. The authors release MedGame Bench, built from 5,000 PMC-Patients cases, with rule-based structural validation and GPT-5.2 LLM-as-a-judge scoring for content-oriented indicators. Experiments compare commercial and open-source LLMs, and LoRA fine-tuned open-source models trained on Gemini-3-Pro reference trajectories. The reported results show that fine-tuning substantially improves open-source models on structural validity and on several content metrics, narrowing the gap with commercial models. A small learner-perception study and expert-revision analyses provide additional evidence. The central quantitative claim is that task-specific fine-tuning improves open-source LLMs on both MedGame Bench tasks.","tokens_in":36838,"tokens_out":7049,"duration_ms":76949,"significance":"The framework and benchmark address a real gap in LLM-based medical education: moving from fragmented QA interactions to whole-case, decision-centered learning trajectories. Concrete strengths include the released interactive platform, machine-checked Pydantic schemas, rule-based validation pipelines, a balanced 5,000-case benchmark, and extensive appendices. The structural validity results are convincing because they rely on automated checks rather than subjective scoring. The related-case overlap analysis is a responsible check on a common benchmark artifact. However, the content-oriented claims that carry the headline conclusion — fine-tuning improves medical accuracy and educational quality — depend on an LLM judge validated on only 50 cases per task, with per-indicator correlations as low as r=0.48. The reference trajectories produced by Gemini-3-Pro without independent medical verification further complicate interpretation. If the judge validity is strengthened or the claims are scoped to structural and stylistic improvements, the contribution is solid; as it stands, the central quantitative claim is conditional on a measurement whose validity is weakest in exactly the indica","major_comments":[{"comment":"The LLM-as-a-judge validation is based on only 50 cases per task. For Story Direction, the average human correlation is r=0.61, and the API Type Selection indicator — which is central to the fine-tuning gains in Table 3 (e.g., Qwen3.5-27B* ATS 9.02 vs. 7.01) — has r=0.48. With n=50, the 95% CI for r=0.48 spans roughly 0.23–0.67, so the judge's ranking of models is nearly indistinguishable from noise for that indicator. The headline claim that fine-tuning 'substantially improves' content quality and narrows the gap with commercial models rests on these scores. Please either provide a much larger human validation (per indicator, with inter-rater reliability) or explicitly downgrade the content-level conclusions to exploratory.","section":"§6.4 / Fig. 18 / Tables 2–3"},{"comment":"Fine-tuned models are trained on Gemini-3-Pro reference trajectories, and content quality is subsequently scored by GPT-5.2 without an independent medical gold standard. Appendix C.2 states that these reference trajectories 'are not treated as direct evidence of clinical correctness.' Because the judge may reward the formatting and narrative conventions of the reference model, the observed content gains (e.g., CSU 5.41→8.50 for Qwen3-32B*) may reflect distillation of Gemini-3-Pro style rather than improved clinical or pedagogical quality. Structural metrics are rule-based and robust to this concern, but CDA/ODA/MEA/QDQ/FQ are not. To support the interpretation, the authors should compare fine-tuned outputs against base outputs with human experts on content dimensions, or use independent expert-constructed references.","section":"§6.3 / Appendix C.2"},{"comment":"The Story Direction track evaluates only image-generation orchestration: the three API types in Appendix D.3 are character_gen, fusion, and modification, and all reported task-reasonability indicators (RA/ATS/PC) concern image tasks. The framework description in §4.2 and the toolset in Appendix A.3, however, include audio and video generation (f_aud, f_vid). No benchmark metric or experiment covers audio/video task planning or dependency modeling for those modalities. Thus the claim of validating 'dependency-aware multimodal orchestration plans' is overstated; MedGame Bench's Story Direction track should be described as visual orchestration only, or the benchmark should be extended to cover the other modalities.","section":"§5 / Appendix A.3 / Appendix D.3"}],"minor_comments":[{"comment":"Eq. (1) defines a linear storyline, which is an explicit design choice. Appendix B's branching extension is presented only as an inference-time construction and is not evaluated. Please state in the main text that the benchmark and experiments cover linear storylines only.","section":"§3 / Appendix B"},{"comment":"The same three medical Ph.D. evaluators who validated the LLM judge also performed the expert-revision study. Please clarify whether the revision evaluators were blinded to model identity and whether they had access to their own earlier scores for the same cases.","section":"§6.4 / Appendix D.4"},{"comment":"Reference formatting contains errors: 'LujieZheng LujieZheng' is duplicated, several entries end with '1 others', and one author name appears as 'Co¸ skun'. Please clean up the bibliography.","section":"References"},{"comment":"The learner-perception study uses eight students and a one-sided paired Wilcoxon test. This is acceptable for a pilot, but the conclusion should more explicitly emphasize the exploratory nature, as the Limitations section already does for long-term outcomes.","section":"§6.5 / Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a strong systems and benchmark contribution, and the structural results are well supported by automated checks. The main obstacle to acceptance is the gap between the strength of the headline content-quality claims and the validity evidence for the judge that produces them. The n=50 human validation and the low per-indicator correlations, especially for API Type Selection, make the fine-tuning 'gap narrowing' claim difficult to trust as stated. I would recommend requiring either a substantially larger human validation study or a careful scoping of the claims to structural and format-level improvements before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"MedGame is a genuine systems contribution: it turns static clinical cases into structured, executable storytelling games, and it ships a 5,000-case benchmark that will be useful to anyone working on LLM-based medical education. The dual-engine split — narrative design vs. story direction — is sensible, and the Pydantic-constrained schema plus the dependency DAG for multimodal orchestration are well thought out. The structural validity gains from fine-tuning are credible because they come from rule-based checks, not from the LLM judge.\n\nThe soft spot is the content-quality evaluation. The headline claim that fine-tuning narrows the gap with commercial models on medical accuracy and educational quality rests on GPT-5.2-as-a-judge, and the human validation is thin: 50 cases per task, with r=0.61 for Story Direction and r=0.48 for API Type Selection. Those correlations are too imprecise to support 1–10 scores used for model ranking. And because the fine-tuning data were produced by Gemini-3-Pro, the judge — a different commercial model — may be rewarding conformity to those reference conventions rather than clinically better content. The paper acknowledges this point in Appendix C.2, and the expert revision study shows raw outputs have 8+ revision regions per draft. So the structural story is solid, but I'd read the content-quality rankings as provisional until there's larger human validation.\n\nThe learner study is a perception pilot with eight students; the authors don't overclaim learning gains, and it's fine as a first signal.\n\nWho should read this: people building interactive medical education systems, and anyone doing benchmark-and-fine-tune evaluations for structured generation in a high-stakes domain. It deserves a serious referee. The benchmark and framework are worth publishing; the eval needs to be hardened or the content claims softened.\n\nRecommendation: send it out. Peer review should ask for a bigger human-judge set and an explicit check for judge bias toward reference-style outputs.","headline":"Solid new benchmark and system for LLM-driven clinical case gamification; content-quality claims need stronger human validation before the fine-tuning results are taken at face value.","tokens_in":37277,"tokens_out":2112,"would_cite":true,"duration_ms":22514,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MedGame transforms static clinical case reports into executable storytelling games, and task-specific fine-tuning lets open-source LLMs approach commercial quality on the two core generation tasks.","keywords":["medical education","storytelling games","large language models","narrative generation","multimodal orchestration","case-based learning","fine-tuning","gamification"],"falsifier":"A blinded review in which independent medical experts score a random sample of 200 generated storylines and orchestration plans using the paper's own rubrics; if their scores fail to correlate with the LLM judge's scores (roughly r < 0.5 for story direction or r < 0.7 for narrative generation), the benchmark's quality claims collapse. A randomized controlled trial comparing knowledge retention in students who play the multimodal game versus those who read the static case could settle whether the engagement advantage translates into learning gains.","tokens_in":36314,"feed_emoji":"🎮","tokens_out":7154,"duration_ms":61626,"temperature":0.7,"pith_summary":"This paper is trying to establish that large language models can convert static clinical case reports into structured, executable storytelling games for medical education, and that the conversion is best done as two separate generation tasks. It introduces MedGame, a dual-engine framework: a Medical Narrative Designer builds a case-grounded storyline with acts, scenes, and decision nodes, and a Story Director turns that storyline into a dependency-aware multimodal execution plan rendered by an interactive platform. The paper reports that fine-tuning open-source models on a few thousand reference trajectories substantially improves their output quality and narrows the gap with commercial models, and that medical students perceive the multimodal game as more engaging and useful than text-only cases. If correct, this offers a scalable, low-cost way to turn existing case libraries into active reasoning practice.","feed_headline":"Open-source LLMs can turn clinical cases into playable games","feed_subtitle":"Fine-tuning on reference trajectories narrows the gap with commercial models; students rate the games more engaging than text.","key_machinery":"The load-bearing object is the Clinical Storyline, a hierarchical, machine-readable structure built from Acts (macro-stages), Scenes (clinical steps), and Decision Nodes (learner-facing checkpoints). The key move is to make this structure the interface between two LLM-driven engines: the Medical Narrative Designer, which generates it from a patient summary under fidelity, pedagogical, and structural constraints, and the Story Director, which consumes it and emits a directed acyclic graph of multimodal generation tasks with dependency placeholders. This separation decouples clinical narrative ideation from technical orchestration, and the graph's dependency propagation is what preserves chara","core_discovery":"The paper's central claim is that case-to-game transformation can be factorized into two tractable generation tasks rather than one free-form request to 'write a game.' The Medical Narrative Designer produces a hierarchical Clinical Storyline — Acts, Scenes, and Decision Nodes — that preserves the factual trajectory of the source case while adding learner-facing choices and feedback. The Story Director then maps that storyline into a directed acyclic graph of multimodal generation primitives (image, audio, video) with explicit identity and scene dependencies, so the same patient and location remain consistent across the rendered game. The paper provides evidence that task-specific fine-tunin","pith_inferences":["Inference: The same dual-engine factorization — narrative design separate from technical orchestration — could transfer to other high-stakes training domains, such as law, aviation, or emergency response, where cases exist as static text but reasoning must be practiced sequentially.","Inference: Because the paper's story-direction quality scores rely on an LLM judge with only moderate human agreement (r=0.61), orchestration-quality claims should be treated as approximate until larger human validation is run.","Inference: A testable extension is to randomize learners across linear versus branching versions of the same case to see whether the branching extension the paper outlines in an appendix improves reasoning outcomes.","Inference: The logical next step, which the authors explicitly leave open, is a larger longitudinal study measuring knowledge retention and clinical reasoning transfer rather than perceived engagement."],"forward_implications":["Existing static case libraries can be repurposed into interactive, decision-centered training without human actors, at a fraction of the cost of standardized patient programs.","A few thousand reference trajectories are enough to make open-source LLMs structurally reliable at both narrative generation and story direction, so institutions can run the pipeline on local hardware.","Medical accuracy remains the hard part: even the strongest models score around 7/10 on clinical accuracy metrics, so expert review is still required before generated content reaches students.","The Story Director's dependency-aware graph makes identity-preserving, causally coherent multimodal rendering feasible, which the paper links to higher perceived engagement and presence in learners.","The benchmark and evaluation protocol give the community a reusable way to measure structured educational storylines, beyond single-turn medical question answering."],"fun_headline_variants":["Two-step LLM pipeline turns medical cases into story games","Fine-tuned LLMs craft clinical story games students prefer","Case to game: LLMs with structured design beat free-form","Medical education: LLM-crafted games boost engagement","Storytelling gamification: LLMs split case-to-game into two tasks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that LLM-as-a-judge scores on 1,000 test cases genuinely measure clinical accuracy and educational quality; human agreement is strong for narrative generation but only moderate for story direction, and the reference trajectories used for fine-tuning were themselves produced by a commercial model without independent medical verification.","fun_headline_variants_meta":{"raw":{"variants":["Two-step LLM pipeline turns medical cases into story games","Fine-tuned LLMs craft clinical story games students prefer","Case to game: LLMs with structured design beat free-form","Medical education: LLM-crafted games boost engagement","Storytelling gamification: LLMs split case-to-game into two tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000636,"raw_usage":{"total_tokens":2732,"prompt_tokens":670,"completion_tokens":2062,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":1992}},"tokens_in":414,"tokens_out":2062,"duration_ms":13469,"temperature":1.0,"reasoning_tokens":1992,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:01:36.916995+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A blinded review in which independent medical experts score a random sample of 200 generated storylines and orchestration plans using the paper's own rubrics; if their scores fail to correlate with the LLM judge's scores (roughly r < 0.5 for story direction or r < 0.7 for narrative generation), the benchmark's quality claims collapse. A randomized controlled trial comparing knowledge retention in students who play the multimodal game versus those who read the static case could settle whether the engagement advantage translates into learning gains.","supporting_citations":[],"review_version":1}