{"id":"6ba9ce62-2af8-4cc8-ba1c-0c15f5877889","arxiv_id":"2411.15593","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A visual analytics system that aligns multimodal medical data to help novice physicians screen, analyze, and review retrospective teaching cases.","lead":"Medillustrator is a visual analytics system that helps novice physicians review past patient cases by aligning MRI images, diagnostic text, and lab indicators into one view. A user study with 13 physicians suggests the system helps them find and analyze valuable teaching cases faster than the hospital's existing system.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The evaluation measures one-session screening and rationale quality, not retrospective learning; without a retention or transfer test, the central claim that Medillustrator improves learning is not established.","rationale":"The reader's conditional verdict is appropriate. I agree with the reader's weakest assumption at a general level: the outcome variable does not establish learning. I would emphasize the conceptual mismatch rather than only the ground-truth definition. The study is internally coherent as a system-usability and screening-effectiveness evaluation: randomized assignment, Mann-Whitney comparisons, ICC for report scoring, and a case study with a clinical collaborator. But the title and abstract claim learning improvement, and none of the dependent variables measure learning. Adding a delayed retention or transfer test would settle this directly. I also note the extra tutorial and exploration time in the treatment arm as a confound; however, the more decisive issue is construct validity. Because the paper is otherwise a reasonable system paper with transparent reporting of modest effects and acknowledged limitations, REJECT is too strong; keeping the verdict CONDITIONAL is the right call, conditional on reframing the claim or adding a learning-outcome measure.","tokens_in":22996,"tokens_out":6167,"duration_ms":59495,"concrete_test":"Run a pre-registered follow-up in which participants return after a one-week washout and repeat a parallel task on a fresh pool of 50 unseen cervical-spine case folders: first, select the 10 most valuable learning cases; second, write one-sentence rationales. Have senior physician raters not involved in benchmark selection, blind to condition, score selection match and rationale accuracy. Compare Medillustrator versus baseline on delayed and transfer performance. If the delayed advantage is not significant, the current study cannot support the \"improves retrospective learning\" claim; if it is significant, the central claim is substantially strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Medillustrator \"enhances physicians' retrospective learning\" (Abstract; Section 7). The controlled study in Section 5.2 measures, within a single one-hour session, how many of 10 senior-curated benchmark cases participants select and the subjective precision of their written rationales. That is a plausible measure of immediate screening and analysis performance, not of learning: there is no pre-test, no delayed retention test, and no transfer task on unseen cases. Even a real between-group difference on these scores would not show that knowledge or skill was acquired, consolidated, or carried forward. The issue is compounded by the benchmark construction: the 10 \"valuable\" cases are selected by two senior physicians with no stated operational criteria or inter-rater reliability, and reports are scored by senior physicians using a rubric that rewards plausible rationales; if the same judges define both the target set and the scoring, the outcome is partly circular. A second internal-validity issue is that the Medillustrator arm received extra system tutorial and exploration time before the task, so the comparison is not purely system versus baseline. Section 6 does not acknowledge this gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Medillustrator, a visual analytics system intended to support novice physicians' retrospective learning from multimodal diagnostic data (MRI images, diagnostic text, and laboratory indicators). The system includes a data processing pipeline (image annotation, Grounding DINO fine-tuning, SAM segmentation), a modeling engine (ResNet-50 and BERT embeddings fused and projected with UMAP, k-NN neighborhood and Jaccard-based glyphs), and a five-view interface organized around discovery-learning levels (overview, detail, retrospection). The authors report a formative study with six physicians, a case study with one intern, and a controlled in-lab between-subject study with 13 participants comparing Medillustrator to a hospital HIS baseline. The user study reports significantly higher scores for the Medillustrator group on report completeness (U=3, p<0.01) and accuracy (U=3.5, p<0.05), as well as on several usability and effectiveness Likert items.","tokens_in":23126,"tokens_out":4288,"duration_ms":38262,"significance":"If the central claim were fully established, the system would address a real and underserved need in continuous medical education: helping novices identify high-value cases and interpret multimodal data in a unified view. The paper's strengths include a transparent description of the modeling pipeline, a realistic baseline comparison (the hospital HIS), a formative study that grounds design requirements, and a case study that illustrates the intended workflow in detail. The statistical reporting is transparent in the sense that U and p values are given. However, the evaluation as designed measures immediate screening and analysis performance, not learning outcomes. There is no retention test and no transfer task, and the ground-truth benchmark selection is not validated with inter-rater reliability. These issues are load-bearing because the abstract and conclusion frame the result as demonstrating enhanced retrospective learning. The system itself is plausible and the qualitative insights are useful, but the evidence does not yet support the broadest claims.","major_comments":[{"comment":"The study's dependent variables are the number of benchmark cases selected during a one-hour session and the precision of the written rationales. These measure immediate screening and analysis performance, not the acquisition, retention, or transfer of knowledge. The Abstract's claim that Medillustrator 'enhance[s] physicians' retrospective learning processes' and the Conclusion's statement that 'user evaluations demonstrate Medillustrator's effectiveness in aiding novice physicians to efficiently identify and analyze research cases' are not supported by the same evidence: the former is a learning claim, the latter is a performance claim. A delayed retention test or a transfer task on unseen cases would be required to support the learning claim; alternatively, the claims should be restricted to immediate performance.","section":"§5.2 (Procedure) and §7 (Conclusion)"},{"comment":"The 10 benchmark cases are described as 'curated by two senior physicians' with no operational definition of what makes a case valuable and no inter-rater reliability reported for this selection. Because the completeness score is defined as the number of selected cases matching this ground truth, the validity of the completeness measure depends entirely on the reliability and validity of that judgment. The ICC value reported (above 0.75) is for the two senior raters who scored the reports, not for the benchmark selection itself. The paper should report the explicit selection criteria (e.g., diagnostic discrepancy, atypical presentation, educational value) and an inter-rater statistic such as Cohen's kappa for the benchmark selection. If the same two senior physicians both selected the benchmarks and scored the reports, the report scoring is also at risk of criterion contamination.","section":"§5.2 (Procedure)"},{"comment":"Participants in the Medillustrator condition received a 5-minute system tutorial followed by 10 minutes of free exploration before the one-hour task, while the baseline condition did not receive an equivalent orientation period. The measured effect may therefore be attributable to additional time on task or differential attention rather than to the system's design. The authors should either give the baseline group an equivalent familiarization activity (e.g., a structured orientation to the HIS features relevant to the task) or otherwise control for time and attention across conditions, and should discuss this confound in Section 6, which currently does not acknowledge it.","section":"§5.2 (Procedure)"},{"comment":"The median completeness scores are 5.71 (Medillustrator) versus 5.57 (baseline) on a 10-point scale. Under the scoring rubric in Appendix B, both values fall in the same band (5-6: 'Identifies 70-84% of cases'), corresponding to the same number of benchmark cases (7 of 10). The reported p<0.01 with U=3 (n1=7, n2=6) indicates a significant rank difference, but the absolute difference of 0.14 points is small and its practical significance is unclear. The paper should report effect sizes (e.g., rank-biserial correlation) and discuss whether the observed difference is educationally meaningful.","section":"Table 2 and Appendix B"}],"minor_comments":[{"comment":"The citation 'Duanm et al.' should read 'Duanmu et al.' (reference [19]).","section":"§2.2"},{"comment":"The model names 'resnet50' and 'bert-base-uncased' should be typeset consistently as 'ResNet-50' and 'BERT-base-uncased'.","section":"§4.2.2"},{"comment":"The three-set similarity formula JA,B,C = |A∩B∩C| / |A∪B∪C| is not the standard Jaccard index for three sets; either define the intended normalized intersection measure explicitly or use the standard formulation.","section":"§4.3.2"},{"comment":"The compensation amount '20 compensation' lacks a currency unit; please specify the currency and amount.","section":"§5.2 (Participants)"},{"comment":"In the description of step 13, 'c2 also be examined as abnormal' should read 'p2 was also examined as abnormal'.","section":"§5.1 Case Study Part II"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central problem is the gap between what is measured (one-session screening and report quality) and what is claimed (enhanced retrospective learning). This is fixable by either adding a retention/transfer measure or substantially tempering the abstract and conclusion. The small sample size (n=13) is a further concern, and the small absolute differences in Table 2 suggest that even the significant results may have limited practical magnitude. The system itself is well-motivated and the paper has useful qualitative content, so major revision rather than rejection is appropriate if the authors are willing to align their claims with the evidence or collect additional data. The paper is within scope for a HCI venue, though the clinical-education claims should be framed more cautiously."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my honest read. Medillustrator is a genuine systems contribution with a thoughtful design process, but the paper's central claim about learning is not supported by the evidence. The formative study is solid, and the integration of Grounding DINO + SAM with the embedding glyphs is new as far as I know. The case study gives a good feel for the intended workflow, and the design alternative discussion is useful transparency. The user study, however, measures screening and report-writing performance in a single session, not retrospective learning. There is no pre-test, no delayed retention test, and no transfer to unseen cases. The abstract's claim that the system 'enhances physicians' retrospective learning processes' overreaches what the data can show. The benchmark construction is also under-specified: the 10 valuable cases were curated by two senior physicians with no stated criteria and no inter-rater reliability for the case labels; the reported ICC only covers the report scoring, not the ground truth. On top of that, the Medillustrator arm received extra tutorial and exploration time, which is a real confound. The alignment model's mAP 0.53 is reported without error analysis, which matters if the system is supposed to support interpretation. Effect sizes are modest — medians of 5.71 vs 5.57 on completeness, 4.16 vs 3.5 on accuracy — and with n=13 the significant p-values are not robust. These are fixable problems. The paper would be much stronger if the authors reframed the claim to 'supports case screening and analysis' or added a retention/transfer measure. As is, it's a decent systems paper overstating its evaluation. I'd send it to peer review, but I'd require the authors to either add a learning measure or tone down the claim. Anyone working on medical visualization or CME will find the system design worth reading; the evaluation is also a good teaching example of why outcome measures need to match the construct.","headline":"Solid system paper whose main claim about learning is untested; worth reviewing, but the evaluation needs a reframing or a retention measure.","tokens_in":23767,"tokens_out":2828,"would_cite":false,"duration_ms":27675,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Medillustrator is a visual analytics system that aligns MRI images, diagnostic text, and lab indicators so novice physicians can screen, analyze, and revisit high-value teaching cases more effectively than with their standard hospital…","keywords":["Retrospective Learning of Physicians","Visual Analytics","Multimodal Data Alignment","Continuous Medical Education","Medical Imaging","Case-Based Learning","Novice Physicians","User Study"],"falsifier":"Check the rating scales with a different panel: have an independent set of senior physicians label the same 50-case set, measure inter-rater reliability, and rerun the task; if the Medillustrator advantage disappears against a differently labeled gold standard, the result is an artifact of the chosen ten. Alternatively, give both groups a delayed retention test or follow-up supervised case discussion weeks later; if the Medillustrator group shows no better recall or diagnostic performance, the \"learning\" improvement is task-specific, not retrospective learning.","tokens_in":22715,"feed_emoji":"🩻","tokens_out":6977,"duration_ms":63181,"temperature":0.7,"pith_summary":"The paper claims that retrospective learning—physicians reviewing past patient cases to improve their skills—can be made substantially easier for novices by presenting multimodal diagnostic data as one aligned, navigable whole. Medillustrator combines an overview of patient mentions and embedding projections, a detail view where diagnostic text is overlaid on the corresponding areas of MRI images, reference ranges for lab indicators, and a record view for saving insights. In a controlled study with 13 novice physicians, the group using Medillustrator selected more of the ten benchmark cases a senior panel had judged valuable and wrote more accurate rationales than the group using the hospital information system, and they also rated the system higher on usability. If the result holds, hospitals and continuing medical education programs could move from senior-physician intuition toward a structured, tool-assisted workflow for turning routine patient records into learning material.","feed_headline":"Aligned MRI and text views help novices find high-value cases","feed_subtitle":"Thirteen novices tested it; report completeness and accuracy beat their usual hospital system.","key_machinery":"The carrying object is the Medillustrator system itself, specifically its multimodal alignment and representation pipeline. The pipeline has four load-bearing parts: a text-to-image grounding step that connects diagnostic phrases to bounding boxes and then to pixel-level segmentation in MRI images; an embedding step that projects image, text, and indicator features into a shared low-dimensional space where each patient is drawn as a glyph whose three segments encode the overlap similarity of nearest-neighbor sets across modality pairs; a detail view with a practice phase (raw image, user annotations, imaging indicators on parallel axes) and a learning phase (aligned diagnostic text layered onto the image by a force-directed layout); and a record view that saves analyzed cases as cards for later review.","core_discovery":"On the paper's own terms, the discovery is that a visual analytics approach built around semantic alignment of multimodal data supports novice physicians' retrospective learning better than the conventional record browser. The system's design follows a discovery-learning structure: start with an overview to find high-value cases, analyze a case in detail with practice-then-learning phases, then record and revisit conclusions. Its evaluation shows the Medillustrator group outperforming the baseline on both measured report dimensions—completeness (median 5.71 vs 5.57 on a 10-point scale) and accuracy (4.16 vs 3.5)—as well as on perceived ease of use, screening help, and confidence, with significance on a non-parametric rank test.","pith_inferences":["The study's completeness scores are close (medians 5.71 vs 5.57), so the larger, more robust benefit may be screening efficiency and user confidence rather than raw identification; a longer-term test with retention measures would clarify whether the learning gain is durable.","A natural testable extension is to replace the senior-physician gold standard with outcome-based labels—for example, cases that later involved diagnostic revision or adverse events—and see whether the system helps novices discover those cases too.","The text-to-image grounding component could be reused as a clinical explanation aid beyond training, for example to highlight what a diagnostic sentence refers to in a scan during case conferences or decision support.","Tracking physicians' record-view entries over many sessions would turn the tool into a portfolio that measures learning curves, something a single one-hour experiment cannot show."],"forward_implications":["If the measured advantage is real, novice physicians can use the embedding glyphs to detect cases where image and indicator signals disagree with the text diagnosis—precisely the cases most likely to be misdiagnosed and most valuable to study.","The practice-then-learning split lets trainees commit to their own reading of an image before the system reveals the physician-aligned diagnosis, turning passive review into an active retrieval exercise.","Because the alignment pipeline runs on physician annotations, the same workflow could be rebuilt for other image-heavy specialties by repeating the annotation and fine-tuning process.","Recorded case cards give trainees a persistent artifact for retrospection, which the formative interviews identified as a bottleneck in current continuing medical education practice."],"supporting_citations":[{"why":"Supplies the discovery-learning theory that structures the system's three-level overview, detail, and retrospection workflow.","marker":"[10]"},{"why":"Motivates the design requirement that reference values be shown alongside indicators for interpretation and learning.","marker":"[13]"},{"why":"Supplies the open-set object detector that the authors fine-tune to link diagnostic text phrases to image regions.","marker":"[29]"},{"why":"Supplies the segmentation step that turns detector boxes into pixel-level masks for aligned image representation.","marker":"[25]"},{"why":"Provides the dimensionality-reduction projection used to place image, text, and indicator features into the shared case view.","marker":"[34]"},{"why":"Supplies the overview-first interaction mantra that the interface follows.","marker":"[50]"},{"why":"Gives the rank-based test used for all group comparisons in the user study.","marker":"[32]"},{"why":"Gives the intraclass-correlation method used to confirm consistency between the two senior raters of the analysis reports.","marker":"[51]"}],"fun_headline_variants":["Visual system aligns scans and notes to improve doctor training","Multimodal alignment helps novices analyze cases effectively","Medillustrator improves novice report completeness and accuracy","Visual analytics for retrospective learning in physician CME","Aligned MRI and text views aid novice physicians' case review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The effectiveness claim rests on treating the ten cases chosen by two senior physicians as the ground truth for what is worth learning, and treating how many of those cases a novice selects (plus the precision of the written rationale) as the measure of learning, without measuring the reliability of that ground truth or showing that the task predicts real clinical improvement.","fun_headline_variants_meta":{"raw":{"variants":["Visual system aligns scans and notes to improve doctor training","Multimodal alignment helps novices analyze cases effectively","Medillustrator improves novice report completeness and accuracy","Visual analytics for retrospective learning in physician CME","Aligned MRI and text views aid novice physicians' case review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001078,"raw_usage":{"total_tokens":4464,"prompt_tokens":853,"completion_tokens":3611,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":3549}},"tokens_in":469,"tokens_out":3611,"duration_ms":23425,"temperature":1.0,"reasoning_tokens":3549,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:06:39.436078+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the rating scales with a different panel: have an independent set of senior physicians label the same 50-case set, measure inter-rater reliability, and rerun the task; if the Medillustrator advantage disappears against a differently labeled gold standard, the result is an artifact of the chosen ten. Alternatively, give both groups a delayed retention test or follow-up supervised case discussion weeks later; if the Medillustrator group shows no better recall or diagnostic performance, the \"learning\" improvement is task-specific, not retrospective learning.","supporting_citations":[],"review_version":1}