{"id":"ab4d3d4f-3626-4fe8-98e1-b81acdf9c23d","arxiv_id":"2504.13023","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"ChatEXAONEPath fine-tunes LLaVA with WSI-level features from TCGA slides, reaching 62.9% acceptance by an AI judge on generated pathology reports, though the judge's reliability is questioned in the paper.","lead":"The authors build ChatEXAONEPath, a whole-slide image language model trained on TCGA slides and LLM-generated pathology captions. They claim a 62.9% acceptance rate from an AI evaluator, but the evaluation and data generation are both LLM-driven, with no human validation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 62.9% headline metric rests on an LLM evaluator that the paper itself concedes may be non-trustworthy; without human validation or evaluator reproducibility, the central diagnostic claim is unsupported.","rationale":"The reader's verdict of REJECT rests on the claim that the AI-evaluator-based acceptance rate is non-trustworthy. My independent reading of the manuscript confirms this as the single most load-bearing concern. The paper's own Evaluation section contains an explicit, self-admitted limitation: the evaluator LLM makes inconsistent decisions and can produce incorrect interpretations, and 'in the worst case, the acceptance rate may be non-trustworthy.' Because the headline result is nothing more than that acceptance rate, the central claim collapses unless the evaluator's decisions are independently validated. The paper provides no human evaluation, no baseline comparison, no confidence intervals, and no reproducibility check for the evaluator. I also note an internal inconsistency between the described chain-of-thought evaluation and the appendix prompt that demands only 'accept' or 'reject' as output, which further undermines confidence that the protocol was executed as claimed. These issues are threats to correctness, not merely disagreements with prevailing consensus, because the metric itself is the evidence offered for the paper's strongest claim. The concrete test of having pathologists adjudicate a sample of cases, and measuring evaluator stability across runs, would settle whether the concern lands. Since the reader already recommended REJECT and my analysis reinforces that recommendation, no verdict change is needed.","tokens_in":12846,"tokens_out":2315,"duration_ms":21820,"concrete_test":"Select a random stratified sample of 200 of the 1,134 test cases, including the model's chosen best answer and the LLM evaluator's accept/reject decision. Have at least two board-certified pathologists independently judge whether the generated answer correctly identifies the primary diagnosis given the WSI and reference report, and compare their judgments to the LLM evaluator's decisions using Cohen's kappa. If the pathologist-evaluator agreement is below 0.6, or if the pathologists' acceptance rate differs from 62.9% by more than 10 percentage points, the headline metric is not a valid measure of diagnostic quality. Also rerun the LLM evaluator 10 times with nonzero sampling temperature on the same 200 cases and report the accept/reject flip rate, since this directly tests the paper's own admission that evaluator decisions are unstable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the 62.9% acceptance rate on 1,134 WSI/test pairs. The metric is defined by an LLM evaluator's accept/reject decision on the best of 10 generated answers. The paper explicitly concedes in the Evaluation section that the evaluator's reasons are 'sometimes inaccurate,' that decisions are 'somewhat unstable,' and that 'in the worst case, the acceptance rate may be non-trustworthy—leaving the whole AI-based evaluation system incorrect.' This is a self-admitted failure of the measurement instrument, not a peripheral caveat. No human-pathologist agreement study, no baseline model, and no evaluator variance estimate are provided, so the 62.9% cannot be interpreted as diagnostic accuracy or 'expert-level' performance. Additionally, the appendix evaluation prompt (Table A2) instructs the evaluator to output only the words 'accept' or 'reject,' which contradicts the paper's claim that chain-of-thought reasoning was elicited during evaluation; either the CoT protocol was not implemented as described, or the published prompt is not the one used. Both issues bear directly on the validity of the headline number.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ChatEXAONEPath, a whole-slide-image (WSI) multimodal large language model for histopathology. The model uses an EXAONEPath patch encoder, a CLAM-based patch aggregator, and a LLaVA-style architecture with LLaMA2:7B as the language backbone. The authors introduce RAIDER, a retrieval-augmented pipeline that generates instruction-tuning data from 10,094 TCGA WSI-report pairs, scaling to 69,544 instruction pairs. They also propose an AI-based evaluation protocol in which LLaMA3.1:70b selects the best of 10 generated answers and then accepts or rejects it. The main reported result is an acceptance rate of 62.9% (713/1,134) for the best model variant, on the basis of which the abstract claims expert-level diagnostic ability.","tokens_in":13069,"tokens_out":4150,"duration_ms":37034,"significance":"If the evaluation were reliable, the paper would offer a useful practical recipe for building WSI-level pathology MLLMs from public TCGA data and for generating instruction datasets without manual annotation. The RAIDER pipeline and the comparison of three model variants are potentially informative. However, the central quantitative claim rests entirely on an LLM evaluator that the paper itself concedes is unstable and possibly non-trustworthy. No human expert validation, no baseline model comparison, no error bars or repeated-evaluation variance, and no patient-level split analysis are provided. As a result, the headline acceptance rate cannot support the paper's 'expert-level' diagnostic claim.","major_comments":[{"comment":"The acceptance rate is the only quantitative evidence for the model's diagnostic ability, but the paper itself states that the evaluator's 'responses and interpretations are somewhat unstable, resulting in incorrect decision making' and that 'in the worst case, the acceptance rate may be non-trustworthy - leaving the whole AI-based evaluation system incorrect.' No human-pathologist agreement study, no repeated-evaluation variance estimate, and no calibration against expert ground truth is reported. The 62.9% figure therefore cannot be interpreted as diagnostic accuracy or as supporting 'expert-level' performance.","section":"Evaluation, Table 3"},{"comment":"The evaluation prompt instructs the evaluator to 'provide only the words accept or reject as your response,' which directly contradicts the Evaluation section's claim that chain-of-thought prompting was used to elicit reasoning steps. Either the published prompt is not the one actually used, making the protocol unreproducible, or CoT reasoning was not actually elicited. Either way, the description of the evaluation protocol is internally inconsistent and this inconsistency bears directly on the validity of the headline metric.","section":"Appendix Table A2"},{"comment":"The manuscript does not clarify whether the 1,134 test pairs were excluded from the RAIDER generation for dataset-v2. If the 69,544 training instruction pairs for v2/v3 were generated from all 10,094 reports, including the test reports, then the reported acceptance rates for v2/v3 are inflated by train/test overlap. This must be stated explicitly, and the evaluation should be re-run on a properly held-out split.","section":"Experiments, Dataset and Table 3"},{"comment":"No baseline model is evaluated under the same protocol, and no human pathologist comparison is provided. The paper compares only its own v1/v2/v3 variants. Without a baseline (e.g., an existing pathology MLLM or an untrained LLaVA) and without human expert agreement, the claim that the model is 'expert-level' and can 'comprehensively understand' WSIs is unsupported even setting aside the evaluator-reliability concern.","section":"Experiments, Evaluation"}],"minor_comments":[{"comment":"The word 'reprsentative' should be 'representative.'","section":"Dataset-v1 section"},{"comment":"The question 'What is the major diagnosis?' appears twice (question 1 and question 7); one instance is presumably redundant.","section":"Table A1"},{"comment":"The phrase 'innately making the whole evaluation process a multi-task performer' is unclear and should be reworded.","section":"Evaluation section"},{"comment":"The statement that the acceptance rate was calculated 'based on the test datasets from dataset-v1' is ambiguous for models trained on dataset-v2; clarify whether the same 1,134 test pairs were used for all three variants and whether those pairs were excluded from dataset-v2 generation.","section":"Experiments, Evaluation section"}],"recommendation":"reject","confidential_remarks":"The central result is undermined by the authors' own admission that the evaluation system may be non-trustworthy, and the appendix prompt contradicts the described CoT protocol. The paper would need a substantially new evaluation framework, including human validation, baseline comparisons, and evaluator reliability analysis, before the diagnostic claim could be considered. The potential train/test overlap for dataset-v2 is an additional concern that should be examined if a future submission is made."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here is the short version: this paper has a genuinely useful data-generation pipeline and a candid limitations section, but its headline number—62.9% acceptance—rests on an evaluator the authors themselves say may be untrustworthy. The appendix prompt for the evaluator asks only for the words \"accept\" or \"reject,\" which contradicts the paper's claim that chain-of-thought reasoning was elicited. That is a real problem, not a quibble.\n\nWhat is new: the RAIDER pipeline, which uses retrieval over OCR'd TCGA reports to build instruction-tuning pairs for WSI-level MLLMs. That is a practical idea, and the paper makes a genuine attempt to move from patch-level to whole-slide understanding. The architecture—CLAM aggregator plus LLaVA-style projector—is not original, but the combination is reasonable. The Table 1 comparison of prior WSI models is useful. Credit also for openly reporting that the augmented dataset (v2) did not beat v1; that kind of negative result is worth knowing.\n\nNow the soft spots. The evaluation section says the evaluator is \"sometimes inaccurate,\" makes \"incorrect or manipulated reasons,\" and \"in the worst case, the acceptance rate may be non-trustworthy.\" That is the authors' own text. There is no human-pathologist agreement, no baseline model, no error bars, and no patient-level split description. The appendix prompt tells the evaluator to respond only with accept or reject, while the main text claims CoT was used. Either the prompt is not the one used, or the CoT claim is false. Both options undermine the metric. The title says \"expert-level,\" but nothing in the paper measures diagnostic accuracy against human-anchored ground truth. The 62.9% is an LLM's opinion filtered through an unstable judge.\n\nIn proportion: the method and pipeline are worth discussing, but the evaluation is load-bearing and broken. The paper does not support its central claim. I would still take it for peer review—there is enough substance in RAIDER and the honest negative result that a careful reviewer could push the authors toward a proper human study. But it is not close to acceptable in current form. I would not cite it yet. For a reading group, it is a good case study in LLM-as-judge pitfalls.","headline":"Useful data pipeline, but the headline acceptance rate rests on an evaluator the authors themselves admit may be untrustworthy—and the published evaluator prompt contradicts the claimed CoT reasoning.","tokens_in":13560,"tokens_out":2582,"would_cite":false,"duration_ms":22493,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The authors claim expert-level diagnosis from whole-slide pathology images, with 62.9 percent of answers accepted by an AI evaluator.","keywords":["ChatEXAONEPath","multimodal large language model","histopathology","whole slide images","pan-cancer","retrieval-augmented generation","AI-based evaluation","TCGA"],"falsifier":"Run the same 1,134 test pairs through a second, independently constructed LLM evaluator, or have a panel of pathologists score the answers, and compare acceptance rates; a large swing would show the 62.9 percent figure reflects the evaluator, not the model's diagnostic ability.","tokens_in":12651,"feed_emoji":"🔬","tokens_out":9284,"duration_ms":77056,"temperature":0.7,"pith_summary":"The paper introduces ChatEXAONEPath, a multimodal large language model that takes a whole-slide histopathology image and answers text questions about the diagnosis. The authors show that a two-phase training scheme, vision-language alignment followed by LoRA instruction tuning, can run with 10,094 WSI-report pairs and still produce answers that a separate LLM evaluator accepts in 62.9 percent of 1,134 test cases. Along the way they introduce RAIDER, a retrieval-augmented pipeline that expands original reports into more instruction-tuning examples, and an AI evaluation protocol that scores answers on seven criteria. The central claim is that a slide-level multimodal LLM trained on modest public data can understand pan-cancer morphology and clinical context well enough to generate diagnostic text, making automated assistance in pathology more practical.","feed_headline":"Whole-slide pathology AI diagnoses cancer, 62.9 percent accepted","feed_subtitle":"From 10,094 WSI-report pairs, a multimodal LLM generates pathologist-style answers on 1,134 test slides.","key_machinery":"The load-bearing components are a frozen patch encoder (EXAONEPath) that embeds 256x256 patches; a CLAM-based patch aggregator (CBPA) with gated attention that learns which patches matter and outputs a single WSI embedding; a vision projector with attention pooling and linear layers that maps the 512-dimensional WSI embedding to 4096-dimensional text space; LLaMA2-7B-Chat as the language generator, fine-tuned with LoRA; and RAIDER, a retrieval-augmented instruction-data generator that OCRs TCGA reports, stores chunks in a vector database, retrieves relevant chunks by cosine similarity, and uses Llama3.1:70B-instruct to write short diagnostic answers. The evaluation mechanism is an LLM-based evaluator that uses chain-of-thought reasoning and seven criteria to accept or reject the best of ten generated answers.","core_discovery":"ChatEXAONEPath is a LLaVA-style model whose vision tower combines a frozen EXAONEPath patch encoder with a CLAM-based attention aggregator to turn thousands of patches into one slide-level embedding; a vision projector aligns that embedding with LLaMA2-7B-Chat, and phase-2 training uses LoRA. The authors report that the best configuration (v3) reaches a 62.87 percent acceptance rate on 1,134 WSI-report pairs, with the primary measurement being an instruction-tuned Llama3.1:70B evaluator that first selects the best of 10 generated answers and then accepts or rejects it using seven criteria: accuracy, relevance, completeness, clarity, appropriateness, consistency, and presentation. They also report that scaling the dataset by text-only augmentation (v2, 69,544 pairs) did not beat the smaller v1 dataset, and that the AI evaluator sometimes produces incorrect rationales, so the acceptance rate may not be trustworthy. The claim is that a modestly sized WSI-level multimodal LLM can diagnose across cancer types and that an interpretable AI evaluation protocol can stand in for human judgment.","pith_inferences":["If the evaluation were repeated with pathologist raters, the 62.9 percent acceptance rate would likely change; the number should be read as accepted by this LLM evaluator, not as a clinically validated diagnostic accuracy.","The success of the frozen patch encoder and attention aggregator suggests the slide embedding produced by the vision tower could be reused for other slide-level tasks, such as survival prediction or biomarker subtyping, which the paper does not test.","The RAIDER result implies that simply paraphrasing existing reports does not create new visual supervision; future gains would likely require new WSI-report pairs or enrichment with genomic or clinical data.","A multimodal evaluator that also sees the WSI, rather than only text answers and references, might make more consistent decisions than the text-only LLM evaluator used here; the paper names this as future work."],"forward_implications":["A WSI-level pathology multimodal LLM can be trained from about ten thousand WSI-report pairs rather than millions, and still answer diagnostic questions in a way that an AI evaluator accepts most of the time.","Text-only augmentation of the same underlying pairs does not reliably improve the model; the v2 dataset with 69,544 pairs performed worse than the v1 dataset with 10,094 pairs, suggesting that alignment between slides and text matters more than quantity.","An LLM evaluator that explains its accept/reject decisions through chain-of-thought can yield interpretable evaluations, but those rationales are sometimes inaccurate, so the evaluation protocol itself needs scrutiny before deployment.","Because the model is pan-cancer, the same system can be applied across multiple cancer types seen in TCGA."],"supporting_citations":[{"why":"EXAONEPath supplies the frozen patch-level encoder that embeds 256x256 patches from each WSI.","marker":"Yun et al. 2024"},{"why":"CLAM provides the attention-based multiple-instance learning aggregator that compresses patches into a single slide embedding.","marker":"Lu et al. 2021"},{"why":"Visual instruction tuning defines the LLaVA-style architecture that concatenates image and text embeddings for the LLM.","marker":"Liu et al. 2024"},{"why":"LLaMA2-7B-Chat is the language backbone that generates answers.","marker":"Touvron et al. 2023"},{"why":"LoRA is the parameter-efficient fine-tuning method used in phase 2.","marker":"Hu et al. 2021"},{"why":"Retrieval-augmented generation is the basis for the RAIDER instruction-data pipeline.","marker":"Lewis et al. 2020"},{"why":"Llama3.1:70B-instruct generates the augmented dataset and serves as the AI evaluator.","marker":"Dubey et al. 2024"},{"why":"PRISM is the prior WSI-level captioning model whose prompts and slide-level approach this work extends.","marker":"Shaikovski et al. 2024"},{"why":"PathChat provides the overall MLLM architecture and prompt format that ChatEXAONEPath is built around.","marker":"Lu et al. 2024b"}],"fun_headline_variants":["WSI-level multimodal LLM achieves 62.9% acceptance","ChatEXAONEPath: multimodal LLM reads WSIs, 62.9% accepted","Whole-slide pathology AI: 62.9% acceptance rate","Pan-cancer WSI LLM hits 62.9% acceptance","Pathology LLM: 62.9% acceptance from 1,134 pairs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the AI evaluator's accept and reject decisions reflect genuine diagnostic quality; the paper itself warns that the evaluator can make inconsistent decisions, so the 62.9 percent acceptance rate may not measure real diagnostic accuracy.","fun_headline_variants_meta":{"raw":{"variants":["WSI-level multimodal LLM achieves 62.9% acceptance","ChatEXAONEPath: multimodal LLM reads WSIs, 62.9% accepted","Whole-slide pathology AI: 62.9% acceptance rate","Pan-cancer WSI LLM hits 62.9% acceptance","Pathology LLM: 62.9% acceptance from 1,134 pairs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001678,"raw_usage":{"total_tokens":6722,"prompt_tokens":1085,"completion_tokens":5637,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":701,"completion_tokens_details":{"reasoning_tokens":5535}},"tokens_in":701,"tokens_out":5637,"duration_ms":37681,"temperature":1.0,"reasoning_tokens":5535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:15:30.361397+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 1,134 test pairs through a second, independently constructed LLM evaluator, or have a panel of pathologists score the answers, and compare acceptance rates; a large swing would show the 62.9 percent figure reflects the evaluator, not the model's diagnostic ability.","supporting_citations":[],"review_version":1}