{"id":"c801a08f-82cf-45b4-9a6b-a3c9d3a2a2e4","arxiv_id":"2607.19201","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"MIRA-Ev is a Spanish/English/Basque benchmark that annotates clinical exam cases with span-level evidence, claims, and support/attack relations so models can be scored on reasoning quality, not just answer accuracy.","lead":"MIRA-Ev is a new benchmark that asks AI systems to find the exact evidence sentences, spans, and support/attack links behind clinical exam answers, not just pick the right option. It is released in Spanish, English, and Basque—the first clinical argumentation dataset in Basque—and is designed to score reasoning quality rather than final-answer accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-lingual equivalence of MIRA-Ev is asserted but unvalidated: Section 3.3 gives no translation/projection protocol and no per-language IAA, so the Basque/English benchmark and 'first Basque clinical argumentation resource' claims are unsupported.","rationale":"The reader's CONDITIONAL verdict appropriately identifies the unvalidated gold standard and the missing cross-lingual protocol as the weakest assumption. Our stress-test agrees that this is load-bearing: the benchmark's central value proposition depends on reliable annotations across Spanish, English, and Basque, and the paper provides no evidence of translation/projection methodology or annotation consistency. The concern is addressable — the authors can supply the dataset, annotation guidelines, and IAA numbers — so the conditionality is correct rather than grounds for rejection. We nevertheless sharpen the requirement: the cross-lingual equivalence must be demonstrated by an explicit protocol and empirically validated, or the 'first clinical argumentation resource in Basque' claim cannot be accepted. The proposed back-translation check would provide direct evidence of whether the released Basque/English versions preserve the argumentative structure, and would settle whether the central multilingual claim actually holds.","tokens_in":7970,"tokens_out":7084,"duration_ms":77387,"concrete_test":"Take a random sample of 50 cases from the released dataset (once accessible). Have a professional clinical translator independently translate the Spanish source into Basque and English, then compare the argumentative structure of the independent translations to the released versions: align gold span boundaries and relations across languages (e.g., via token alignment) and compute span F1 and relation agreement (e.g., Cohen's kappa). If alignment F1 falls below ~0.85 or relation agreement below an acceptable threshold (e.g., kappa < 0.6), the parallel-equivalence claim fails. An analytical prerequisite: verify the annotation guidelines document the cross-lingual construction method; if it is absent, the burden of proof is already unmet.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MIRA-Ev is a valid multilingual clinical argumentation benchmark, with novelty resting in part on being 'the first clinical argumentation resource in Basque.' Section 3.3 ('Multilingual Construction') consists of a single sentence asserting parallel English and Basque versions, with no description of how they were created: no human-translation protocol, no machine-translation plus post-editing, no back-translation step, no clinician validation, and no per-language inter-annotator agreement. No IAA is reported for the Spanish source either. Without this, the gold annotations' correctness and cross-lingual equivalence are unverified. If the Basque/English versions were projected from Spanish via automatic alignment, sub-sentential evidence spans and directed support/attack relations may not be preserved, making the multilingual evaluation claims hollow and the Basque novelty a labeling artifact rather than a clinical reasoning resource. This is load-bearing because the abstract and contributions (i) and (v) foreground multilingual parallelism and Basque; if the non-Spanish versions are not validated, the claimed first Basque clinical argumentation resource loses its evidentiary basis and the benchmark's cross-lingual comparability breaks.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces MIRA-Ev, a clinical argument-mining benchmark built from Spanish MIR exam cases, re-annotated by clinicians with span-level premises and claims plus directed support/attack relations. It proposes a three-tier task hierarchy (evidence sentence retrieval, component extraction, relation classification) and parallel Spanish, English, and Basque releases, claiming the first clinical argumentation resource in Basque. The experimental section reports a single EriBERTa encoder pipeline on the Spanish version across the three tiers, with strict and relaxed matching. The authors report Task 1 macro-F1 73.48, Task 2 macro-F1 62.63/68.02 strict/relaxed, Task 3 macro-F1 7.58/10.61, and an aggregate 'Pipeline Overall' of 47.90 repeated across all result tables. The discussion interprets the high claim-detection F1 as a positional-shortcut artifact and attributes low relation F1 to cascading extraction errors.","tokens_in":8258,"tokens_out":6132,"duration_ms":73798,"significance":"If properly validated, MIRA-Ev would address a real gap: MCQA benchmarks cannot verify whether a correct diagnosis is grounded in relevant evidence, and Basque has essentially no clinical argumentation resources. The task hierarchy and strict/relaxed matching protocol in §3.1–§4.1 are clearly defined, and the deliberate preservation of inverted claim/premise mappings is a useful probing device. The promise of a trilingual, token-level clinical argumentation benchmark with expert annotations is significant for multilingual clinical NLP. However, as submitted, the central resource claims are not supported by the reported validation: there is no inter-annotator agreement, no multilingual construction protocol, and no experimental evidence for the LLM and cross-lingual evaluations promised in the introduction. The novelty rests heavily on the Basque and multilingual claims, so the missing validation is load-bearing.","major_comments":[{"comment":"The abstract and contributions claim parallel Spanish/English/Basque versions and 'the first clinical argumentation resource in Basque,' but §3.3 consists of a single sentence with no translation/projection protocol. There is no description of human translation, machine translation plus post-editing, span alignment, or clinician validation per language; Table 1 reports only aggregate counts. Because the non-Spanish versions are the basis of the novelty claim, the manuscript must specify how the English and Basque versions were constructed, how sub-sentential spans and directed relations were preserved, and provide per-language validation.","section":"§3.3; contributions (i), (v); abstract"},{"comment":"The introduction and contribution (iv) promise benchmarking with 'a range of medical-pretrained and general-purpose generative LLMs' across tiers, languages, and the standard/inverted mapping. The Results section contains only one 'Encoder + Classifier' pipeline, with no LLM results, no per-language results, and no explicit inverted-mapping results. Either add the missing tables or revise the abstract and contributions to state that LLM and multilingual evaluation are future work. As written, the stated empirical scope is unsupported.","section":"§1; §5; Tables 2–4"},{"comment":"The same 'Pipeline Overall' value 47.90 appears in all six rows of Tables 2–4, including both strict and relaxed rows. The explanation that the score is 'shared across all three tables' because it reflects the end-to-end pipeline run does not define the metric; moreover, relaxed Task 2 (68.02) and Task 3 (10.61) differ from their strict counterparts, so an unchanged aggregate is unexpected. Specify the formula (e.g., macro-average of task F1s) and report strict and relaxed aggregates separately, or remove the duplicate column.","section":"§5, Tables 2–4"},{"comment":"No inter-annotator agreement is reported for the Spanish source, nor for the English and Basque versions. Since MIRA-Ev's value is as gold ground truth for span-level premises/claims and directed relations, the reliability of these labels is central. Please report IAA on at least a held-out subset (span-boundary agreement and relation-level agreement, separately by language), and describe the annotation guidelines, adjudication, and clinician background.","section":"§3.2–§3.3; §4.1"},{"comment":"The Discussion attributes the low Task 3 F1 (7.58 strict / 10.61 relaxed) to cascading upstream extraction errors, but no experiment isolates this claim. An oracle relation-classification run with gold spans would provide an upper bound; without it, the observed scores could equally reflect relation-model deficiency. Add an oracle/upper-bound experiment or soften the causal interpretation.","section":"§6, 'Error Propagation in Relation Classification'; Table 4"},{"comment":"The Discussion asserts that performance 'degraded sharply' on inverted-mapping cases where the claim is embedded in the narrative and options serve as premises, and that models relied on positional heuristics. No table or statistic reports the size of the inverted subset or the standard/inverted results. This is the central evidence for the argumentative-role ambiguity claim; please report subset counts and strict/relaxed F1 separately for standard and inverted mappings.","section":"§6, 'Claim Detection: Reliance on Structural Heuristics'"}],"minor_comments":[{"comment":"The source data description says prior annotations and clinician explanations were 'stripped' from CasiMedicos, but no details are given on licensing, redistribution rights, or data availability. Since the manuscript claims the resource is released, include a repository URL and license.","section":"§3.2"},{"comment":"The notation in Tables 2–4 is dense and sometimes unclear (e.g., 'Overall Component Detection' vs 'Pipeline Overall'). Consider adding a legend that explains which columns are macro-averaged and which are per-class.","section":"§5"},{"comment":"Task 1 is defined as binary sentence classification, but the implementation uses the [SEP] token representation. Clarify how the relevance label for the sentence is obtained from the token-classification head in practice.","section":"§4.2.1"},{"comment":"A few references are incomplete or informal (e.g., 'Clinical trials' under World Health Organization, 2021). Please verify all citations and expand them where needed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's abstract and contributions substantially overstate the experimental content: the results section covers a single Spanish encoder pipeline, while LLM, cross-lingual, and inverted-mapping results are promised but absent. The multilingual construction and annotation reliability are not documented, which is central to the Basque-novelty claim. I believe these issues are fixable within the scope of a major revision by adding the missing protocols, IAA, and either experiments or clearly scoped claims. I do not see a core technical error that would require rejection, but the current version should not be published without addressing the evidence gap."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core of this is worth taking seriously: MIRA-Ev gives clinical NLP a reasoning-trace evaluation that MCQA can't, and the token-level argumentative annotation of MIR cases—especially the deliberately inverted claim/premise mapping—is a real step beyond CasiMedicos-Arg and ClinArgES. The Basque version is new, and the task hierarchy in §3.1 is clearly specified. The authors also deserve credit for honest discussion in §6: they openly call the high claim-detection score a structural artifact and attribute Task 3 failures partly to cascading pipeline error. That is the right way to read your own results.\n\nThe soft spots are real but mostly fixable. The abstract promises LLM evaluation across languages and inverted mappings, but Section 5 delivers only one encoder run. No LLM tables, no Basque/English tables, no inverted-subset numbers, no dataset link, no inter-annotator agreement anywhere. The identical 47.90 across Tables 2–4 is actually explained in the text—it is the end-to-end pipeline aggregate, not a repeated task metric—so that specific charge is off. The bigger problem is §3.3: the parallel Basque and English versions are asserted in a single sentence with no translation or projection protocol. That is load-bearing, because the Basque novelty is a primary contribution. If the non-Spanish versions were machine-aligned without validation, the cross-lingual claims collapse. The stress-test note is right about that.\n\nI would not call the resource circular: the gold labels are human judgments independent of the evaluated model. But the paper's empirical section is currently too thin to verify the benchmark's usefulness for cross-lingual work, and the abstract overstates what is demonstrated.\n\nWho benefits: people building clinical argument mining resources, especially in low-resource languages, and anyone evaluating whether LLMs ground diagnoses in relevant evidence. The paper deserves a serious referee—it should be sent to review, not desk-rejected—but it needs revision before publication: report IAA, release or at least describe the multilingual construction protocol, and either add the LLM/inverted-mapping results or remove those claims from the abstract.","headline":"A genuinely useful clinical reasoning resource, but the paper's empirical and multilingual claims run ahead of the evidence actually shown.","tokens_in":8743,"tokens_out":1261,"would_cite":false,"duration_ms":16621,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MIRA-Ev: a clinical reasoning benchmark that checks whether a model's diagnosis is actually grounded in the right evidence, not just whether the final answer is correct.","keywords":["clinical argument mining","MIR exam","multilingual clinical NLP","evidence sentence retrieval","argumentative component extraction","support/attack relations","Basque clinical NLP","reasoning trace evaluation"],"falsifier":"A direct falsifier would be an inter-annotator agreement study on a random sample of the corpus: if expert clinicians disagree on more than a small fraction of span boundaries or relation polarities, or if translated versions fail to align with the native annotations on a per-span basis, the reliability of the benchmark as a gold standard collapses. Similarly, if the Basque version was produced by translation rather than native annotation, and a controlled study shows that models perform differently on translated versus native Basque text, the claim of a 'native' resource would be weakened.","tokens_in":7899,"feed_emoji":"🩺","tokens_out":1247,"duration_ms":16476,"temperature":0.7,"pith_summary":"MIRA-Ev is a new benchmark built from Spanish medical licensing exams that evaluates clinical reasoning by looking at the evidence a model uses, not just the answer it gives. Expert clinicians re-annotated exam cases with fine-grained argumentative structure: spans of text marked as premises or claims, and directed relations of support or attack between them. The dataset is released in parallel Spanish, English, and Basque versions, making it the first clinical argumentation resource in Basque. The paper argues that this resource can reveal failure modes that standard multiple-choice benchmarks miss, such as a model selecting the right diagnosis while relying on irrelevant, absent, or contradictory evidence.","feed_headline":"New benchmark tests whether clinical AI reasons from the right evidence","feed_subtitle":"MIRA-Ev scores evidence spans and support/attack relations in Spanish, English, and Basque exams, exposing wrong-reasoning-with-right-answer","key_machinery":"The central object is the argumentative structure annotated over clinical case vignettes: spans of text typed as Premise or Claim, connected by directed Support or Attack relations. The key mechanism is the deliberate inversion of the standard mapping between candidate options and claims in a subset of cases. Normally, options are claims and the case text provides premises; in inverted cases, an option functions as a premise and the claim is embedded in the case stem. This inversion prevents models from using the positional heuristic 'options are claims' and forces them to infer argumentative role from content, allowing the benchmark to directly probe whether a system understands argumentati","core_discovery":"The central claim is that clinical NLP evaluation should move beyond answer-only accuracy to inspect the reasoning trace, and that argument mining provides a workable formalism for doing so at scale. The paper introduces a three-tier task hierarchy—evidence sentence retrieval, argumentative component extraction, and relation classification—and shows empirically that current encoder-based systems perform well on retrieving relevant sentences (macroF1 ~73) but collapse on relation classification (strict macroF1 ~7.6). A key finding is that claim detection is inflated by a positional shortcut: because candidate options usually map to claims, models can 'look up' claims from options, and perform","pith_inferences":["The benchmark's emphasis on the reasoning trace could be extended to evaluate explanation generation tasks: instead of scoring a generated explanation as a whole, one could use MIRA-Ev's annotation schema to check whether each stated premise is grounded in the case and whether the claimed support/attack relations match gold evidence.","The inverted-mapping design suggests a general principle for building diagnostic benchmarks: deliberately break syntactic heuristics (position, punctuation) to force deeper semantic processing. The same trick could be applied to other NLP tasks where models exploit positional shortcuts, such as reading comprehension or fact verification.","The pattern of over-predicting sentence relevance (high recall, lower precision on the Not-Relevant class) implies that clinical evidence retrieval models might be biased toward treating all case content as relevant; a calibration-aware evaluation could reveal whether this is a dataset effect or a model bias.","If the relation classification bottleneck is partly due to cascading span errors, an end-to-end architecture trained jointly on extraction and relation classification might substantially outperform the sequential pipeline, a testable hypothesis the paper does not explore."],"forward_implications":["Benchmarks like MIRA-Ev can detect 'right answer, wrong reasoning'—a failure mode that multiple-choice QA cannot see—and thus offer a more informative signal for clinical model evaluation.","The inverted-mapping subset provides a concrete diagnostic for whether a model has learned positional shortcuts; a model that performs well on standard instances but poorly on inverted ones is not reasoning about argumentative roles.","The strict-versus-relaxed boundary findings show that span-boundary noise, not semantic misunderstanding, is the main bottleneck in component extraction, suggesting that evaluation metrics and model architectures should focus on boundary robustness.","The very low relation classification scores indicate that directed support/attack reasoning in clinical text is a largely unsolved problem, pointing to a clear research target for clinical NLP.","The Basque version of the dataset opens a new direction for low-resource clinical reasoning evaluation, allowing progress in a language with essentially no prior annotated reasoning data."],"fun_headline_variants":["Clinical AI retrieves evidence but can't reason with it","MIRA-Ev benchmark tests evidence retrieval and argument relations","First multilingual clinical argument-mining benchmark includes Basque","Models good at finding evidence, bad at judging support vs attack","AI cheats claims by using answer options, not clinical reasoning"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The gold annotations—both the span boundaries and the support/attack relations—must be correct and consistent across the Spanish, English, and Basque versions; the paper does not report inter-annotator agreement or an annotation/projection protocol for the parallel versions.","fun_headline_variants_meta":{"raw":{"variants":["Clinical AI retrieves evidence but can't reason with it","MIRA-Ev benchmark tests evidence retrieval and argument relations","First multilingual clinical argument-mining benchmark includes Basque","Models good at finding evidence, bad at judging support vs attack","AI cheats claims by using answer options, not clinical reasoning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000703,"raw_usage":{"total_tokens":2958,"prompt_tokens":643,"completion_tokens":2315,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":387,"completion_tokens_details":{"reasoning_tokens":2244}},"tokens_in":387,"tokens_out":2315,"duration_ms":19595,"temperature":1.0,"reasoning_tokens":2244,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T13:07:25.974626+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifier would be an inter-annotator agreement study on a random sample of the corpus: if expert clinicians disagree on more than a small fraction of span boundaries or relation polarities, or if translated versions fail to align with the native annotations on a per-span basis, the reliability of the benchmark as a gold standard collapses. Similarly, if the Basque version was produced by translation rather than native annotation, and a controlled study shows that models perform differently on translated versus native Basque text, the claim of a 'native' resource would be weakened.","supporting_citations":[],"review_version":1}