{"id":"1120bcca-534e-49f3-b48c-f4006651a5cb","arxiv_id":"2607.26834","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A Dash-based V0 web app integrates cohort filtering, PyRadiomics extraction, and guardrailed inference for brain-tumor radiomics with inspectable intermediate artifacts.","lead":"Researchers built a lightweight web dashboard that ties together patient tables, radiomic feature extraction, and guarded ML predictions for brain tumors. It matters as a practical attempt to make opaque radiomics pipelines inspectable on ordinary clinical computers.","discovery_kind":"incremental","skeptic_critique":{"model":"grok-4.5","headline":"No stronger load-bearing concern than the reader's: benefit claims rest on module description, not measured usability or responsible-use outcomes.","rationale":"The paper is a cs.SE systems/engineering report. Its implementable content (Dash modules, NIfTI+CSV path, PyRadiomics config upload, inference guardrails, lightweight deployment) is coherent and honestly limited (§V: DICOM, slow extraction, incomplete coverage). The only place the strongest claim can fail is the unevidenced jump from ‘we exposed intermediate artifacts and added guardrails’ to ‘this improves traceability/interpretability/responsible use.’ That is precisely the reader’s weakest_assumption and correctly drives CONDITIONAL: fine as an early V0 description if code is released and benefit language is toned down; not yet evidence of clinical improvement. I did not find a more load-bearing technical flaw (e.g., contradictory pipeline scope, broken data model, or hidden training claims). Novelty and missing artifacts already noted by the reader do not independently sink the architecture claim. Hence verdict stays CONDITIONAL and agreement is full.","tokens_in":9592,"tokens_out":547,"duration_ms":11057,"concrete_test":"Require a minimal user study on the released V0: N≥6 target clinicians complete the same three tasks (filter cohort, extract+append radiomics for one case, run guarded vs mismatched inference) on V0 vs their current tool; pre-register success as significant gains on traceability checklist score and inappropriate-inference attempts blocked, plus SUS/task-time. If no reliable gains, narrow Abstract/Conclusion to implementation-only claims.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that V0 improves traceability, interpretability, and responsible AI use (Abstract; strongest_claim) is load-bearing on the assumption that building three inspectable modules (Data, Radiomics, Prediction) plus guardrails and clinician feedback loops is itself evidence of those improvements. §IV only describes UI behavior (filters, append-to-database, orange mismatch banner); §III.B step 5 mentions qualitative feedback sessions without protocols, themes, or before/after measures; no task-time, error-rate, trust, or auditability metrics appear, and no comparison to 3D Slicer / existing radiomics viewers is given. If that assumption fails, the paper remains a credible engineering report of a Dash prototype but does not support the translational/benefit language. This is exactly the reader's weakest_assumption; I find no deeper internal inconsistency (architecture, formats, or guardrail logic) that would overturn the claim on its own.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript presents V0 of a lightweight, Dash-based web visual analytics system for radiomics-driven ML inference in neuro-oncology. It integrates three modules in one interface—Data (CSV/XLSX cohort ingest, filtering, export, subject registration), Radiomics (NIfTI image/mask upload, PyRadiomics extraction, append-to-database by patient ID), and Prediction (on-demand inference with pre-trained models plus mismatch guardrails that block incompatible model–dataset pairs). Design followed an iterative user-centred process (pipeline decomposition, format analysis, component design, deployment under clinical hardware constraints, clinician feedback). The system is demonstrated on BraTS 2020 (n=236 glioblastoma) and a proprietary INCB skull-base meningioma cohort (n=104). The central claim is that explicit exposure of intermediate workflow artifacts plus guardrails improves traceability, interpretability, and responsible AI use, offering a portable foundation for clinically oriented brain-tumor AI.","tokens_in":9776,"tokens_out":1468,"duration_ms":31557,"significance":"If the contribution is read as an engineering systems report—a portable, inspectable end-to-end radiomics–ML dashboard matched to real clinical workstation constraints—it is a useful and under-served piece of translational infrastructure. Related work correctly notes that most tools stop at dashboards, labelling, or segmentation (e.g., 3D Slicer, MONAI Label) rather than chaining cohort management, feature extraction/fusion, and guarded inference in one interactive system. Strengths include a coherent three-module architecture, explicit intermediate artifacts, inference guardrails, and a deployment model (Dash, no heavy backend) aligned with limited clinical hardware. The paper does not ship quantitative usability evidence, open code/artifacts in the text, or controlled comparisons; significance therefore rests on the systems design and the gap it targets, not on measured clinical impact.","major_comments":[{"comment":"Abstract and §I claim the system was “evaluated on” BraTS 2020 and the INCB cohort and that explicit exposure of intermediate artifacts “improves traceability, interpretability, and responsible use of AI.” §IV (Results) only describes UI behaviour (filters, preview tables, Append to Database, orange mismatch banner); no task-time, error-rate, trust, auditability, or usability metrics, and no before/after or comparison to 3D Slicer / existing radiomics viewers, are reported. §III.B step 5 mentions clinician feedback sessions without protocol, themes, or outcomes. Either add a minimal evaluation (even qualitative thematic summary or small task-based study) or revise the abstract/contribution language to “demonstrated/deployed on” and frame benefit claims as design goals rather than established results.","section":"Abstract; §IV; §III.B.5; §V"},{"comment":"The Prediction module is described as enabling “on-demand inference using the embedded models developed in this project,” yet the manuscript never specifies which models, targets (survival? volumetric response?), training protocol, performance, or how SHAP/explainability (promised in §I contributions and pipeline analysis) is surfaced in V0. §III.B.3 states models are assumed pre-trained and reusable, and preprocessing/training are out of scope—fine for a systems paper—but without naming the embedded artifacts, inputs/outputs, and what the user actually sees at inference time, the “guarded inference” and “explainable pipelines” claims cannot be assessed or reproduced. Add a short subsection or table listing model cards (task, features required, performance on the two cohorts, explanation modality if any).","section":"§I contributions; §III.B.3; §IV Prediction module"},{"comment":"§II asserts that a coherent system implementing the full inspectable radiomics–ML chain “remains unaddressed.” That gap claim is load-bearing for novelty. Several cited and uncited tools (3D Slicer radiomics extensions, Severn et al.’s explainable radiomics pipeline [23], various research dashboards) partially cover extraction + explanation. The manuscript should more carefully delineate what is new in V0 (single web app, cohort↔radiomics join, inference guardrails, clinical-hardware deployment model) versus what is incremental, so the contribution boundary is falsifiable rather than absolute.","section":"§II Related Works"}],"minor_comments":[{"comment":"Title and abstract say “Explainable” / “explainability (e.g., via SHAP),” but §IV does not describe any explanation UI in V0. Align title/abstract with implemented scope or move SHAP to future work explicitly in Results.","section":"Title; Abstract; §I; §IV"},{"comment":"Table I header says “BRATS2020” and reports n=236; BraTS 2020 training set is conventionally larger—clarify the exact subset and inclusion criteria (e.g., subjects with survival labels only).","section":"Table I; §III.A"},{"comment":"Figure 1 is described but not available in the text package; ensure the bottom-right screenshot legibly shows the three modules and guardrail banner for camera-ready.","section":"Figure 1"},{"comment":"Minor typos/spacing: “cli nical,” “tran s-parency,” “inf erence,” “structur ed,” “pre-tra ined,” “coh ort,” “workﬂow,” “F ondazione,” author list formatting. Copy-edit for line-break artifacts from the PDF.","section":"Throughout"},{"comment":"Index terms and keywords are appropriate; consider adding “visual analytics” and “clinical decision support systems” for discoverability in cs.SE / medical informatics venues.","section":"Index Terms"},{"comment":"§V limitations correctly flag DICOM, extraction latency, and incomplete coverage—good. A brief sentence on data governance / PHI handling for a hospital-deployed web app would strengthen the translational discussion.","section":"§V"}],"recommendation":"major_revision","confidential_remarks":"Fit is reasonable for a software-engineering / clinical-informatics venue as a systems/tool paper, but overstated evaluation and explainability language will draw pushback. If the journal expects empirical HCI or clinical validation, this needs major revision; if it accepts well-motivated tool descriptions with clear limitation statements, toning claims and adding model cards may suffice. No integrity concerns; novelty is incremental but real on the integration + guardrails + light deploy axis. I did not find internal inconsistency in architecture or guardrail logic beyond the evidence gap the reader flagged."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a systems paper, not a biomarker paper. What is new is a single lightweight Dash interface that joins clinical tables, PyRadiomics extraction from NIfTI+masks, and pre-trained inference with an explicit mismatch block (orange banner) so you cannot run the wrong model on the wrong cohort. That assembly is real plumbing, not vapor.\n\nThey do several things cleanly. The design process matches stated clinical hardware limits (no heavy backend). Intermediate artifacts stay visible: filtered cohort, feature preview, append-by-patient-ID into the active table. Related work is honest about 3D Slicer, MONAI Label, oncology dashboards, and Severn-style explainability viewers; they correctly note that full radiomics-ML chain inspectability in one interactive system is still thin for brain tumors. Limitations (DICOM, slow online extraction, incomplete V0 coverage) are stated without spin. BraTS2020 (n=236) and the INCB meningioma cohort (n=104) are real test beds for the pipeline, not invented entities. Circularity is low; this is not a fitted law dressed up as prediction.\n\nThe soft spot is exactly where the reader and stress-test put it, and it is not deeper than that. Abstract and discussion claim that exposing artifacts “improves” traceability, interpretability, and responsible use. Results §IV only describe modules. Clinician feedback is mentioned as qualitative sessions with no protocol, themes, task times, error rates, or comparison to existing tools. So the translational benefit language is asserted, not measured. That does not break the architecture claim; it means the paper should be read as “we built V0 with these guardrails,” not as evidence of clinical superiority. Code/models are not shipped here, which caps reproducibility for now.\n\nWho it is for: people building clinical-informatics tooling for radiomics translation, and methodologists who care about deployment constraints. Not for someone hunting a new survival model or IDH predictor. I would send it to peer review as an early engineering report; a serious referee can force claim narrowing and a code release. I would not cite it yet for a scientific result, but I would watch the next version if they open the repo and add even light usability numbers.","headline":"Credible V0 engineering report of a Dash radiomics dashboard with sensible guardrails; benefit language outruns the evidence, but the architecture itself is coherent and worth a referee look if claims are narrowed.","tokens_in":10454,"tokens_out":549,"would_cite":false,"duration_ms":10381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A single lightweight web dashboard runs brain-tumor radiomics from cohort tables through feature extraction to guarded ML prediction while keeping intermediate steps visible.","keywords":["brain tumors","radiomics","machine learning","clinical decision support","visual analytics","neuro-oncology","explainable AI"],"falsifier":"A controlled clinician study comparing this dashboard to current tools on the same radiomics-to-prediction tasks, scoring mismatched-model error rates, time to verify intermediate features, and trust ratings—if those measures do not improve, the central claim fails.","tokens_in":10429,"feed_emoji":"🧠","tokens_out":762,"duration_ms":30027,"temperature":0.7,"pith_summary":"AI and radiomics for brain tumors often fail to reach the clinic because pipelines are scattered, hard to inspect, and poorly matched to how clinicians work. This paper presents version zero of a portable web visual-analytics system that puts three jobs in one interface: manage cohorts from clinical tables, extract radiomic features from MRI volumes and segmentation masks, and run pre-trained models only under explicit safety checks. Intermediate artifacts stay on screen so users can see what was filtered, what features were produced, and how they were joined to clinical data. The system was shaped through iterative clinician feedback and exercised on a public glioblastoma cohort and a proprietary meningioma cohort. The authors argue that this inspectable, low-infrastructure design is a practical bridge from lab models to responsible clinical use.","feed_headline":"One web app runs brain-tumor radiomics with safety locks","feed_subtitle":"Cohort filters, feature extraction, and blocked mismatched models sit in a single portable dashboard.","key_machinery":"The V0 three-module dashboard (Data, Radiomics, Prediction): it appends PyRadiomics features into the active clinical table by patient ID and blocks inference with a visible warning when the selected model is incompatible with the chosen dataset.","core_discovery":"The authors establish that a scalable web-based visual analytics system can integrate cohort management, radiomic feature extraction and fusion, and guarded inference with pre-trained models in one interface, and that explicitly exposing the intermediate workflow artifacts improves traceability, interpretability, and responsible use of AI in brain-tumor analysis.","pith_inferences":["The guardrail pattern—block and explain mismatch rather than emit a score—could transfer to other imaging ML settings where wrong model application is a safety risk.","Without reported task-time or error-rate metrics, claims of improved responsible use will stay hard to compare against existing imaging workbenches.","Caching and asynchronous extraction, already flagged as future work, are likely required before routine use on ordinary clinical workstations."],"forward_implications":["Clinics with limited CPU and RAM can run radiomics inference without a heavy dedicated backend.","Explicit model–dataset guardrails reduce silent wrong predictions in neuro-oncology.","Inspectable cohort filters and feature tables create an auditable path from data to prediction.","The same lightweight web pattern can be extended once DICOM support and faster extraction are added.","Pre-trained models become usable at the point of care without forcing clinicians to reassemble fragmented scripts."],"fun_headline_variants":["Web system unites cohort tools, radiomics, and guarded ML for brain tumors","One dashboard exposes radiomics steps for traceable brain-tumor AI","Scalable app locks mismatched models in brain-tumor radiomics pipelines","Visual analytics platform merges features and safe inference for gliomas","Inspectable web workflow ties cohorts to guarded brain-tumor ML"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Describing the modules on two datasets and iterating with clinician feedback is enough evidence that exposing intermediate steps truly improves traceability, interpretability, and responsible use.","fun_headline_variants_meta":{"raw":{"variants":["Web system unites cohort tools, radiomics, and guarded ML for brain tumors","One dashboard exposes radiomics steps for traceable brain-tumor AI","Scalable app locks mismatched models in brain-tumor radiomics pipelines","Visual analytics platform merges features and safe inference for gliomas","Inspectable web workflow ties cohorts to guarded brain-tumor ML"]},"model":"grok-4.5","effort":"low","cost_usd":0.002426,"raw_usage":{"total_tokens":908,"prompt_tokens":701,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":24264000,"prompt_tokens_details":{"text_tokens":701,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":133,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":701,"tokens_out":74,"duration_ms":4609,"temperature":1.0,"reasoning_tokens":133,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T19:41:42.358576+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled clinician study comparing this dashboard to current tools on the same radiomics-to-prediction tasks, scoring mismatched-model error rates, time to verify intermediate features, and trust ratings—if those measures do not improve, the central claim fails.","supporting_citations":[],"review_version":1}