{"id":"a80c0e69-0194-4391-a91f-62bbb084c2f2","arxiv_id":"2501.08744","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Timeline, ridgeline, and split-violin plots can visualise bevacizumab's randomised trial evidence across seven cancer types and support judgements about borrowing information across indications.","lead":"This paper develops new charts, called evidence maps, to show all clinical trial results for the cancer drug bevacizumab across seven cancer types at once. It aims to help health agencies see how evidence accumulates and decide whether results from one cancer can inform another.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Apparent transcription errors in extracted hazard ratios (Table S2) propagate into the cumulative meta-analyses and the cross-indication overlap claim; the case-study demonstration is not robust until these are corrected.","rationale":"The paper's central claim is methodological: graphical displays of multi-indication evidence can help decision-makers judge when cross-indication borrowing is plausible. That claim does not require a perfectly exhaustive evidence base, so the reader's weakest assumption—non-comprehensive searches—weakens the case-study illustration but does not by itself undermine the proposed visualisation approach. However, the paper goes further and makes a specific empirical assertion from the ridgeline plots: that the treatment effect of bevacizumab is similar across indications. That assertion depends entirely on the extracted HRs and CIs being correct. The apparent errors in Table S2 are directly load-bearing because they alter the precision of key trials before the cumulative meta-analyses are run. E3200, for instance, is one of the earliest and largest colorectal cancer trials; replacing its upper CI bound from 1.89 to 0.89 changes the variance by roughly an order of magnitude and therefore changes its contribution to every subsequent colorectal cumulative analysis. The same applies to the other suspect entries. The lack of accompanying code or data means a reader cannot check whether these are isolated typos or systematic extraction problems. In that sense the concern is not about methodology but about reproducibility and the validity of the demonstration. I do not think this overturns the paper's core methodological contribution, and it does not require rejecting the manuscript; it does reinforce the CONDITIONAL verdict, because the authors should correct the data tables, provide the dataset and code, and confirm that the visual conclusions survive correction. I therefore keep the reader's verdict unchanged.","tokens_in":43916,"tokens_out":7133,"duration_ms":78376,"concrete_test":"Correct Table S2 (at minimum E3200 OS to 0.75 (0.63, 0.89), RIBBON-2 OS to 0.90 (0.71, 1.14), and AVF2107 OS to 0.66 (0.52, 0.84)), then rerun the cumulative IP/CP/HMA analyses for OS and PFS and compare the pooled log-HRs and credible interval widths against Tables S4/S5. If any pooled estimate shifts by more than 0.05 on the log scale or any credible interval width changes by more than 20%, the visual 'overlap' and cross-indication similarity conclusion are not robust to data corrections. Regenerate Figures 4-6 from the corrected data to confirm whether the displayed patterns persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest quantitative conclusion—that the OS and PFS ridgeline curves overlap, suggesting similar bevacizumab effects across indications (Section 5.2)—rests on the extracted HRs in Table S2. Several entries appear to be transcription errors. E3200 OS is listed as 0.75 (0.63, 1.89), but the published E3200 OS HR is 0.75 (0.63, 0.89); RIBBON-2 OS is listed as 0.90 (0.71, 1.33), whereas the published value is 0.90 (0.71, 1.14); and AVF2107 OS appears as '066 (0.52, 0.84)', missing the leading zero and decimal. These errors change standard errors and therefore meta-analytic weights: E3200 enters the colorectal cumulative OS meta-analysis early, and its erroneous CI substantially underweights a key trial. The cumulative IP/CP/HMA estimates in Tables S4/S5 and the density curves in Figures 4-6 are all built from such entries. Because no code or machine-readable dataset is provided, the extent to which these errors propagate into the displayed 'overlap' and into the split-violin comparisons cannot be independently assessed. The visualisation methodology may still be sound, but the case-study evidence for cross-indication similarity is only as reliable as the data feeding it. This is more concrete than the acknowledged non-comprehensive search: even within the assembled trials, the numeric evidence base appears corrupted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops and demonstrates visualisation tools—timeline, ridgeline, and split-violin plots—for comparing randomised controlled trial evidence across multiple indications of a single oncology drug, using bevacizumab as a case study. The authors assemble a dataset of 41 trials across seven licensed cancer types, extract hazard ratios for overall and progression-free survival, fit three Bayesian hierarchical meta-analysis models (independent, common, and hierarchical), and display cumulative meta-analysis results over time. The central claim is that such graphical representations help analysts judge whether treatment effects are similar across indications and whether cross-indication borrowing is appropriate; the authors assert that the ridgeline plots show overlapping curves within and across indications, suggesting similar bevacizumab effects.","tokens_in":44166,"tokens_out":2642,"duration_ms":27818,"significance":"If the underlying data and analyses are correct, the paper offers a genuinely useful and practical visualisation toolkit for health technology assessment, where multi-indication drugs are increasingly common. The use of standard Bayesian meta-analysis models with weakly informative priors is methodologically defensible, and the case-study demonstrates how these displays can support decisions about evidence borrowing. The value of the contribution depends on the integrity of the extracted data and the reproducibility of the figures. The paper does not provide code or a machine-readable dataset, and the supplementary tables contain several apparent transcription errors that propagate into the cumulative meta-analyses and the cross-indication overlap claim. These issues are fixable but are currently load-bearing for the case-study demonstration.","major_comments":[{"comment":"Table S2 contains multiple apparent transcription errors in extracted hazard ratios. E3200 OS is reported as 0.75 (0.63, 1.89), whereas the published Giantonio (2007) result is 0.75 (0.63, 0.89). RIBBON-2 OS is listed as 0.90 (0.71, 1.33), while the published Brufsky (2011) value is 0.90 (0.71, 1.14). AVF2107 OS appears as '066 (0.52, 0.84)', missing the leading zero and decimal point. These errors directly change the standard errors used in the cumulative meta-analyses, and because E3200 enters the colorectal cumulative OS analysis early, its inflated CI materially underweights a key trial. The pooled estimates in Tables S4/S5 and the density curves in Figures 4-6 are built from these entries, so the cross-indication overlap conclusion in Section 5.2 is not robust to the current data. The authors must correct the table, re-run the analyses, and verify that the displayed patterns and conclusions remain unchanged.","section":"Supplementary Table S2 (data extraction, OS)"},{"comment":"The cumulative meta-analysis results and the split-violin comparisons depend entirely on the extracted HRs and their CIs. Because no code or machine-readable dataset is provided, the effect of the transcription errors above cannot be independently assessed. Given that at least three OS entries in Table S2 are demonstrably wrong, the authors should either supply the cleaned dataset and analysis code or report a full re-analysis of all results, including Tables S4-S7 and Figures 4-6, after correcting the extracted data.","section":"Section 5.3 and Supplementary Tables S4/S5"},{"comment":"The authors acknowledge that 'due to time and resource constraints the searches conducted were not comprehensive.' This limitation is appropriate, but it applies to the central claim that the observed overlap of treatment effects across indications supports similarity. If important trials were missed, the ridgeline plot overlap and the cumulative meta-analysis comparisons could change. The manuscript should explicitly temper the conclusion that bevacizumab's effect is similar across indications, presenting it as conditional on the assembled, non-exhaustive evidence base rather than as a general finding.","section":"Section 6 (Discussion, limitations)"},{"comment":"The maturity visualisations are severely limited by the very low number of trials reporting event counts—as the authors note in Section 5.1 and 6, most entries in Tables S2 and S3 are 'NR' for events. This means the maturity plots convey little comparative information across indications. Since the paper presents maturity as one of the key features to display, the authors should either present a quantitative summary of how many trials contributed usable maturity data or explicitly downgrade the maturity visualisation from a demonstrated tool to a prototype that requires more complete reporting.","section":"Section 3.3.1 and Figures 3(d), S2, S3"}],"minor_comments":[{"comment":"The entry '066 (0.52, 0.84)' should read '0.66 (0.52, 0.84)'; the missing leading zero and decimal point is a typographical error that would confuse any reader attempting to reproduce the data.","section":"Supplementary Table S2 (AVF2107)"},{"comment":"The PFS entry for E3200 is given as '0.61 (0.48, 078)', which is missing a decimal point before 78; it should be '0.61 (0.48, 0.78)'.","section":"Supplementary Table S3 (E3200 PFS)"},{"comment":"The within-indication SD entry '0.231 (0.015, 0.910 0.166' is missing a closing parenthesis and appears to concatenate two numbers; this should be corrected.","section":"Table S5 (Glioblastoma, 31/12/2012)"},{"comment":"The text states that an arbitrary gap of 2 months was added between reporting points to avoid overlap; this should be clearly noted in the figure caption, not only in the body text, so that readers do not interpret the timeline positions as exact dates.","section":"Figure 2 and Section 5.1"},{"comment":"The ridgeline plots for colorectal, breast, and ovarian cancers are described as 'difficult to interpret' due to clustering; the supplementary ordered plots (Figure S6) help, but the main-text figures could benefit from an explicit visual cue (e.g., colour or faceting) to make the overlap claim easier to verify.","section":"Section 5.2, Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable fit for a statistical applications journal, but the case-study data extraction errors are serious enough that the current numerical results cannot be trusted. I would ask the authors to correct the supplemental tables, re-run all analyses, and ideally provide the cleaned dataset and analysis code. If the corrected analyses change the cross-indication overlap conclusion, the manuscript's central demonstration would need to be substantially revised, possibly to a more modest claim about the visualisation methodology alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a worthwhile methods paper. The authors combine timeline, ridgeline, and split-violin displays with cumulative meta-analysis under three evidence-sharing models to picture how evidence for a multi-indication drug accumulates across cancer types. That cross-disease evidence-map angle is new as far as I know, and the charts are genuinely informative for HTA discussions about when borrowing across indications is plausible. The modelling follows Singh et al., which is honest, and the limitations section is unusually candid about the non-comprehensive search and the poor event reporting.\n\nThe soft spot is concrete and sits in the data. Table S2 lists E3200 OS as 0.75 (0.63, 1.89); the published E3200 OS HR is 0.75 (0.63, 0.89). RIBBON-2 OS is listed as 0.90 (0.71, 1.33) versus the published 0.90 (0.71, 1.14). AVF2107 OS appears as '066 (0.52, 0.84)'. Table S3 has a similar missing decimal in E3200 PFS. These are not cosmetic. The CIs feed the cumulative meta-analyses and the density curves that underlie the 'curves overlap across indications' claim. The wrong E3200 CI makes a key colorectal trial look much less precise than it is, changing its weight in the early colorectal OS meta-analysis and the CP/HMA results. I can't tell how far the errors propagate because the paper ships no code or machine-readable data. The visualisation framework may survive correction, but the case-study demonstration is not robust until the tables are fixed and the analyses rerun.\n\nOther concerns are minor by comparison. The search is acknowledged non-comprehensive, maturity data are mostly NR, and there is no formal usability assessment. None of that kills the methods contribution. The citation pattern is fine; relying on Singh et al. for the models is appropriate.\n\nVerdict: send it to peer review, but with a request for corrected tables and code/data. The core idea deserves referee time; the current case-study numbers do not.","headline":"Useful visualisation framework for multi-indication oncology evidence, but the bevacizumab case study rests on a supplementary table with apparent transcription errors that need correcting before the quantitative claims can be trusted.","tokens_in":44744,"tokens_out":2498,"would_cite":true,"duration_ms":25599,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Clear graphical maps of a drug's evidence across all its licensed cancer types can reveal when trial results are exchangeable and support decisions about borrowing evidence across indications.","keywords":["evidence visualisation","multi-indication drugs","bevacizumab","health technology assessment","cumulative meta-analysis","ridgeline plots","evidence synthesis","oncology"],"falsifier":"Re-run the data assembly with a comprehensive systematic search, redraw the ridgeline plots, and formally estimate between-indication heterogeneity; if the log hazard-ratio densities separate by indication or if the between-indication standard deviation in the hierarchical model is large, the paper's visual case for exchangeability fails.","tokens_in":43705,"feed_emoji":"📊","tokens_out":5414,"duration_ms":54306,"temperature":0.7,"pith_summary":"This paper argues that a drug licensed for several cancer types should not be appraised using only the evidence from one indication. It develops three kinds of plots, timelines, ridgeline plots, and split-violin plots, that lay out the full bevacizumab evidence base across seven licensed cancer types, showing when trials ran, when results were reported, how mature and precise the estimates are, and how pooled estimates change under different evidence-sharing assumptions. The maps are meant to give health technology assessment analysts a visual basis for judging whether the drug's effect is similar enough across indications to borrow information across them. The case study suggests that the log hazard-ratio densities for overall and progression-free survival overlap within and across indications, so borrowing across indications is plausible for bevacizumab.","feed_headline":"Bevacizumab's effect looks similar across seven cancers","feed_subtitle":"Timeline, ridgeline and split-violin maps show when health technology assessment can borrow evidence across cancers.","key_machinery":"The carrying objects are three plot types and three synthesis models. Timeline plots put every trial on a shared time axis so that the accumulation, size, precision, and maturity of evidence can be seen as it appears; ridgeline plots stack the distribution of each reported log hazard ratio by year so that the overlap of densities across indications can be judged directly; split-violin plots place two distributions on either side of a central line to compare overall and progression-free survival across models. The quantitative machinery is the cumulative meta-analysis under three sharing assumptions: the independent-parameter model (no borrowing), the common-parameter model (complete borrowing), and the hierarchical meta-analysis model (borrowing moderated by between-indication heterogeneity), all run in a Bayesian random-effects framework. The plots work by substituting entire densities for point estimates, so the visible overlap of curves becomes the evidence for exchangeability.","core_discovery":"Using 41 randomised controlled trials across seven licensed indications of bevacizumab, the paper constructs evidence maps that display the evolution of overall survival and progression-free survival estimates over more than two decades. The central visual claim is that ridgeline plots, which draw the full density of each trial's reported log hazard ratio instead of only a point estimate and confidence interval, show the curves overlapping within and across indications; the authors read this as suggesting that the treatment effect of bevacizumab is similar across indications. They then fit three cumulative meta-analysis models, no borrowing, complete borrowing, and hierarchical partial borrowing, and show with split-violin plots that the model results are largely consistent, with the no-borrowing model being the least precise. The paper's conclusion is that such graphical summaries give a better understanding of the whole evidence base and can inform judgements about which cross-indication assumptions to make in evidence synthesis for health technology assessment.","pith_inferences":["The same display grammar could be applied to other multi-indication drugs, and the visual overlap of densities would provide a quick screening test for whether cross-indication borrowing is worth modelling.","A natural extension would be to add a quantitative rule of thumb to the split-violin plots, such as the posterior probability that indication-specific effects differ, so that the visual judgement can be audited.","Because progression-free survival is reported earlier and often more precisely than overall survival, the displays could support borrowing on progression-free survival while overall survival evidence is still immature, with the overall survival timeline as a running check on that choice.","The ridgeline overlap is read from an incomplete evidence base, so the plots should be treated as a decision-support display rather than a formal exchangeability test."],"forward_implications":["Health technology assessment analysts can use timeline maps to see at a glance how many trials, how much follow-up, and how precise the results are for each indication before deciding whether to borrow evidence.","The overlapping ridgeline curves give a visual, model-free argument that bevacizumab's effect is exchangeable across indications, making cross-indication borrowing more defensible.","Cumulative meta-analysis displays show that after roughly three studies per indication the pooled estimate stabilises, so later results mostly add precision rather than changing the effect.","Because the common-parameter model gives the most precise estimates but rests on the strongest assumption, the split-violin comparisons make the precision-bias trade-off explicit and place it before the analyst for discussion.","The displays are updateable: as trials report interim or final outcomes, new points and new pooled densities can be added without changing the plotting machinery."],"supporting_citations":[{"why":"Defines the seven licensed indications of bevacizumab that structure the whole case study.","marker":"[18]"},{"why":"Known systematic reviews used to identify the bevacizumab trials that make up the evidence base.","marker":"[9, 20]"},{"why":"Introduce panoramic meta-analysis, the motivation and basis for the hierarchical model that borrows evidence across indications.","marker":"[15-17]"},{"why":"Companion methods paper whose Bayesian model code and discussion of the appropriateness of assumptions the cumulative meta-analyses adapt.","marker":"[30]"},{"why":"Supply the cumulative meta-analysis framework used to display how pooled estimates evolve over time.","marker":"[31, 32]"},{"why":"Provides the events-over-patients maturity measure used to plot evidence maturity in the timeline displays.","marker":"[23]"}],"fun_headline_variants":["Bevacizumab's effect overlaps across 7 cancers","Evidence maps show bevacizumab works across cancers","Bevacizumab: one effect across seven cancers","Ridgeline plots reveal bevacizumab's shared effect"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The visual case for cross-indication similarity rests on the 41 trials found by searches the authors themselves call non-comprehensive, so the overlapping densities could change if missing trials reported different effects.","fun_headline_variants_meta":{"raw":{"variants":["Bevacizumab's effect overlaps across 7 cancers","Evidence maps show bevacizumab works across cancers","Bevacizumab: one effect across seven cancers","Ridgeline plots reveal bevacizumab's shared effect"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000644,"raw_usage":{"total_tokens":3002,"prompt_tokens":1027,"completion_tokens":1975,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":1910}},"tokens_in":643,"tokens_out":1975,"duration_ms":13802,"temperature":1.0,"reasoning_tokens":1910,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:17:56.651711+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the data assembly with a comprehensive systematic search, redraw the ridgeline plots, and formally estimate between-indication heterogeneity; if the log hazard-ratio densities separate by indication or if the between-indication standard deviation in the hierarchical model is large, the paper's visual case for exchangeability fails.","supporting_citations":[{"cited_title":"Efficacy and safety of bevacizumab plus chemotherapy in Chinese patients with metastatic colorectal cancer: a randomized phase III ARTIST trial","cited_arxiv_id":null,"evidence_quote":"Companion methods paper whose Bayesian model code and discussion of the appropriateness of assumptions the cumulative meta-analyses adapt."},{"cited_title":"Bevacizumab in combination with oxaliplatin-based chemotherapy as first-line therapy in metastatic colorectal cancer: a randomized phase III study","cited_arxiv_id":null,"evidence_quote":"Provides the events-over-patients maturity measure used to plot evidence maturity in the timeline displays."}],"review_version":1}