{"id":"ac7b9a7b-ba7c-42a6-bb5b-777064e93956","arxiv_id":"2506.11587","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of HTA guidelines and 195 French Transparency Committee ITCs shows that indirect treatment comparisons rarely influence reimbursement decisions, and acceptance varies by disease area and method.","lead":"This paper reviews international guidelines on indirect treatment comparisons (ITCs) and analyzes 195 ITCs submitted to the French HAS Transparency Committee between 2021 and 2023. It finds that only 13.3% of these ITCs influenced reimbursement decisions, with higher acceptance in genetic diseases and for methods using individual patient data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 13.3% 'influenced' figure counts 'not rejected' ITCs (20/195) together with truly 'accepted' ITCs (6/195); if influence is reserved for accepted ITCs, the rate is 3.1%.","rationale":"The reader identified the reliance on the written ASMR summary as the weakest assumption. That is a reasonable concern, but the manuscript contains a more direct, internal problem: the headline 13.3% is obtained by pooling 'Acceptable' and 'Not rejected' ITCs, even though the Methods define 'Not rejected' as considered with 'clear stated issues'. The paper's own breakdown in Figure 1 gives 6 accepted and 20 not-rejected among the 26 'considered' ITCs, so the influence rate depends entirely on whether 'not rejected' counts as influence. This is a load-bearing concern because the abstract's central quantitative claim changes by a factor of more than four under the narrower, arguably more natural reading. It is not an external consensus dispute or a stylistic complaint; it is an inconsistency between the Methods, the Results, and the Abstract. The recommended remedy is a re-analysis that reports both rates and defines 'influence' operationally. Because the qualitative conclusion (ITCs rarely affect TC decisions) would likely survive either reading, and because the reader already issued a CONDITIONAL verdict, I do not recommend changing the verdict; I do recommend that the conditional be explicitly tied to the accepted-versus-not-rejected decomposition.","tokens_in":11930,"tokens_out":7054,"duration_ms":64305,"concrete_test":"Ask the authors to release the per-ITC extraction table (or re-code from the 138 opinions) and to recompute all 'influenced' and 'acceptance' rates with 'Acceptable' and 'Not rejected' kept separate. In particular, recompute the abstract's 13.3% using only the 'Acceptable' category; if it falls to approximately 3.1% (6/195), the headline must be revised to report the two rates separately. As a secondary check, compare ASMR outcomes for the 6 accepted versus the 20 not-rejected ITCs to see whether 'Not rejected' ITCs actually changed any decision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central statistic in the abstract, 'Only 13.3% of these ITCs influenced TC decision-making,' is not supported by the paper's own category definitions. The Methods define three acceptability outcomes: 'Acceptable', 'Not rejected' (considered for the decision-making but with clear stated issues), and 'Not acceptable'. The Results then report 26/195 (13.3%) ITCs as 'considered'. Using the paper's own Figure 1 breakdown, only 6 ITCs were 'accepted' (3/92 main-source plus 3/103 non-main-source, i.e. 3.1%), while 20 ITCs were 'Not rejected' (17/92 plus 3/103, i.e. 10.3%). The 13.3% numerator is therefore a composite of accepted plus not-rejected ITCs. If 'influence' means the committee positively relied on the ITC, the headline should be 3.1%; if 'influence' includes ITCs explicitly described as having 'clear stated issues', then the Methods definition conflicts with the ordinary meaning of 'influenced'. Either reading makes the headline number an artifact of category labeling rather than a direct measurement. The reader's concern about reliance on the ASMR summary is related but secondary: even granting that the summary is a faithful record, the paper has already merged two different categories in its lead statistic.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a pragmatic review of indirect treatment comparison (ITC) guidelines from seven HTA and multistakeholder bodies, together with a systematic analysis of all French Transparency Committee (TC) opinions published between 2021 and 2023 that mention at least one ITC. The authors extracted 195 ITCs from 138 TC opinions, classified each ITC's acceptability on the basis of how it was treated in the TC's ASMR summary, and reported acceptance rates by disease area, by statistical method, and by whether the ITC was the primary source of evidence. The paper's central quantitative claim is that only 13.3% of submitted ITCs influenced TC decision-making, with higher acceptance in genetic diseases and for IPD-based methods, and it also describes the most common limitations identified by the TC.","tokens_in":12166,"tokens_out":5187,"duration_ms":50041,"significance":"If the classification issues were resolved, this would be a valuable and comprehensive descriptive study of ITC use in French HTA decision-making. Its strengths include a census-like sampling frame (all TC opinions in a three-year window), a structured 41-variable extraction template, manual review with a documented 10% quality-assurance check, and a transparent list of the most frequently cited limitations. The guideline comparison in Table 1 is a useful synthesis. The paper is, however, descriptive rather than methodological, and the headline influence rate is not currently supported by the paper's own category definitions. The dataset and extraction approach could serve as a reproducible baseline for future evaluations if the terminological and classification ambiguities are corrected.","major_comments":[{"comment":"The headline 'Only 13.3% of these ITCs influenced TC decision-making' conflates the category 'Acceptable' with 'Not rejected.' The Methods define three distinct outcomes: 'Acceptable,' 'Not rejected' (considered for the decision-making but with clear stated issues), and 'Not acceptable.' The Results report 26/195 (13.3%) as 'considered,' and the text around Figure 1 implies 6 accepted ITCs (3.1%) and 20 not rejected ITCs (10.3%). If 'influence' is reserved for ITCs the TC positively relied on, the rate is 3.1%; if it includes ITCs described as having 'clear stated issues,' the abstract's wording needs an explicit definition. The abstract, Results, and Discussion should report the three categories separately and use consistent terminology such as 'considered,' 'accepted,' and 'not rejected' rather than using 'influenced' for the combined category.","section":"Abstract; Methods (Overall acceptability); Results (paragraph beginning 'Among the 195 indirect comparisons')"},{"comment":"The sentence 'the percentage of accepted indirect comparisons is higher when the ITC is the main source for comparing effectiveness (21.6% vs 5.8%)' directly contradicts the preceding sentence reporting 3% acceptance in both the main-source and non-main-source groups. The values 21.6% (20/92) and 5.8% (6/103) correspond to the combined 'accepted or not rejected' category, not to 'accepted' alone. This incorrect label propagates to the discussion of IPD-based versus NMA-based acceptance rates and should be corrected, with the accepted and not-rejected subcategories reported separately throughout.","section":"Results, paragraph after Figure 1"},{"comment":"The acceptability classification depends entirely on whether and how the ITC is mentioned in the TC's ASMR summary. This assumes that the written summary fully captures the role of the ITC in the actual decision. A TC could weigh an ITC without citing it, or mention it without substantive influence. This measurement-validity assumption should be acknowledged as a limitation and, if possible, tested by comparing a sample of classifications against the full TC dossiers rather than only the opinion summaries.","section":"Methods (Overall acceptability definition)"}],"minor_comments":[{"comment":"The terms 'accepted,' 'considered,' and 'influenced' are used interchangeably; please standardize the terminology to match the definitions in the Methods.","section":"Throughout"},{"comment":"The abstract refers to 'important clinical benefit' while the Results report 'important SMR'; use one consistent term for this outcome.","section":"Abstract and Results"},{"comment":"The sentence 'This article are based on two complementary works' should read 'This article is based on two complementary works.'","section":"Methods, first sentence"},{"comment":"The automated keyword algorithm is described only as 'based on key words' with a reference to Supplementary Methods; please provide the exact search strategy and validation results in the main text or a fully available supplement.","section":"Methods, TC opinion screening"},{"comment":"The table uses both 'X' and 'O' but the legend explains only 'X' and 'o'; unify the case and clearly define all symbols.","section":"Table 1 legend"},{"comment":"Some acceptance-rate comparisons are based on very small denominators (e.g., unadjusted comparisons 3/8); please report exact counts with confidence intervals or explicit cautions about small-sample comparisons.","section":"Figure 2 and accompanying text"},{"comment":"The statement 'This is in line with HTA recommendations... (HAS, 2019)' cites a year not present in the reference list; either add the reference or correct the citation.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a descriptive policy review rather than a methodological advance, so fit with a methods-oriented journal may be a scope consideration for the editor. Given the industry funding and the reliance on manual classification, the authors should also report inter-rater reliability statistics and state the funder's role in study design, analysis, and publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2506.11587. First, it is a genuinely useful descriptive dataset: 1,218 French Transparency Committee opinions from 2021-23, 138 with at least one indirect treatment comparison (ITC), 195 ITCs manually extracted with a 10% QA check, and a taxonomy of limitations. That is new and worth having. Second, the headline result—'Only 13.3% of these ITCs influenced TC decision-making'—is not supported by the paper's own categories. The Methods define three outcomes: acceptable, not rejected (considered but with clear stated issues), and not acceptable. The Results report 26/195 'considered'. From the paper's own numbers, only 6 ITCs were actually 'accepted' (3.1%); 20 were 'not rejected'. The abstract's 'influenced' merges those two categories. In the Discussion, the same 26 are described as 'taken into account', which is closer to the truth, but the abstract and several other sentences say 'accepted' when they mean 'accepted or not rejected' (e.g., the 21.6% vs 5.8% for main vs non-main source is clearly the combined rate, not the accepted rate). So the headline number is an artifact of category labeling. The qualitative conclusion—ITCs rarely drive reimbursement decisions—survives, because even 13.3% is low and the accepted-only rate is 3.1%. But the paper needs to be explicit about which number it is reporting.\n\nWhat is good: the guideline comparison table is a handy synthesis, the disease-area and method acceptance patterns are interesting (genetic diseases 34.4% vs oncology 10.0%; IPD methods 23.1% vs NMA 4.2%), and the limitation frequencies (heterogeneity/bias 59%, lack of data 48%) are practically useful for manufacturers. The extraction protocol is described well enough to be reproduced, though the underlying data and algorithm are not released.\n\nSoft spots: the acceptability classification relies on whether the TC summary mentions the ITC, a reasonable proxy but subjective; there are no confidence intervals or statistical tests (fine for a descriptive review, but should be labeled as such); and the internal inconsistency in 'accepted' vs 'accepted or not rejected' needs fixing. The Astra Zeneca funding is disclosed, and the findings are not flattering to industry, so I do not see a major bias.\n\nWho this is for: HTA practitioners, manufacturers, and methodologists who want a snapshot of current French practice. I would send it to peer review—a useful empirical contribution that needs careful revision, mostly to make the outcome definitions consistent and the headline match the data. I would bring it to a reading group as a case study in how HTA evidence is actually used, but I would not quote the 13.3% figure as-is.","headline":"Useful French HTA dataset, but the 13.3% 'influenced' headline loses the paper's own distinction between accepted and merely not-rejected ITCs.","tokens_in":12716,"tokens_out":5667,"would_cite":true,"duration_ms":44243,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Only 13.3% of indirect treatment comparisons submitted to the French Transparency Committee from 2021 to 2023 influenced reimbursement decisions, with acceptance ranging from 34.4% in genetic diseases to 4.2% for network meta-analyses.","keywords":["indirect treatment comparison","health technology assessment","French Transparency Committee","network meta-analysis","reimbursement decisions","single-arm trials","evidence synthesis","real-world data"],"falsifier":"Look for a sample of the 169 ITCs classified as not influencing decisions in the committee's internal records or in interviews with committee members: if a substantial share were actually discussed and used in the SMR or ASMR reasoning, the 13.3% estimate is an artifact of written reporting. A simpler check would re-run the same dataset coding influence from whether the ITC changed the final rating relative to the manufacturer's requested rating and compare the two acceptance rates.","tokens_in":11731,"feed_emoji":"💊","tokens_out":8281,"duration_ms":76112,"temperature":0.7,"pith_summary":"Indirect treatment comparisons (ITCs) compare a new drug against a comparator using data from separate trials rather than a head-to-head study. The paper tries to establish that, despite their growing use in French reimbursement submissions, these comparisons rarely shape the final decision. In 138 Transparency Committee opinions published between 2021 and 2023, the authors identified 195 ITCs and found that only 26 (13.3%) influenced the committee's decision-making. Acceptance was uneven: 34.4% in genetic diseases, 11.1% in autoimmune diseases, and 10.0% in oncology, and methods using individual patient data performed better (23.1%) than network meta-analyses (4.2%). The paper also reviews international guidelines and concludes they largely agree, so the gap is not mainly conflicting regulatory standards but practical failure to meet existing methodological expectations.","feed_headline":"Indirect comparisons sway French payers 13% of the time","feed_subtitle":"Acceptance ranged from 34% in genetic diseases to 4% for network meta-analyses.","key_machinery":"The central instrument is an acceptability classification built from the Transparency Committee's published summaries. Each ITC was tagged as acceptable, not rejected, not acceptable, or unclear depending on whether and how the decision summary mentioned the ITC, with the ITC rather than the drug opinion as the unit of analysis. A 41-variable extraction grid and a three-branch taxonomy of limitations (data, methodology, uncertainty) let the authors quantify acceptance by therapeutic area, by statistical method, and by whether the ITC was the primary evidence source.","core_discovery":"Across all French Transparency Committee opinions from 2021 to 2023, 138 contained at least one indirect comparison, yielding 195 ITCs for analysis. Only 13.3% were considered in the ASMR decision; 86.7% were not. When an ITC was the primary source of comparative effectiveness, the proportion of important clinical benefit fell to 60.9% versus 73.4% when randomized controlled trials anchored the comparison, and the proportion of insufficient benefit rose from 9.6% to 18.8%. The committee's most frequent criticisms were heterogeneity or risk of bias (59%), lack of or unclear data (48%), statistical-methodology problems (29%), study-design concerns (27%), small sample size (25%), and variability in outcome definition or timing (20%). The authors read the low acceptance of network meta-analyses and Bucher comparisons (4.2%) as a sign that strict homogeneity and consistency conditions are often unmet, while the higher acceptance of unadjusted comparisons reflects settings of extreme unmet medical need rather than methodological preference.","pith_inferences":["The 13.3% figure is only as good as the written summaries: a natural validation would compare the summary-based classification against the committee's internal deliberation records or manufacturers' own accounts of what was discussed.","The higher acceptance of unadjusted comparisons is probably confounded with disease context, since companies tend to submit them only where unmet need is extreme and treatment effects are dramatic; matching on disease area could separate method quality from context.","If the 2023 HAS doctrine changes ITC requirements, the acceptance rate among opinions from 2024 onward would be a direct test of whether clearer guidance actually raises the influence of indirect comparisons.","The guideline comparison suggests that a common European ITC reporting template could reduce duplication, a policy implication the paper notes but does not test."],"forward_implications":["Manufacturers submitting ITCs to French reimbursement reviews should expect most to be ignored in the final ASMR decision, even when the ITC is the only comparative evidence.","The method gradient means network meta-analyses and Bucher comparisons are unlikely to influence decisions unless homogeneity and consistency are convincingly demonstrated, whereas IPD-based or unadjusted comparisons in rare and genetic diseases have a better chance.","Relying on an ITC as the primary source of comparative evidence is associated with a lower probability of an important clinical benefit rating and a higher probability of an insufficient rating.","Because international guidelines largely agree, a single well-conducted ITC dossier could in principle serve multiple European HTA bodies, with the main national divergence centered on population-adjusted methods."],"supporting_citations":[{"why":"Documents the increasing use of external comparators in HTA submissions, the trend the paper set out to examine.","marker":"[1]"},{"why":"Shows the rising use of population-adjusted indirect comparisons in the NICE submission process, defining the method-specific landscape.","marker":"[2]"},{"why":"Provides the French HTA guidance on indirect comparisons that anchors the Transparency Committee's expectations.","marker":"[5]"},{"why":"Defines the SMR and ASMR evaluation principles that structure the paper's outcome variables.","marker":"[9]"},{"why":"Supplies the EUnetHTA methodological guideline for direct and indirect comparisons, one of the international standards compared.","marker":"[13]"},{"why":"Provides ISPOR good research practices for indirect treatment comparisons and network meta-analysis used in the guideline comparison.","marker":"[31]"},{"why":"Gives the IQWiG general methods that take a strict stance against unanchored population-adjusted comparisons, used to interpret acceptance differences.","marker":"[35]"}],"fun_headline_variants":["Only 13% of indirect comparisons sway French payers","Indirect evidence rarely decides French reimbursement","French payers accept indirect comparisons 13% of the time","Genetic diseases see 34% acceptance for indirect comparisons","Indirect comparison influence on French decisions: 13%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The classification of an ITC as influencing the decision depends entirely on whether the Transparency Committee's written ASMR summary mentions it, so an ITC that was weighed but not cited is counted as not influencing the decision.","fun_headline_variants_meta":{"raw":{"variants":["Only 13% of indirect comparisons sway French payers","Indirect evidence rarely decides French reimbursement","French payers accept indirect comparisons 13% of the time","Genetic diseases see 34% acceptance for indirect comparisons","Indirect comparison influence on French decisions: 13%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1356,"prompt_tokens":1067,"completion_tokens":289,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":212}},"tokens_in":683,"tokens_out":289,"duration_ms":3533,"temperature":1.0,"reasoning_tokens":212,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:18.486218+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look for a sample of the 169 ITCs classified as not influencing decisions in the committee's internal records or in interviews with committee members: if a substantial share were actually discussed and used in the SMR or ASMR reasoning, the 13.3% estimate is an artifact of written reporting. A simpler check would re-run the same dataset coding influence from whether the ITC changed the final rating relative to the manufacturer's requested rating and compare the two acceptance rates.","supporting_citations":[{"cited_title":"Use of External Comparators for Health Technology Assessment Submissions Based on Single-Arm Trials","cited_arxiv_id":null,"evidence_quote":"Documents the increasing use of external comparators in HTA submissions, the trend the paper set out to examine."},{"cited_title":"CO163 The Increasing Use of Population- Adjusted Indirect Comparisons in the NICE Health Technology Assessment (HTA) Submission Process and the Response to These Methods","cited_arxiv_id":null,"evidence_quote":"Shows the rising use of population-adjusted indirect comparisons in the NICE submission process, defining the method-specific landscape."},{"cited_title":"Indirect comparisons: Methods and validity","cited_arxiv_id":null,"evidence_quote":"Provides the French HTA guidance on indirect comparisons that anchors the Transparency Committee's expectations."},{"cited_title":"Doctrine de la commission de la transparence (CT) Principes d’évaluation de la CT relatifs aux médicaments en vue de leur accès au remboursement","cited_arxiv_id":null,"evidence_quote":"Defines the SMR and ASMR evaluation principles that structure the paper's outcome variables."},{"cited_title":"Methods Guideline-D4.3.2-Methodological-Guideline-on-Direct-and-indirect- comparisons-V1.0.pdf","cited_arxiv_id":null,"evidence_quote":"Supplies the EUnetHTA methodological guideline for direct and indirect comparisons, one of the international standards compared."},{"cited_title":"Conducting indirect- treatment-comparison and network-meta-analysis studies: report of the ISPOR Task Force on Indirect Treatment Comparisons Good Research Practices: part 2","cited_arxiv_id":null,"evidence_quote":"Provides ISPOR good research practices for indirect treatment comparisons and network meta-analysis used in the guideline comparison."},{"cited_title":"Yes”,”No","cited_arxiv_id":null,"evidence_quote":"Gives the IQWiG general methods that take a strict stance against unanchored population-adjusted comparisons, used to interpret acceptance differences."}],"review_version":1}