{"id":"ecbc9ce9-7aa7-46dc-b7a0-0ad1c43b574e","arxiv_id":"2501.09906","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In healthcare research, closed LLMs like GPT-4 are used mainly for high-accuracy imaging and diagnostics, while open LLMs like LLaMA are used for customized applications such as mental health chatbots.","lead":"This position paper analyzes 201 language models and 6,198 arXiv papers to compare how open and closed AI models are used in healthcare. It reports that closed models dominate high-performance diagnostic research, while open models enable specialized fine-tuning in areas like mental health and patient communication.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The complementary-roles claim is not supported by the topic analysis as run: Figure 4 reports raw topic patterns without normalizing for closed models' much larger overall mention count, so 'closed models dominate radiology' may just be a popularity effect.","rationale":"The paper's contribution is a descriptive observation about where open and closed LLMs are mentioned in the healthcare literature, not a benchmark evaluation. The reader is rightly cautious about using arXiv mentions as a proxy for actual use. My stress test goes one step further and asks whether the paper's own comparison is valid even under that proxy. The data in Figures 2 and 3 show that closed models have a much larger base of mentions, both overall and in medicine. A topic analysis that reports which topics are most common inside each class is not evidence of specialization unless it is compared with the distribution expected from base rates or from the full medical corpus. For example, if radiology is simply the largest medical topic on arXiv, closed models' higher raw count in that topic would be predicted without assuming any particular affinity. The paper gives no effect sizes, confidence intervals, or statistical tests for Figure 4, and it does not state how the 933 papers that mention both open and closed models are treated. The additional wording in Section 4 about fine-tuning also goes beyond the data: count-based topic modeling cannot distinguish fine-tuning from evaluation or API use. None of this invalidates the position statement's plausibility, but it does mean the central complementary-roles claim currently rests on an unnormalized and unvalidated topic interpretation. The reader's CONDITIONAL verdict is the right level; my recommendation is unchanged, with the condition that the authors release the 404 medical paper IDs and topic assignments and demonstrate that the claim survives normalization.","tokens_in":5043,"tokens_out":5107,"duration_ms":56415,"concrete_test":"Recompute Figure 4 as a normalized contrast: for each BERTopic topic, estimate P(topic | closed-only mention) versus P(topic | open-only mention) from the 404 medical papers, excluding or separately coding the 933 papers that mention both open and closed models, and report bootstrap confidence intervals for the difference. If the closed-vs-open topic differences vanish after normalization, the central complementary-roles claim is not established. As a secondary check, manually inspect a random sample of 50 open-model mental-health papers to verify whether they actually fine-tune the model weights.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is Section 3's move from Figures 3 and 4 to the summary's claim that closed LLMs dominate high-performance tasks such as radiology and medical imaging, while open LLMs enable fine-tuned applications such as mental health support and patient communication. The authors themselves report that closed models receive substantially more overall arXiv attention (Figure 2) and more medical mentions (Figure 3). Yet the topic analysis in Figure 4 is presented as raw topic concentration within each model class, with no denominator, no baseline distribution, and no statistical contrast. Under a null model in which healthcare researchers mention whichever model is more visible or more capable uniformly across topics, closed models would appear more often in every high-volume specialty simply because of their higher base rate. The observed radiology/imaging pattern is therefore indistinguishable from a popularity effect and does not demonstrate a healthcare-specific division of labor. The fine-tuning statement is an additional unsupported inferential step: mentions of LLaMA in mental-health papers are counted, but the paper never checks whether those papers actually fine-tune the weights rather than evaluating or using the base model. Both gaps sit in Sections 2.2 and 3 and are inherited by the abstract and Section 4.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper compares how open-source and closed-source large language models are discussed in the healthcare literature. Using the Stanford CRFM Ecosystem Graph for model counts and arXiv metadata for papers mentioning specific model names, the authors report that closed LLMs (e.g., GPT-4) receive more scientific attention overall and, in the medical subset, appear predominantly in papers on radiology and medical imaging, while open LLMs (e.g., LLaMA) appear more in papers on mental health and patient communication. From these patterns, they conclude that closed models lead high-performance diagnostic applications while open models enable specialized fine-tuned applications, and they suggest hybrid approaches for the future.","tokens_in":5278,"tokens_out":2646,"duration_ms":28807,"significance":"The question the paper addresses—whether open and closed LLMs occupy complementary roles in healthcare—is timely and of genuine interest to the NeurIPS community. The descriptive dataset assembled from 6,198 arXiv papers and 201 models is a useful resource, and the paper is commendably transparent about its retrieval strategy and data sources. However, the central qualitative claim currently rests on unnormalized topic counts and unvalidated topic labels, so the significance of the specific complementary-roles conclusion is not yet established. The paper reads more as a set of hypotheses grounded in a preliminary descriptive analysis than as a demonstrated result, and the analyses would need to be strengthened before the conclusions can be relied upon.","major_comments":[{"comment":"The key claim that closed LLMs dominate high-performance tasks such as radiology and medical imaging while open LLMs dominate mental health and patient communication is not supported by the analysis as presented. Figure 4 reports raw topic concentrations within each model class without normalizing for the substantially different base rates: Figures 2 and 3 show that closed models have many more total medical mentions. Under the null model that researchers mention whichever model is more visible or more capable uniformly across all topics, closed models would appear more often in every high-volume specialty purely because of their higher base rate. To support the complementary-roles claim, the authors should report topic proportions within each class relative to a baseline (e.g., the model class share of all medical-LLM mentions), or perform a statistical contrast such as a chi-square or log-odds ratio per topic that accounts for the overall popularity difference.","section":"Section 3, Figures 3 and 4"},{"comment":"The assertion that the number of open LLMs grows exponentially while the number of closed LLMs increases linearly is made on the basis of visual inspection of Figure 1, with no model fitting, no growth-rate estimate, and no uncertainty quantification. Since this claim is used to motivate the 'democratization' narrative, it is load-bearing and should be substantiated. For example, the authors could fit count models (e.g., Poisson or negative binomial regressions with a log link) to the annual model counts and report rate ratios and confidence intervals, or explicitly state that the distinction is qualitative and based on visual inspection rather than a fitted trend.","section":"Section 3, Figure 1"},{"comment":"The paper uses the presence of a model name in titles and abstracts as a proxy for 'using, evaluating, or mentioning' the model, and then interprets topic co-occurrence as evidence of actual use and fine-tuning. In particular, the claim that open LLMs 'enable researchers to fine-tune models for specific domains, such as mental health and patient communication' is not verified anywhere: a mental-health paper that mentions LLaMA may be evaluating the base model, comparing it with GPT-4, or mentioning it only in a literature-review sentence. The analysis should either include a validation step that checks for fine-tuning-related terms (e.g., 'fine-tun', 'LoRA', 'instruction-tuned') in the open-model medical papers, or the conclusions should be weakened to state that open models receive attention in these topic areas without claiming that the papers actually fine-tune them.","section":"Sections 2.2 and 3"},{"comment":"The BERTopic labels in Figure 4 are presented without any validation. No topic coherence scores, manual inspection details, representative terms per topic, or inter-annotator agreement are reported. Because the topic labels are the direct evidence for the paper's central healthcare claim, the authors should at least show the top terms defining each topic cluster and report a validation measure (e.g., topic coherence or a manual labeling exercise), so that a reader can judge whether labels such as 'radiology and medical imaging' accurately reflect the underlying documents.","section":"Section 3, BERTopic analysis"}],"minor_comments":[{"comment":"The abstract states that the paper 'analyzes the evolving roles' and reaches firm conclusions, but the evidence is only descriptive. Consider adding a phrase such as 'suggest' or 'provide evidence for' to align the abstract with the strength of the analysis.","section":"Abstract"},{"comment":"The notation '147 (103)' and '54 (44)' is not explained. Clarify that the first number is the total identified models and the second is the number matched in arXiv titles/abstracts.","section":"Table 1"},{"comment":"The left panel is labeled 'cumulative ratio' but the denominator is not defined in the text. Specify whether it is the ratio of medical LLM papers to all LLM papers or to all medical papers, and state the base population.","section":"Figure 3"},{"comment":"Some references are incomplete: the GPT-4 technical report [9] lacks a version or date beyond '2023', and reference [5] (Vicuna blog) would benefit from a stable URL or a DOI if available.","section":"References"},{"comment":"The paper uses only arXiv data, which is a significant limitation for healthcare research, where much applied work is published in clinical journals not indexed by arXiv. This limitation is not acknowledged; at minimum, it should be stated in Section 2.2 or Section 4.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a position piece with a plausible but not yet supported central claim. The main fixes—normalizing the topic analysis, validating topic labels, and checking the fine-tuning interpretation—are all within the scope of a revision and would substantially strengthen the paper. I see no circularity in the technical sense, but the interpretive leap from mentions to capability dominance is a correctness risk. The fit between the paper and the journal is acceptable for a position paper, though the authors should be encouraged to frame the conclusions as hypotheses if the additional analyses do not fully confirm them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for the stress-test note. I read the paper, and I think the stress-test is right on target. The paper is a position piece with a small original analysis: 6,198 arXiv papers mentioning open or closed LLMs, drawn from the CRFM ecosystem graph, with topic modeling of the 404 medical subset. That corpus assembly is legitimate and useful, and the paper writes clearly about the open/closed distinction. The descriptive finding that closed models get more overall arXiv attention, and that the mix differs by field, is probably real.\n\nThe problem is the jump from topic co-occurrence to a division of labor in healthcare. Figure 4 is presented as raw topic concentration within each class, with no denominator and no baseline. Since closed models have far more mentions overall, radiology/imaging can dominate the closed-model topic list simply because more papers mention GPT-4 in general. The paper never shows that the proportion of closed-model papers about radiology exceeds the proportion of open-model papers about radiology, or that the topic distribution differs from a random draw given base rates. So the sentence in the abstract and Section 4 that closed models dominate high-performance diagnostic tasks is not supported by the evidence as analyzed.\n\nThere are smaller issues too. The exponential vs linear growth claim in Section 3 is asserted from Figure 1 without a fit or statistical test. The BERTopic labels (e.g., 'mental health', 'patient communication') are not validated against the actual titles. And the fine-tuning claim about LLaMA in mental-health papers relies on mentions rather than any check that weights were actually fine-tuned. No code or data is shipped, which makes the analysis hard to check.\n\nNone of this makes the paper a waste. It is an honest position statement, and the underlying dataset could be a useful starting point. But the central claim needs either normalized topic comparisons or direct performance/deployment evidence. As it stands, the paper is a plausible hypothesis supported by a popularity effect.\n\nWho is it for? Readers interested in bibliometric snapshots of LLM adoption, or workshop attendees. It doesn't deserve a hard reject; it deserves a serious referee who can push the authors to fix the normalization and either soften the claims or add data. I'd accept it for peer review with major revision expected.","headline":"The paper's descriptive dataset is useful, but its central claim about closed LLMs dominating high-performance healthcare rests on unnormalized topic counts and looks like a base-rate effect.","tokens_in":5763,"tokens_out":2033,"would_cite":false,"duration_ms":20695,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that open and closed large language models have settled into complementary roles in healthcare: closed models such as GPT-4 dominate high-stakes diagnostic tasks like radiology and medical imaging, while open models such…","keywords":["large language models","open-source AI","closed-source AI","healthcare AI","medical imaging","mental health","topic modeling","arXiv"],"falsifier":"A benchmark study that compared open- and closed-weight models on a standard radiology diagnostic suite and found open models matching or exceeding closed models on accuracy, or a survey of deployed clinical systems showing open models widely used in imaging, would undermine the paper's central complementarity claim.","tokens_in":4853,"feed_emoji":"🏥","tokens_out":5915,"duration_ms":54096,"temperature":0.7,"pith_summary":"The paper seeks to establish a division of labor in medical AI: closed-weight language models are winning the high-performance, high-stakes tasks such as radiology and multimodal diagnostics, while open-weight models are powering specialized, cost-effective applications such as mental health support and patient communication. The evidence is a quantitative analysis of 6,198 arXiv papers that mention specific LLMs in their titles and abstracts, with a focus on 404 medical papers. Topic modeling shows that papers mentioning closed models concentrate on imaging and diagnostics, whereas papers mentioning open models concentrate on conversational and mental-health applications. If the pattern holds, it means a research community has sorted model choices by task demands: top reasoning where errors are costly, adaptability where personalization matters.","feed_headline":"Closed LLMs dominate imaging; open LLMs power mental-health AI","feed_subtitle":"Topic analysis of 404 medical arXiv papers: closed models lead diagnostics, open models enable personalized care.","key_machinery":"The argument is carried by a quantitative comparison of 6,198 arXiv papers that mention specific open- or closed-weight LLMs in titles and abstracts, drawn from a foundation-model ecosystem graph that classifies each model's access type. From the 404 medical-related papers, BERTopic, a neural topic model using class-based TF-IDF, clusters research themes, and the resulting topic distributions for papers mentioning closed versus open models are compared. The load-bearing claim is that these mention distributions reflect how each type of model is actually being used in healthcare research.","core_discovery":"The paper's central claim is that the scientific community's use of LLMs in healthcare has split along the open/closed axis into complementary roles. Closed-weight models, exemplified by GPT-4, lead in high-complexity, high-stakes tasks such as radiology and medical imaging, where strong reasoning and precision are paramount. Open-weight models, exemplified by the LLaMA series, are the tools of choice for specialized, lower-cost applications such as mental health support, conversational agents, and patient communication, because their released weights allow fine-tuning on targeted datasets. This complementarity, the paper argues, is visible in the topics of medical arXiv papers: those naming closed models concentrate on diagnostic imaging, while those naming open models concentrate on personalized, interactive care.","pith_inferences":["A testable extension: measure actual model use in clinical deployments or code and API logs rather than paper mentions; the mention-based pattern may overstate closed models' real-world share in diagnostics.","If the boundary is set by task stakes rather than capability ceiling, improvements in open models will shrink but not erase the closed-model niche, since closed deployment also offers accountability and data-handling assurances that matter in clinical settings.","The same complementarity may apply beyond healthcare, in regulated domains such as legal advice or finance, where closed models handle high-stakes analysis and open models power accessible customer-facing tools.","The paper's 'complementary roles' framing suggests a resourcing implication: institutions may reasonably support both open and closed routes rather than treating open source as a universal substitute."],"forward_implications":["If the division of labor is real, closed models will remain the default tools for radiology and other high-accuracy diagnostic tasks as long as they lead in reasoning ability.","Open models will continue to proliferate in niche healthcare areas because released weights make fine-tuning cheap, which was the driver of their exponential growth since LLaMA's release.","The complementary pattern points toward hybrid pipelines that pair a closed model's diagnostic reasoning with open models customized for patient-facing communication.","Research attention is now concentrated differently: closed models draw more overall scientific interest, especially in fields that use rather than build models, while open models dominate in machine-learning subfields that optimize training and adaptation."],"supporting_citations":[{"why":"Supplies the ecosystem graph data that classifies the 201 foundation models into open and closed access types.","marker":"[3]"},{"why":"Provides the BERTopic method used to derive the topic distributions showing complementary healthcare roles.","marker":"[8]"},{"why":"Is the closed-model exemplar whose capabilities anchor the high-performance diagnostic claim.","marker":"[9]"},{"why":"Is the open LLaMA model whose release catalyzed the open-model surge and fine-tuning ecosystem.","marker":"[15]"},{"why":"Marks the point where the most capable models began releasing weights only through APIs, establishing the closed-model lineage.","marker":"[4]"},{"why":"Is the open LLaMA 2 model used as the representative open-source LLM in the background.","marker":"[1]"}],"fun_headline_variants":["Closed LLMs for imaging, open LLMs for mental health","Healthcare AI splits: closed for scans, open for therapy","Open and closed LLMs pick complementary medical roles","Closed wins diagnostics, open wins personalized care","Complementary AI: closed for imaging, open for dialogue"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis treats a model being named in an arXiv paper's title or abstract as evidence that the model is actually used and valued in that healthcare area; if researchers mention closed models mainly because they are well-known, or report clinical work outside preprint servers, the claimed division of labor is not established.","fun_headline_variants_meta":{"raw":{"variants":["Closed LLMs for imaging, open LLMs for mental health","Healthcare AI splits: closed for scans, open for therapy","Open and closed LLMs pick complementary medical roles","Closed wins diagnostics, open wins personalized care","Complementary AI: closed for imaging, open for dialogue"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1150,"prompt_tokens":769,"completion_tokens":381,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":385,"completion_tokens_details":{"reasoning_tokens":304}},"tokens_in":385,"tokens_out":381,"duration_ms":3862,"temperature":1.0,"reasoning_tokens":304,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:31:22.558834+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A benchmark study that compared open- and closed-weight models on a standard radiology diagnostic suite and found open models matching or exceeding closed models on accuracy, or a survey of deployed clinical systems showing open models widely used in imaging, would undermine the paper's central complementarity claim.","supporting_citations":[{"cited_title":"Gpt-4 technical report","cited_arxiv_id":null,"evidence_quote":"Is the closed-model exemplar whose capabilities anchor the high-performance diagnostic claim."}],"review_version":1}