{"id":"48138a6b-363e-4af5-975d-4bcc28ca9a54","arxiv_id":"2608.04171","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":11,"one_line_summary":"MHD papers have longer abstracts than PIC and gyrokinetic papers in arXiv plasma physics, and PIC papers dominate GPU mentions, based on regressions over 5,522 papers.","lead":"This paper analyzes 5,522 computational plasma physics papers from arXiv (2010 to 2025) and reports that MHD papers have longer abstracts than PIC or gyrokinetic papers, while PIC papers are much more likely to mention GPUs. The authors use abstract word count as a proxy for methodological complexity and apply regression models to link method, team size, and GPU mentions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Tobit 'robustness' results report Cox-model hazard ratios with signs opposite to the OLS estimates, so the paper's team-size and temporal claims are internally contradicted.","rationale":"I read the paper as a descriptive bibliometric study whose central empirical content is: (i) MHD abstracts are longer than PIC/gyrokinetic; (ii) abstract length increases with team size; (iii) PIC/gyrokinetic prevalence and GPU mentions rise over time. Claim (iii) is supported by the logistic regression and descriptive figures, and claim (i) appears in both OLS and Tobit narratives, so those are less fragile. Claim (ii) is the one that collapses: the OLS section says additional authors increase length, while the 'Tobit' section reports a hazard ratio below 1 and states latent length decreases with author count. Since a Tobit model of a word count does not output hazard ratios, the reported numbers cannot originate from Eq. 2–3; they are almost certainly from a survival model or a copy-paste error. This is not a matter of proxy validity or external consensus; it is an internal inconsistency in the paper's own exhibits. The reader's stated weakest assumption (abstract length as complexity proxy) is a legitimate concern, but it is secondary: even if the proxy were valid, the team-size effect is not robust across the paper's own specifications. My proposed check—refitting the Tobit and checking the censoring fraction—would settle whether the sign flip is real or an artifact. If the check confirms the OLS sign, the paper must be revised to remove or correct the Tobit section; if it confirms the Tobit sign, the abstract must be revised. Either way, the present version should not be accepted; the reader's REJECT is appropriate, so I recommend no change to the verdict.","tokens_in":7283,"tokens_out":4411,"duration_ms":42072,"concrete_test":"Obtain the released dataset and scripts (or reconstruct the 5,522-paper dataset via the arXiv API with the §2.1 filter) and refit Eq. 2–3 with a standard Tobit estimator (e.g., statsmodels 'Tobit' or R 'AER::tobit') on abstract_len_cens, recording the n_authors and year coefficients. Then compare to Fig. 6. If a correct Tobit yields positive author and year coefficients close to OLS, the published negative hazard ratios are from a different model and the robustness check is invalid. If it yields negative coefficients, the abstract's 'abstract length increases modestly with team size' is contradicted by the paper's own preferred censored specification. Also report the fraction of abstracts censored at 500 words; if that fraction is small, any large sign flip cannot be attributed to censoring.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that abstract length increases with team size and that temporal trends are identifiable rests on the OLS and Tobit regressions agreeing. They do not. In §3.1, Fig. 6 is described as a Tobit fit of Eq. 2–3, but it reports a 'hazard ratio' (HR=0.9907 for n_authors, logHR=-0.0189 for year). Tobit models do not produce hazard ratios; those outputs come from a survival/Cox model. Even taken at face value, the reported effects have the opposite sign from the OLS results: OLS says each additional author increases abstract length and lengths rise over time, while the Tobit section says additional authors decrease latent length and years decrease latent length. The Discussion (§5) even concedes this split ('simple regression suggested a slight increase... models accounting for censored abstracts indicated a shift toward more concise abstracts'). Because the censoring cap at 500 words should barely bind for most arXiv abstracts, a correct Tobit should approximate OLS; a sign flip plus hazard-ratio language indicates a different or mis-specified model was run. Therefore the claimed robustness check does not support the headline team-size result, and the central claim that abstract length increases modestly with team size is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a bibliometric analysis of 5,522 arXiv plasma physics papers published between 2010 and 2025, using abstract length as a proxy for methodological and algorithmic complexity. It applies OLS, Tobit, and logistic regressions to estimate how computational method category, number of authors, and publication year relate to abstract length and to GPU mentions. The central claims are that MHD papers have longer abstracts than PIC and gyrokinetic papers, that abstract length increases modestly with team size, that temporal trends in abstract length are detectable, and that PIC and gyrokinetic methods have grown in prevalence and are more strongly associated with GPU adoption.","tokens_in":7556,"tokens_out":3832,"duration_ms":42974,"significance":"If the quantitative claims were supported, the paper would provide a useful field-level descriptive contribution on how research communication and HPC adoption vary across computational methods in plasma physics. The paper's strengths include a large curated corpus, explicit inclusion and exclusion criteria, complementary regression models, and clear visualizations of temporal trends. However, the central team-size and temporal-length claims are currently undermined by an internal inconsistency between the OLS and Tobit analyses, and the proxy interpretation linking abstract length to methodological complexity is asserted rather than validated. The paper therefore has the ingredients of a useful empirical study, but the load-bearing statistical results and their interpretation need substantial reworking.","major_comments":[{"comment":"The Tobit results are reported with hazard ratios (HR = 0.9907 for n_authors and logHR = -0.0189 for year), but a Tobit model does not produce hazard ratios; these quantities come from a survival or Cox proportional hazards model. Moreover, the reported Tobit effects have opposite signs from the OLS results: OLS says each additional author increases abstract length and lengths rise over time, while the Tobit text says additional authors and later years decrease the latent abstract length. Because the 500-word cap should bind for only a small fraction of arXiv abstracts, a correctly specified Tobit should approximate the OLS estimates, not reverse their signs. This internal contradiction means the paper's robustness claim is invalid and the headline finding that abstract length increases modestly with team size is not established.","section":"§3.1, Fig. 6, Eqs. (2)–(3)"},{"comment":"The 500-word cap is introduced as a censoring mechanism, but arXiv abstracts are not naturally censored at 500 words; the authors artificially cap the dependent variable and then apply a Tobit model. A Tobit for this constructed variable estimates the effect of the authors' own transformation, not an underlying censorship mechanism. The text also uses 'censored' and 'truncated' interchangeably, which are distinct concepts. The authors should either justify that arXiv enforces a real upper limit that binds in their data, or remove the Tobit and present the OLS and logistic results as the primary analyses.","section":"§2.2, §2.3, Eqs. (2)–(3)"},{"comment":"The paper assumes without validation that 'abstract length is a proxy for methodological and algorithmic complexity.' The abstract and Discussion then interpret longer MHD abstracts as indicating 'more extensive methodological and physical exposition.' This is a labeling tautology: the conclusion restates the proxy definition rather than providing evidence for a link between length and complexity. To support the interpretation, the authors should validate the proxy against an external measure, such as code size, number of equations, full-text methods-section length, or expert ratings, or alternatively limit the conclusions to abstract verbosity itself.","section":"§1, §2.3, Abstract"},{"comment":"The method-category assignment is underspecified. The filtering criteria list inclusive and exclusive keywords, but no rule is given for assigning each paper to exactly one of PIC, gyrokinetic, MHD, or other. Table 1 already shows ambiguous examples: a paper on the Boltzmann equation and current density is classified as 'Other,' and a paper on Rayleigh–Taylor turbulent mixing is classified as 'PIC.' Because every method comparison in the paper depends on this classification, the assignment pipeline must be described in sufficient detail, or the data and code should be made available so that the classifications can be checked.","section":"§2.1, Table 1"}],"minor_comments":[{"comment":"The title reads 'A Empirical Analysis' and should be corrected to 'An Empirical Analysis.'","section":"Title"},{"comment":"The derived variable 'Abstract length (number of words)' should specify whether the count includes the title, author list, or only the abstract body, and whether standard tokenization rules were applied.","section":"§2.2"},{"comment":"The OLS model is reported with R² = 0.043 and p < 0.0001, but no F-statistic, intercept value, or confidence intervals are given; these should be included for reproducibility and to help readers assess effect sizes beyond the small R².","section":"§3.1, Fig. 5"},{"comment":"The logistic regression results report the PIC coefficient and the year coefficient, but omit the intercept, the gyrokinetic coefficient, and the baseline MHD log-odds; complete coefficient tables should be provided.","section":"§3.1, Fig. 7"},{"comment":"The Discussion concedes that 'simple regression suggested a slight increase in abstract lengths over time, models accounting for censored abstracts indicated a shift toward more concise abstracts,' which directly reproduces the contradiction noted in the major comments; this conflict must be resolved before the paper can be considered internally consistent.","section":"§5"},{"comment":"No data or code availability statement is included, even though the paper relies on a curated CSV dataset and custom filtering and regression code; for an empirical bibliometric study, sharing these artifacts is recommended.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The topic is timely and the dataset is a potentially useful resource, but the regression section as written is not internally coherent: the Tobit model appears to have been replaced by a survival model, and the resulting signs contradict the OLS findings. I would be willing to consider a revised version that either removes the invalid Tobit or re-estimates it correctly, validates or explicitly weakens the complexity proxy, and makes the classification pipeline reproducible. If the authors cannot provide corrected analyses, the central claims will not be supportable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper compiles a new filtered corpus of 5,522 computational plasma physics papers and documents method prevalence and GPU mentions over time. That part is useful and, as far as I can tell, not done elsewhere. The finding that MHD abstracts run longer than PIC and gyrokinetic ones is consistent across the OLS and Tobit outputs and is likely real. The GPU adoption story (PIC leading, rising over time) is also plausible and matches the field's direction.\n\nThe problem is the team-size and temporal-trend claims. The OLS model in Fig. 5 says each author adds words and abstracts get longer over time. The Tobit model in Fig. 6 reports hazard ratios (HR=0.9907 for authors, logHR=-0.0189 for year), which are not Tobit outputs, and has signs opposite to OLS: additional authors and later years are associated with shorter latent length. The paper's own Discussion admits the split. Since the 500-word cap almost never binds for arXiv abstracts, a correct Tobit should approximate OLS; a sign flip means the model is either misspecified or the figure is from a survival model. Either way, the robustness check does not support the abstract's claim that 'abstract length increases modestly with team size.'\n\nThere's also the proxy assumption. Abstract word count is treated as 'methodological and algorithmic complexity' without any validation. That could be fine as a descriptive measure, but the paper leans on it for interpretation. The method assignment filter is only partially specified, and the dataset and code aren't released, so the numbers can't be checked. The reference list is mostly relevant case studies and econometric texts; nothing suspicious there.\n\nWhat the paper does well is the descriptive trend analysis: the corpus construction, the heatmaps, the method shares over time, and the GPU logistic results. Those are new and worth having if cleaned up.\n\nBottom line: the paper isn't ready as is. The central claim about team size is not established, and the Tobit section is a red flag. A serious editor should send it to peer review, because the corpus and descriptive findings have value, but the authors need to fix the regression analysis, release the data, and soften the proxy language. I wouldn't cite it until that happens.","headline":"Useful new corpus and descriptive trends, but the team-size and temporal claims are contradicted by the paper's own Tobit results.","tokens_in":8100,"tokens_out":2588,"would_cite":false,"duration_ms":24672,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A text-mining study of 5,522 arXiv papers reports that magnetohydrodynamics papers carry systematically longer abstracts than particle-in-cell or gyrokinetic papers, even though the latter two are the methods most associated with GPU…","keywords":["Computational Plasma Physics","Plasma Simulations","High-Performance Computing","GPU Acceleration","OLS Regression","Particle-in-Cell","Gyrokinetic","Magnetohydrodynamics"],"falsifier":"Select 100 MHD and 100 PIC papers from the same dataset, count equations and algorithm terms in the full-text methods sections, and check whether the MHD-vs-PIC ordering matches the abstract-length ordering; if the direct full-text complexity measure does not reproduce the abstract-length gap, the proxy underlying the paper's central claim is unsupported.","tokens_in":7028,"feed_emoji":"⚡","tokens_out":6734,"duration_ms":60707,"temperature":0.7,"pith_summary":"This paper tries to establish that the way computational plasma physicists write their abstracts is not random: it depends systematically on the simulation method used, the number of co-authors, and the publication year. Analyzing 5,522 arXiv papers from 2010 to 2025, it reports that magnetohydrodynamic (MHD) papers have longer abstracts than particle-in-cell (PIC) or gyrokinetic papers, and that larger teams write slightly longer abstracts. It also finds that PIC and gyrokinetic methods have become relatively more common over the past decade and are much more likely to mention GPUs, while MHD's share has declined. The paper reads this as a decoupling between methodological verbosity and methodological prevalence: the methods that are growing fastest are not the ones writing the most words.","feed_headline":"5,522 papers: MHD abstracts longest; PIC leads GPU adoption","feed_subtitle":"A 15-year scan of 5,522 arXiv papers maps method choice, team size, and GPU mentions across plasma simulations.","key_machinery":"The analytical engine is a set of three regression models applied to a curated arXiv-derived dataset: ordinary least squares (OLS) for abstract word count, tobit regression for abstract length censored at 500 words, and logistic regression for the binary GPU-mention indicator. Abstract length is the operational proxy for methodological and algorithmic complexity, and method category dummies (PIC, gyrokinetic, with MHD as baseline), author count, and publication year serve as the predictors. The models work by testing whether method, team size, and time explain significant variation in how papers describe themselves, with the tobit model guarding against truncation and the logistic model connecting method choice to GPU adoption.","core_discovery":"The central discovery is that abstract length, used as a proxy for methodological and algorithmic complexity, differs systematically across computational methods in plasma physics: MHD papers carry significantly longer abstracts than PIC and gyrokinetic papers after controlling for team size and publication year. Regression models (OLS, tobit, logistic) on 5,522 arXiv papers confirm the MHD verbosity advantage, and show that the number of authors has a modest positive effect on abstract length in the OLS specification, while the tobit model indicates a marginal negative effect on the latent length once censoring at 500 words is accounted for. The same models show that PIC papers are strongly associated with GPU mentions while gyrokinetic papers are not significantly associated, and that GPU mentions have increased over time. The paper concludes that methodological verbosity and method prevalence have decoupled: the most verbose method is not the one growing fastest.","pith_inferences":["The paper's own results contain an internal tension: OLS says author count increases abstract length, while the tobit model says the latent length decreases marginally with each additional author; if the tobit specification is more trustworthy, the headline team-size result may be an artifact of uncensored OLS.","Abstract length is a crude proxy; a direct test would compare abstracts from the same team or same code across methods, or measure the actual number of equations and solver details in full texts, to see whether the MHD ordering survives on direct complexity measures.","The logistic finding that gyrokinetic papers are not significantly associated with GPU mentions, despite being grouped with PIC as computationally intensive, suggests the GPU-adoption story may be PIC-specific; separating gyrokinetic GPU usage in specific codes from field-wide trends could sharpen the conclusion.","The method-categorization pipeline assigns each paper to a single category based on keywords; papers that combine methods (e.g., PIC-MHD) are forced into one bucket, which could bias the abstract-length comparisons; a multi-label extension would test whether the ordering is robust."],"forward_implications":["If abstract length tracks methodological complexity, field-wide bibliometric scans can use abstracts alone to chart how computational complexity evolves across plasma physics subfields.","The strong PIC-GPU link implies that GPU acceleration in plasma physics is being driven mainly by PIC and gyrokinetic simulations, so accelerator investments and code-porting efforts should target those methods first.","The decoupling between verbosity and prevalence means that publication-count metrics under-represent the decline of MHD as a share of computational plasma physics.","Since abstract length increases with team size, large collaborations may systematically produce more information-dense abstracts, which should be accounted for in any text-based meta-analysis of the literature.","Temporal trends in the tobit model suggest abstracts have been getting more concise in recent years once censoring is accounted for, implying that conciseness norms are shifting even as the method mix changes."],"supporting_citations":[{"why":"Supplies the tobit regression framework used to handle abstract length censored at 500 words.","marker":"[1]"},{"why":"Informs the application of tobit models in empirical research settings.","marker":"[2]"},{"why":"Supports the robustness of the logistic regression model used for the binary GPU-mention outcome.","marker":"[3]"},{"why":"Provides the logit/probit/tobit toolkit for models with dummy dependent variables.","marker":"[7]"},{"why":"Serves as an example of a logit-tobit model that the authors draw on for their empirical modeling approach.","marker":"[10]"}],"fun_headline_variants":["MHD papers wordier, PIC leads GPU surge in plasma arXiv","In 5,522 plasma papers, MHD verbose, PIC GPU-driven","Plasma arXiv: MHD longest abstracts, PIC top GPU mention","15-year scan: MHD verbosity, PIC GPU adoption decouple","Study: MHD abstracts longest, PIC accelerates on GPUs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that abstract word count is a reliable proxy for the methodological and algorithmic complexity of a plasma physics paper, and that each paper can be cleanly assigned to a single method category.","fun_headline_variants_meta":{"raw":{"variants":["MHD papers wordier, PIC leads GPU surge in plasma arXiv","In 5,522 plasma papers, MHD verbose, PIC GPU-driven","Plasma arXiv: MHD longest abstracts, PIC top GPU mention","15-year scan: MHD verbosity, PIC GPU adoption decouple","Study: MHD abstracts longest, PIC accelerates on GPUs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000665,"raw_usage":{"total_tokens":3050,"prompt_tokens":971,"completion_tokens":2079,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":1986}},"tokens_in":587,"tokens_out":2079,"duration_ms":14460,"temperature":1.0,"reasoning_tokens":1986,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:20:58.918156+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select 100 MHD and 100 PIC papers from the same dataset, count equations and algorithm terms in the full-text methods sections, and check whether the MHD-vs-PIC ordering matches the abstract-length ordering; if the direct full-text complexity measure does not reproduce the abstract-length gap, the proxy underlying the paper's central claim is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the tobit regression framework used to handle abstract length censored at 500 words."},{"cited_title":"Global Strategy Journal11(3), 331–355 (2021)","cited_arxiv_id":null,"evidence_quote":"Informs the application of tobit models in empirical research settings."},{"cited_title":"Journal of the Royal Statistical Society: Series B (Methodological)55(3), 693–706 (1993)","cited_arxiv_id":null,"evidence_quote":"Supports the robustness of the logistic regression model used for the binary GPU-mention outcome."},{"cited_title":"International Journal of Computa- tional and Experimental Science and Engineering6(1), 63–74 (2020)","cited_arxiv_id":null,"evidence_quote":"Provides the logit/probit/tobit toolkit for models with dummy dependent variables."},{"cited_title":"Research policy30(2), 245–262 (2001)","cited_arxiv_id":null,"evidence_quote":"Serves as an example of a logit-tobit model that the authors draw on for their empirical modeling approach."}],"review_version":1}