{"id":"7504db6f-57bf-4014-8a2c-7f32c3e38186","arxiv_id":"1908.10214","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MeSH hierarchy link changes show preferential attachment by node out-degree and descendant count, anti-preference by ancestor count, and non-trivial preferences for target properties.","lead":"This paper tracks yearly changes in the MeSH medical term hierarchy and measures whether new or deleted links favor nodes with particular properties. It finds that links tend to attach to nodes with many children and many descendants, while avoiding nodes with many ancestors, and that rewiring and deletion are as common as growth.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Binomial error model in Eq. (2) assumes independent events; MeSH updates are curator batches, and Table I shows years with massive coordinated restructuring, so the error bands around W_rand are too narrow and preference classifications may be false positives.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the binomial error model in Eq. (2) assumes independent Bernoulli trials, but MeSH updates are coordinated curator batches. This concern is supported directly by the paper's Table I, which lists years (notably G 2008 and N 2008) with hundreds of simultaneous link/node changes, far beyond what independent random events would produce. If the error bands are too narrow, the qualitative classifications—especially the weak-preference and weak-anti-preference labels that feed into the aggregated Table 3—could be false positives, undermining the central claim of non-uniform preferences. A block-bootstrap or outlier-exclusion test would settle whether the strong-preference findings survive. Because the reader's verdict already conditions acceptance on addressing this statistical reliability issue, and because the concern is addressable without breaking the paper's empirical contribution, the appropriate verdict remains CONDITIONAL. No change to the reader's verdict is needed.","tokens_in":20449,"tokens_out":5142,"duration_ms":58147,"concrete_test":"Perform a block bootstrap over the 15 yearly time steps: resample years with replacement, recompute W_emp(x) and the null W_rand(x) for each resample, and construct 95% confidence intervals for the difference W_emp(x) - W_rand(x). Repeat the analysis excluding the 2008 G and N restructuring years (and any other year with more than 20% link turnover). If the strong-preference cells (s+/s-) in Tables 2-3 no longer exclude zero at a family-wise error rate of 0.05, the independence assumption is load-bearing and the error model must be replaced; if they remain significant, the central claim is robust to the batch-update concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that new links preferentially attach to source nodes with many children and that detachment preferentially hits nodes with many children/descendants, while ancestor count shows anti-preference. Every classification in Tables 2-3 rests on comparing W_emp(x) to the null W_rand(x) plus or minus sigma from Eq. (10), which assumes each link change is an independent Bernoulli trial with fixed probability per year. MeSH is curated: the NCBI applies annual updates in coordinated batches, so events within a year are not independent. The paper's own Table I shows extreme batch years: hierarchy G in 2008 has 632 deleted nodes, 1059 deleted links, and 345 new links from new to old nodes; hierarchy N in 2008 adds 254 nodes and 244 new links between new nodes. Such coordinated restructuring violates the independence assumption, making sigma too small and producing false-positive 'preference' or 'anti-preference' labels. The effect is strongest for weak labels (w+, w-), but even some s+/s- cells could be affected if a single batch dominates the aggregate. The authors do not test for overdispersion, cluster by year, or exclude outlier restructuring years. Since the aggregated Table 3 averages these labels, the load-bearing statistical foundation of the paper's empirical conclusions is the validity of Eq. (2)-(10) under curated batch updates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the temporal evolution of the hierarchical networks formed by PubMed MeSH terms. Using yearly snapshots of 16 hierarchies (and 7 hierarchies with more than 1000 nodes during the whole period), the authors classify link changes into five types: additions involving old/new sources and targets, and deletions between old nodes. For each change type and for four node properties (number of children, number of parents, total descendants, total ancestors), they compare the observed complementary cumulative distribution of selected nodes against a random null model, computing W_emp(x) and comparing it to W_rand(x) plus or minus a binomial standard deviation. The central finding is that attachment events preferentially select source nodes with many children and many descendants, while the number of ancestors of the source node shows anti-preference across essentially all link-change types; deletion events also preferentially strike nodes with many children and descendants. The authors aggregate per-hierarchy classifications into a summary table and argue that the observed preferences are consistent across hierarchies.","tokens_in":20616,"tokens_out":4315,"duration_ms":49267,"significance":"If the statistical results hold, the paper provides a useful empirical characterization of how a large curated hierarchical ontology evolves, with concrete evidence that restructuring is not uniform but is biased by topological and hierarchical node properties. The analysis uses publicly available data, the methodology is transparent, and the authors validate the W(x) framework on simulated attachment events, which are strengths. The claimed preferences are potentially relevant for modeling the evolution of hierarchical systems beyond MeSH. However, the significance is conditional on the validity of the independence assumptions underlying the error bars and on the reproducibility of the qualitative classification scheme, both of which need strengthening before the empirical claims can be fully trusted.","major_comments":[{"comment":"The binomial error model in Eq. (2) treats each attachment or deletion event as an independent Bernoulli trial with a fixed probability u(x), and Eq. (10) sums variances over years. However, MeSH updates are coordinated annual curation batches, not independent per-link decisions. Supporting Information Table I shows hierarchy G in 2008 undergoing 632 node deletions, 1059 link deletions, and 345 old-to-new link additions in a single year, and hierarchy N in 2008 adding 254 nodes and 244 new-to-new links. Such coordinated restructuring induces positive correlations among events, so the variance around W_rand is underestimated and the classification labels, especially w+ and w- but possibly some s+/s- cells dominated by one batch, may be false positives. The authors should test for overdispersion, use a year-clustered bootstrap, or repeat the analysis excluding the major restructuring years.","section":"Data and methods, Eqs. (2), (10); Supporting Information Table I"},{"comment":"The distinction between 'strong' and 'weak' preference is not reproducible: the categories are defined by whether W_emp(x) exceeds W_rand(x) + sigma(W_rand(x)) by 'a large amount' or 'a small amount', with no numerical threshold. Table 2 and the aggregated Table 3 therefore depend on an unreported judgment call. A quantitative rule (for example, W_emp above W_rand + k sigma for a stated constant k, or a formal test statistic applied uniformly to all cells) is needed.","section":"Results, category definitions"},{"comment":"The analysis classifies on the order of hundreds of cells (7 hierarchies, 5 link-change types, and 8 property-by-endpoint combinations per hierarchy), yet no correction for multiple testing is applied. Under the null hypothesis, a substantial number of w+ and w- labels would be expected to appear by chance. The authors should report adjusted significance levels or quantify the expected number of false classifications.","section":"Tables 2, S8-S14, and Table 3"},{"comment":"The aggregated Table 3 averages labels with weights s+=1, w+=p+=0.5, s0=0, w-=p-=-0.5, s-=-1, irrespective of the number of events or the statistical power underlying each label. A strong label from a hierarchy with few events contributes the same weight as one from hierarchy D with many events, and cells with more than three i.s. entries are simply marked i.s., potentially hiding genuine signals in smaller hierarchies. Weighting by event counts or by a confidence measure would make the aggregation more defensible.","section":"Table 3 and aggregation method"}],"minor_comments":[{"comment":"The term 'MeSH' is misspelled as 'MesH' in several places, including the abstract and the introduction.","section":"Abstract and throughout"},{"comment":"The sentence 'the total number of descendants of the individual roots ... varies roughly between a 1,00 and a 10,000 nodes' contains a typo and should read 'between 100 and 10,000 nodes'.","section":"Data and methods, basic properties"},{"comment":"The phrase 'the likelihood for nodes to take part in restructuring events can be effected by their properties' should use 'affected' rather than 'effected'.","section":"Discussion"},{"comment":"The column headers for the link-change types are difficult to parse because of repeated 'source:' and 'target:' lines; clearer labels such as 'add new->new', 'add new->old', 'add old->new', 'add old->old', and 'delete old->old' would improve readability.","section":"Tables 2 and 3"},{"comment":"The text refers to colors 'orange' and 'blue' for the curves; if the journal does not guarantee color printing, the figure should also use distinguishable line styles or markers.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about batch dependencies in the MeSH update process is legitimate and should be the primary focus of the revision; the paper's empirical claims are otherwise interesting and within the scope of the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is worth reading: it extends the usual preferential-attachment analysis of MeSH growth to deletion and rewiring, and it finds consistent empirical patterns across seven hierarchies. Second, its error bars are more optimistic than the data allow, because yearly MeSH updates are coordinated batches, not independent draws. That problem is real but it does not sink the main conclusions.\n\nThe genuinely new content is the systematic study of detachment and rewiring. Previous MeSH work focused on how new terms attach under old ones; here the authors treat link deletion and old-to-old rewiring as first-class events and show they happen at the same scale as growth. The method is clearly specified and validated on simulated attachment events, and the central preferences — strong preference for high child count at the source, strong anti-preference for high ancestor count at the source — appear in multiple hierarchies and in the aggregated table. That consistency, plus the fact that the strongest curves exceed the neutral expectation by an order of magnitude, gives me confidence the qualitative directions are right.\n\nThe soft spot is statistical, not conceptual. Equations (2)–(10) model each yearly event as an independent Bernoulli trial. MeSH is curated by NCBI staff in large annual batches, and Table I shows the footprint: hierarchy G in 2008 deletes 1059 links and adds 345 new-source-to-old-target links in a single year; hierarchy N adds 244 new-to-new links in 2008. Positive correlation within a batch makes the variance in Eq. (10) too small, so some weak-preference labels, and conceivably a marginal strong label, could be false positives. The authors do not cluster by year, test for overdispersion, or report a sensitivity analysis with restructuring years excluded. The strong effects are large enough to survive, but the precise magnitudes and the weak-effect cells should not be trusted as reported.\n\nTwo smaller issues: the s+/w+/s0 classification is qualitative — no numerical threshold is given for \"large\" versus \"small\" excess — and there is no multiple-comparison correction across the hundreds of cells in Tables 2 and S8–S14. These are minor, not fatal.\n\nWho gets value: network scientists interested in ontology evolution or in validating hierarchy-growth models, and anyone building predictive models of MeSH changes. A serious referee should engage with it; the batch-independence concern is addressable with clustering or year-level bootstrapping, and the paper would be stronger for it. My recommendation is conditional acceptance: ask for the independence issue to be handled explicitly and the weak-effect classifications to be flagged as exploratory.\n\nI would bring this to the reading group and would probably cite it, though with a caveat about the error bars.","headline":"A genuinely useful empirical study of how MeSH hierarchies grow and rewire, with a real but addressable statistical problem: the error bars assume independence across events that actually arrive in curator batches.","tokens_in":21211,"tokens_out":2351,"would_cite":true,"duration_ms":28238,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["89.75.Hc","89.75.Fb"],"model":"deepseek-v4-flash","headline":"New MeSH links attach by preference, not chance; child-rich nodes are favoured while high-ancestor nodes are avoided.","keywords":["MeSH hierarchies","preferential attachment","network evolution","directed acyclic graphs","link deletion","rewiring","PubMed","taxonomy restructuring"],"falsifier":"Run the same preference analysis with a null model that reshuffles which existing nodes are edited within each yearly update while preserving the number and type of edits, and check whether the strong preference for child-rich sources and the anti-preference for ancestor-rich sources remain outside the widened confidence intervals; a negative answer would overturn the central claim.","tokens_in":20194,"feed_emoji":"🕸️","tokens_out":6988,"duration_ms":67727,"temperature":0.7,"pith_summary":"This paper uses the yearly public updates of the 16 MeSH hierarchies behind PubMed to ask whether the addition and deletion of links in a hierarchical network is random or biased. It finds that new links pointing from existing terms to newly introduced terms select source nodes with many children at a rate far above uniform chance, and that link deletion also favours nodes with many children and many descendants. At the same time, the total number of ancestors of a source node is avoided in essentially every type of link change, so shallow terms near the root are under-used as origins of restructuring. The authors conclude that rewiring and deletion matter as much as growth, and that hierarchy evolution mixes preferential and anti-preferential attachment over several topological properties at once.","feed_headline":"New MeSH links attach by preference, not chance","feed_subtitle":"A 14-year look at PubMed's subject headings shows rewiring and deletion matter as much as growth.","key_machinery":"The carrying device is the ratio $W(x)=w(x)/Q(x)$, in which $Q(x)$ is the complementary cumulative distribution of a node property $x$ (number of children, parents, descendants, or ancestors) among available nodes, and $w(x)$ is the number of actually chosen nodes whose property is at least $x$. Under uniform random choice $W(x)$ is flat, an increasing $W(x)$ signals preference for large $x$, and a decreasing $W(x)$ signals anti-preference; the expected value and standard deviation of $W(x)$ are derived from a binomial model (Eq. 2) and used as error bands around the neutral value $W_{\\mathrm{rand}}(x)$. For deletions the null model is selecting a uniformly random link, implemented by re-weighting $Q(x)$ by node degree, so that hubs are not mistakenly counted as preferred.","core_discovery":"The central discovery is that the growth and restructuring of the MeSH hierarchies are not uniformly random. When a new link is added from an old node to a new term, source nodes with more children are chosen with significantly higher probability than uniform random selection would give; deletion events likewise strike nodes with many children and many descendants. Conversely, the total number of ancestors of the source node displays anti-preference across nearly all change types, meaning broad, shallow terms are less likely than chance to be the origin of a rewiring or a new link. Properties of the target node have a smaller influence than properties of the source node, and across the seven largest hierarchies the same combination of change type and property never shows preference in one hierarchy and anti-preference in another. The authors present these patterns as evidence that taxonomy evolution is shaped by an interplay of multiple non-uniform, hierarchy-specific attachment rules rather than by a single preferential-attachment law.","pith_inferences":["The observed anti-preference for ancestor count may reflect curator intent to place new terms under the most specifically relevant existing parent, so the pattern could be an emergent signature of expertise-driven classification rather than a purely structural law.","If year-by-year edits are applied in coordinated batches, the independent-Bernoulli error bands in Eq. (7) are too narrow; a year-level permutation test would tell which reported preferences are robust to batch structure.","The same $W(x)$ statistic can be exported to other curated hierarchies such as gene ontology or Wikipedia category trees, offering a direct test of whether organised knowledge systems share this preference pattern.","A generative model linking parent choice to a power of child count multiplied by a decreasing function of depth could reproduce the joint pattern; fitting that function would turn the present qualitative findings into a quantitative evolution rule."],"forward_implications":["Deletion and rewiring between old nodes occur at the same magnitude as links to new nodes, so realistic models of hierarchy evolution must treat restructuring as a first-order process, not a perturbation.","A single preferential-attachment rule cannot explain the data; predicting where the next change lands requires simultaneous preferences over out-degree, descendant count, and ancestor count.","The anti-preference for high ancestor counts implies that reorganisation concentrates at intermediate depth, so future growth models should include a depth penalty for parent choice.","Because no hierarchy displays the opposite preference on the same cell, the qualitative pattern is generalisable across the largest MeSH hierarchies and plausibly to curated hierarchies elsewhere.","Source-node properties dominate target-node properties, so early indicators of upcoming rewiring should be measured on the parent side of a link."],"supporting_citations":[{"why":"Supplies the method of comparing the distribution of a property among chosen nodes to the distribution among available nodes to detect preference.","marker":"[23]"},{"why":"Provides the preferential attachment rule that the paper's growth results are compared to as the established baseline model.","marker":"[43]"},{"why":"Empirical method for measuring preferential attachment in growing networks that the paper parallels for hierarchical link addition.","marker":"[45]"},{"why":"Further empirical evidence of preferential attachment that justifies the comparison baseline.","marker":"[46]"},{"why":"An earlier empirical detection of preferential attachment in collaboration networks used for comparison.","marker":"[44]"},{"why":"Prior work on MeSH classification and growth that the paper extends by showing restructuring is equally important.","marker":"[49]"}],"fun_headline_variants":["MeSH hierarchies evolve by non-random rewiring","PubMed taxonomy growth is far from random","Link preference shapes MeSH network evolution","MeSH link changes show clear topological bias","Hierarchy evolution follows preference, not chance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The binomial error model in Eq. (2) treats every link addition and deletion as an independent Bernoulli trial with a fixed probability $u(x)$, so if MeSH updates are applied as coordinated batches by curators, the error bands around $W_{\\mathrm{rand}}$ are too narrow and some apparent preferences could be false positives.","fun_headline_variants_meta":{"raw":{"variants":["MeSH hierarchies evolve by non-random rewiring","PubMed taxonomy growth is far from random","Link preference shapes MeSH network evolution","MeSH link changes show clear topological bias","Hierarchy evolution follows preference, not chance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":1229,"prompt_tokens":955,"completion_tokens":274,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":209}},"tokens_in":571,"tokens_out":274,"duration_ms":3679,"temperature":1.0,"reasoning_tokens":209,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:48:51.341671+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same preference analysis with a null model that reshuffles which existing nodes are edited within each yearly update while preserving the number and type of edits, and check whether the strong preference for child-rich sources and the anti-preference for ancestor-rich sources remain outside the widened confidence intervals; a negative answer would overturn the central claim.","supporting_citations":[{"cited_title":"Preferential attachment of communities: The same principle, but a higher level","cited_arxiv_id":null,"evidence_quote":"Supplies the method of comparing the distribution of a property among chosen nodes to the distribution among available nodes to detect preference."},{"cited_title":"Emergence of scaling in random networks","cited_arxiv_id":null,"evidence_quote":"Provides the preferential attachment rule that the paper's growth results are compared to as the established baseline model."},{"cited_title":"Measuring preferential attachment in evolving networks","cited_arxiv_id":null,"evidence_quote":"Empirical method for measuring preferential attachment in growing networks that the paper parallels for hierarchical link addition."},{"cited_title":"Clustering and preferential attachment in growing networks","cited_arxiv_id":null,"evidence_quote":"Further empirical evidence of preferential attachment that justifies the comparison baseline."},{"cited_title":"Evolution of the social network of scientiﬁc collaborations","cited_arxiv_id":null,"evidence_quote":"An earlier empirical detection of preferential attachment in collaboration networks used for comparison."},{"cited_title":"A Maximum-Entropy approach for accurate document annotation in the biomedical domain","cited_arxiv_id":null,"evidence_quote":"Prior work on MeSH classification and growth that the paper extends by showing restructuring is equally important."}],"review_version":1}