{"id":"69ac8f73-aae2-414e-8141-ab9b7ad848c4","arxiv_id":"2412.18757","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A first-year last-authorship anomaly is proposed as a warning signal for authorship disambiguation quality in bibliometric databases.","lead":"This paper finds that over 60% of biomedical researchers in the OpenAlex database appear as last authors in their first year of publishing, a pattern that is implausible for genuine career trajectories. The authors interpret this anomaly as a signal of poor name disambiguation and use it to compare data quality across metadata groups and between genders.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The anomaly signal is not yet causally tied to disambiguation errors: missing early publications can produce the same pattern, and the ORCID control group still shows ~40% anomaly, so validation against ground truth is missing.","rationale":"The reader's weakest assumption—that incomplete early-publication coverage, rather than disambiguation error, could produce the anomaly—is exactly the load-bearing concern. I add one piece of internal evidence: the ORCID group, described as a gold standard, still shows a roughly 40% anomaly rate, which is difficult to reconcile with the claim that the anomaly specifically marks poor disambiguation. The paper's empirical patterns are real and the proposed signal is plausible, but the causal attribution is not validated. Since the reader already made acceptance conditional on resolving this attribution, my read does not change the verdict. If the proposed ORCID-based test later shows the anomaly persists with complete publication histories, acceptance would be strengthened; if the censoring simulation reproduces the anomaly, the central claim would need substantial revision.","tokens_in":6467,"tokens_out":6143,"duration_ms":62601,"concrete_test":"Extract a stratified random sample of about 500 biomedical authors from OpenAlex who have linked ORCID iDs. For each, use the ORCID public API to obtain their complete works list, and determine the true year of first publication and the true year of first last-author publication. Compare these to the OpenAlex-based values, and recompute the debut-year anomaly using ORCID-confirmed publication years. Then, to stress-test the mechanism, artificially remove the first 1, 3, 5, and 10 ORCID-confirmed publications from each profile and recompute the anomaly rate. If the deletion simulation alone pushes the anomaly toward the observed 60-80% levels, incomplete coverage can produce the headline pattern without disambiguation errors. If the anomaly remains high even with full ORCID-confirmed histories, the signal is robust to this confounder.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the debut-year last-author anomaly rate is a reliable warning signal for poor authorship disambiguation. The load-bearing condition is that the anomaly is caused specifically by conflating distinct authors, and not by other reasons an author's first recorded publication can be later than their true first publication. The paper never tests this against a gold standard. It argues from contrasts: authors without affiliation or ORCID have higher anomaly rates, but those same authors are also more likely to have incomplete early-career records. Even more telling, the ORCID group—whose profiles the paper treats as a gold standard for disambiguation—still shows an anomaly rate near 40%. If ORCID profiles are well disambiguated, 40% cannot be a disambiguation-error signal; if ORCID profiles are incomplete, then incomplete coverage—not disambiguation—is a first-order cause. Either way, the paper's assertion that 'errors in authorship disambiguation highly contribute to the observed anomalies' (Results) is not established. The authors' own limitations mention last-author conventions and OpenAlex-only data, but they do not address the confounder of missing early publications.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes OpenAlex metadata for roughly 5.8 million biomedical authors and defines each author's career debut as the year of their first recorded publication. It reports that over 60% of authors have their first last-author paper in the same year as their debut, which the authors call an anomaly. The paper proposes this anomaly rate as a warning signal for poor authorship disambiguation, supports this with higher anomaly rates among authors lacking affiliation or ORCID metadata, and applies the metric to argue that disambiguation quality was lower for female than male authors before roughly 2008. The central claim is that the anomaly is caused primarily by authorship name disambiguation errors.","tokens_in":6699,"tokens_out":4600,"duration_ms":38933,"significance":"If the causal interpretation were established, the anomaly metric would be a cheap, scalable diagnostic for disambiguation quality in large bibliometric databases, and the gender-specific application would have direct implications for reinterpreting pre-2010 gender disparity studies. The paper's use of open data at very large scale is a strength, and the proposed metric is simple enough to be applied across databases. However, the current evidence is correlational rather than causal: the anomaly is never validated against a manually curated gold standard, and the ORCID group still shows an anomaly rate near 40%, which is difficult to reconcile with the claim that the anomaly mainly reflects disambiguation errors. The framework is promising, but the core attribution needs substantially stronger support before the metric can be called reliable.","major_comments":[{"comment":"The central causal claim is not established because the anomaly measure is never validated against a curated gold standard of disambiguated author profiles. The career debut year is defined as the year of the earliest publication in the database; if an author's early publications are missing, the recorded debut shifts later and can coincide with a genuine first last-author paper, inflating the anomaly without any disambiguation failure. This alternative mechanism is especially relevant for the authors without affiliation or ORCID, who are also more likely to have incomplete early-career records, so the associations in Fig. 2 are confounded. The Discussion lists other limitations but does not address incomplete early publication coverage.","section":"Results, 'Poor authorship disambiguation is the cause of the anomaly'"},{"comment":"The ORCID comparison undermines the paper's interpretation rather than supporting it. Authors with ORCID are treated as a gold standard for disambiguation, yet this group still shows an anomaly rate near 40%. If ORCID profiles are well disambiguated, 40% cannot be a disambiguation-error signal; if ORCID profiles are incomplete, then incomplete coverage, not disambiguation, is a first-order cause. Either way, the statement that 'errors in authorship disambiguation highly contribute to the observed anomalies' (Results) is not established.","section":"Fig. 2b"},{"comment":"The claim that female authors exhibit 'minor yet statistically significant' discrepancies is unsupported by any reported statistical test. Materials and Methods describes the gender prediction procedure and thresholds but does not describe a hypothesis test, effect size, or confidence interval for the comparisons in Fig. 3. Given the small differences and the multiple subgroup, continent, and period comparisons, the statistical evidence for the gender claim needs to be reported explicitly.","section":"Results, gender analysis (Fig. 3)"}],"minor_comments":[{"comment":"There is a typo: 'authorshup disambiguation' should be 'authorship disambiguation'.","section":"Introduction, final paragraph"},{"comment":"The caption contains 'stiking anomaly'; this should be 'striking anomaly'.","section":"Fig. 1 caption"},{"comment":"The phrase 'last author positon' should read 'last author position'.","section":"Results, first paragraph"},{"comment":"The title 'Discrepancy in anomaly occured between male and female authors before 2008' has a grammatical error; 'occured' should be 'occurred'.","section":"Fig. 3 title"},{"comment":"The terms 'affiliation information' and 'country metadata' are used interchangeably (Fig. 2 and text); the paper should define what exactly is measured, since country and institution are not the same.","section":"Materials and Methods, 'Authorship Metadata'"},{"comment":"No data or code availability statement is provided; for a methods-oriented claim, releasing the analysis scripts would substantially improve reproducibility.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The gender-disparity finding is likely to attract attention beyond the specialist readership, so I would encourage the editor to require a validation step against a gold-standard author set or an appropriate sensitivity analysis before publication. The paper may also be better framed as a descriptive anomaly report rather than a validated disambiguation-quality metric. Additionally, the absence of a code/data availability statement is a reproducibility concern for a cs.DL methods paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a genuinely new diagnostic, the empirical pattern is real, and the paper is worth taking seriously. But the central claim — that the anomaly rate is a warning signal for poor authorship disambiguation — is not yet established, because the ORCID control group still shows ~40% of authors with an anomaly, and the alternative explanation of missing early publications is never tested.\n\nWhat's new: using last-author analysis (van Dijk et al. 2014) to look at when researchers first appear as last authors, the authors find that over 60% of 5.8 million biomedical researchers in OpenAlex had their first last-author paper in their debut year. That is a striking pattern. They show it is more common among authors without affiliation metadata (above 80%) and without ORCID (about 70%) compared with those who have such identifiers (around 60% and 40%, respectively). They also report a small but persistent pre-2010 gender gap, with female authors more likely to show the anomaly. The work is large-scale, uses a public snapshot, and the authors are upfront about several limitations (last-author conventions, OpenAlex-only, name-based gender inference).\n\nThe soft spots are real. The paper interprets the anomaly as disambiguation error, but the contrary explanation is equally plausible: a researcher's first recorded paper in the database may not be their true first paper. If early publications are missing, their recorded debut shifts later, and it can coincide with a genuine first last-author paper, inflating the anomaly without any conflation of distinct authors. The ORCID comparison cuts both ways: if ORCID is a good gold standard, then 40% of well-disambiguated authors still show the anomaly, so the signal is not specific to disambiguation. If ORCID profiles are incomplete, then coverage gaps — not name conflation — are a first-order cause. The section header \"Poor authorship disambiguation is the cause of the anomaly\" is more confident than the evidence supports, since the results are correlations with metadata proxies, not a causal demonstration.\n\nThat said, the descriptive finding is valuable, and the proposed signal could be useful if validated. The fix is not beyond reach: compare a sample against manually curated author profiles, or test the missing-publications confounder directly (for instance, using authors whose early records are complete via ORCID or Scopus). Given that the paper openly acknowledges some limitations but misses this central confounder, my recommendation is to send it to peer review, but with a strong request either to validate the signal against ground truth or to soften the causal language and present it as a hypothesis. The right audience is bibliometricians and any OpenAlex user concerned about data quality. It deserves serious refereeing.","headline":"A genuinely new anomaly signal for disambiguation quality, but the causal claim is overconfident and the ORCID comparison undermines it; worth refereeing with a required validation step.","tokens_in":7163,"tokens_out":2868,"would_cite":true,"duration_ms":24746,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the fraction of researchers whose first last-author paper appears in their debut year can serve as a reliable warning signal for poor authorship disambiguation, and demonstrates the signal on 5.8 million biomedical…","keywords":["authorship disambiguation","last author analysis","career independence","debut-year anomaly","gender disparities","bibliometric databases","researcher identifiers","anomaly detection"],"falsifier":"Build a curated set of biomedical author profiles whose true publication histories are known; if the true histories show the same roughly 60% debut-year last-author rate, then the anomaly is not a disambiguation signal. More directly, check whether authors flagged as anomalous have their earliest recorded papers missing from the database: if many anomalous authors have no publications before the first last-author paper despite known earlier work, the signal tracks coverage gaps rather than name conflation.","tokens_in":6293,"feed_emoji":"📊","tokens_out":8146,"duration_ms":71471,"temperature":0.7,"pith_summary":"The paper tries to establish that a single, easy-to-compute number—the share of researchers whose first last-author paper falls in the same year as their first publication—can flag poor authorship disambiguation in large bibliographic databases. Using the publication histories of 5.8 million biomedical researchers drawn from a major open bibliometric database, it finds that roughly 62% of authors show this \"immediate independence\" anomaly, a pattern the authors argue is far too common to reflect real careers. The anomaly is higher among authors lacking affiliation data and lower among authors with a persistent self-reported researcher identifier, matching what would be expected if disambiguation errors were the driver. If the claim holds, the anomaly rate offers a scalable, gold-standard-free diagnostic for data quality and reveals that disambiguation quality is not gender-neutral in historical records.","feed_headline":"The 62% first-year 'PI' anomaly flags bad name disambiguation","feed_subtitle":"Missing affiliations push the rate above 80%; persistent IDs cut it to 40%, exposing errors that can bias gender analyses.","key_machinery":"The load-bearing object is the \"immediate-independence anomaly rate\": the share of authors in a cohort whose first last-author paper is dated in the same calendar year as their debut publication. It is computed from the distance in years between first publication and first last-author publication, under the field-specific convention that last author means senior or corresponding author. The mechanism carries the argument because it converts an unobservable quality, disambiguation correctness, into an observable distribution that has a plausible real-world ceiling; any rate far above that ceiling signals that author profiles are contaminated. The authors use metadata subsets, such as the presence of affiliation or of a persistent researcher identifier, as controlled comparisons to show that the anomaly moves with disambiguation difficulty.","core_discovery":"On the paper's own terms, last-author position in biomedical publishing marks the principal investigator, so the time from a researcher's first publication to their first last-author publication estimates career independence. The authors show that more than half of the 5.8 million researchers in their cohort are recorded as reaching that milestone in their debut year, and that this share stays above 60% across entry cohorts. They interpret the anomaly as evidence of name-disambiguation error, because authors without affiliation metadata show rates above 80%, while authors with a persistent self-reported identifier show rates around 40%. On that basis they claim the anomaly rate can serve as a reliable warning signal for disambiguation quality, and they apply it to show that female authors were more likely to be affected than male authors before roughly 2010. The paper's central claim is not a new disambiguation algorithm but a diagnostic: a high anomaly rate indicates poor disambiguation, and group comparisons built on such data should be re-examined.","pith_inferences":["The same anomaly metric could be recalibrated for other fields where the corresponding-author convention differs, since disciplines with rapid PI transitions or non-standard author ordering may have higher natural baselines.","A direct testable extension is to link anomalous author profiles to known-correct identity records and verify whether the apparent first last-author paper is in fact authored by a different person with the same name.","If incomplete early-career coverage rather than name conflation drives part of the anomaly, the metric could also function as a coverage diagnostic, not just a disambiguation diagnostic.","The paper's gender finding suggests that historical gender-disparity results may need to be re-examined even when disambiguation quality appears acceptable on average, because subgroup-specific bias can hide behind aggregate rates."],"forward_implications":["The anomaly rate can be computed from any large bibliographic database with author-order information, so it offers a scalable quality check without requiring manually curated gold standards.","Datasets with high anomaly rates should be treated as risky for career-timing studies, because debut-year last-authorship likely reflects record contamination rather than unusually fast promotion.","Gender comparisons based on pre-2010 biomedical publication data may be biased by differential disambiguation quality between female and male authors, and those results warrant rechecking.","Improving metadata completeness, such as adding affiliation details or persistent researcher identifiers, should reduce the anomaly and strengthen downstream science-of-science analyses.","Because the same signal can be stratified by any author attribute, it provides a general way to test whether disambiguation quality varies across subgroups."],"supporting_citations":[{"why":"Supplies the open-access bibliometric database whose snapshot provides all publication histories and author metadata for the study.","marker":"[12]"},{"why":"Provides the last-author analysis method used to estimate when a researcher becomes an independent principal investigator.","marker":"[11]"},{"why":"Defines the persistent self-reported researcher identifier treated as a high-quality disambiguation benchmark in the comparison.","marker":"[9]"},{"why":"Supplies the name-based gender classification model used to separate male and female author groups.","marker":"[19]"},{"why":"Supports the field convention that last authors are corresponding or senior authors, which the anomaly definition relies on.","marker":"[20]"},{"why":"Documents the conflation and splitting risks in author name disambiguation that motivate interpreting the anomaly as an error signal.","marker":"[5]"},{"why":"Notes the imprecision of name-based gender inference, used to qualify the reported gender differences in disambiguation quality.","marker":"[24]"},{"why":"Motivates the restriction to careers starting after 1999 by describing the historical cap on listed authors that could otherwise distort early publication records.","marker":"[15]"}],"fun_headline_variants":["First-year PI anomaly flags poor name disambiguation","60% of researchers become PIs in year one? Data flaw","Name disambiguation quality exposed by career timing anomaly","Missing affiliations push PI anomaly rate above 80%","ORCID cuts first-year PI anomaly to 40% — signal of error"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the abnormally high first-year last-author rate is caused by disambiguation errors and not by other database problems, especially incomplete coverage of each author's early publications.","fun_headline_variants_meta":{"raw":{"variants":["First-year PI anomaly flags poor name disambiguation","60% of researchers become PIs in year one? Data flaw","Name disambiguation quality exposed by career timing anomaly","Missing affiliations push PI anomaly rate above 80%","ORCID cuts first-year PI anomaly to 40% — signal of error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1492,"prompt_tokens":1009,"completion_tokens":483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":401}},"tokens_in":625,"tokens_out":483,"duration_ms":4890,"temperature":1.0,"reasoning_tokens":401,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:29:16.504126+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a curated set of biomedical author profiles whose true publication histories are known; if the true histories show the same roughly 60% debut-year last-author rate, then the anomaly is not a disambiguation signal. More directly, check whether authors flagged as anomalous have their earliest recorded papers missing from the database: if many anomalous authors have no publications before the first last-author paper despite known earlier work, the signal tracks coverage gaps rather than name conflation.","supporting_citations":[{"cited_title":"Publication metrics and success on the academic job market","cited_arxiv_id":null,"evidence_quote":"Provides the last-author analysis method used to estimate when a researcher becomes an independent principal investigator."},{"cited_title":"Orcid: a system to uniquely identify researchers","cited_arxiv_id":null,"evidence_quote":"Defines the persistent self-reported researcher identifier treated as a high-quality disambiguation benchmark in the comparison."},{"cited_title":"An open-source cultural consen- sus approach to name-based gender classification","cited_arxiv_id":null,"evidence_quote":"Supplies the name-based gender classification model used to separate male and female author groups."},{"cited_title":"Contributorship and division of labor in knowledge production","cited_arxiv_id":null,"evidence_quote":"Supports the field convention that last authors are corresponding or senior authors, which the anomaly definition relies on."},{"cited_title":"Author name disambiguation","cited_arxiv_id":null,"evidence_quote":"Documents the conflation and splitting risks in author name disambiguation that motivate interpreting the anomaly as an error signal."},{"cited_title":"Name-based demographic inference and the unequal distribution of misrecognition","cited_arxiv_id":null,"evidence_quote":"Notes the imprecision of name-based gender inference, used to qualify the reported gender differences in disambiguation quality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the restriction to careers starting after 1999 by describing the historical cap on listed authors that could otherwise distort early publication records."}],"review_version":1}