REVIEW 5 major objections 5 minor 4 references
Reviewer comments reveal 49 unwritten rules for bibliometric studies
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Peer review comments on bibliometric studies were systematically classified, yielding 49 recommendations that overlap partially with three existing reporting guidelines.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A genuinely useful inductive method for mining reporting standards from peer reviews, but the headline coverage numbers contradict the results and the sample is too narrow to support 'community standards.' the 5 major comments →
Implicit reporting standards in bibliometric research: what can reviewers' comments tell us about reporting completeness?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On its own terms, the paper’s central claim is that open peer reviews encode the reporting norms bibliometricians actually apply when judging a manuscript, and that those norms can be systematically recovered. The evidence is a three-sample corpus: articles from one library-and-information-science journal with open reviews, submissions to one science-and-technology-indicators conference with open reviews, and a small set of non-LIS articles, totalling 182 reviews of 85 studies. From 968 in-scope comments the authors inductively built 11 thematic categories—description of data and methods alone accounts for 29.1% of comments—and 68 sub-categories, then phrased each sub-category as a recommend
What carries the argument
The carrying mechanism is an inductive coding pipeline: 968 reviewer comments are sorted into broad themes, refined into 68 sub-categories, and each sub-category is translated into an “Authors should…” recommendation, producing 49 items. The same items are then coded against PRIBA, GLOBAL, and BIBLIO using a complete/partial/no-match scheme, which lets the paper quantify how much of the review-derived standard the guidelines already cover and where they are silent.
Load-bearing premise
The load-bearing premise is that reviewers' comments are a valid proxy for the reporting standards of the whole bibliometric community, even though nearly all reviewed studies come from one journal and one conference and every study in the sample was accepted.
What would settle it
Collect open reviews from a broader set of bibliometric venues and check whether the same 49 recommendations emerge. If reviewers in other venues ask for different details—such as funding disclosure, software parameters, or title and abstract identification—or if the frequency ranking of topics changes materially, then the 'community standard' label fails and the list instead describes two venues' review cultures. A second check: score a fresh batch of published bibliometric papers against the 49 items and test whether low-scoring papers draw more clarification requests in later reviews.
If this is right
- The 49 recommendations can be used by authors as a pre-submission checklist alongside existing guidelines, covering specifics such as author-name disambiguation, fractionalisation, normalisation, sample sizes, and figure captions.
- Developers of future reporting guidelines can treat open peer reviews as a complementary inductive source to Delphi and expert-consensus methods, capturing details that experts may not think to list.
- The nine topics common to all four lists—research question, databases, data-collection dates, search terms, inclusion/exclusion criteria, flowchart, results description, discussion against the literature, and limitations—can be read as a minimal core for complete reporting.
- Review-derived recommendations focus on academic content and omit title/abstract guidance, software identification, funding declarations, and study strengths; those functions require guidelines or journal policies rather than reviewer comments.
- Because the reviews were not based on a fixed checklist, the frequency numbers describe what reviewers happened to mention, not the true prevalence of reporting problems in bibliometric studies.
Where Pith is reading between the lines
- The authors’ “community standards” claim is broader than the sample can fully support: with nearly all journal articles coming from one venue and all conference papers from one conference, the recovered standards are most directly those of two review cultures; reproducing the analysis in other venues would test whether the list is genuinely community-wide.
- A testable extension: turn the 49 recommendations into a scoring rubric and apply it to a fresh batch of published bibliometric papers; if low-scoring papers receive more reviewer clarification requests, the recommendations capture what reviewers actually need.
- Absence of a topic in reviewer comments does not prove it is not a standard—reviewers rarely comment on what is routinely done well, so guidelines may legitimately include items reviewers never mention, such as title and abstract details.
- The overlap percentages are unstable because GLOBAL is still being piloted; re-running the comparison with the final published GLOBAL version could shift the reported coverage figures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a qualitative and quantitative analysis of 968 reviewer comments extracted from 182 open peer reviews of 85 bibliometric studies (39 LIS articles from Quantitative Science Studies, 41 STI 2023 conference papers, and 5 non-LIS articles). The comments are inductively coded into 11 broad themes and 68 sub-categories, from which the authors derive 49 reporting recommendations for bibliometric studies. These recommendations are then compared with three existing reporting guidelines (GLOBAL, PRIBA, BIBLIO) by qualitative matching. The paper reports that reviewers focus on completeness and clarity of data, methods, and results, and concludes that the derived recommendations represent implicit community standards for reporting bibliometric studies and could inform future guideline development.
Significance. The study has clear strengths: the dataset is shared, the three-stage coding process is transparent, and the bottom-up, review-derived perspective is a genuine complement to the Delphi-based process used to develop existing guidelines. The 49 concrete recommendations are a useful resource for authors and guideline developers. However, the paper's central claims are currently overreaching: the sample is drawn almost entirely from two venues with two review policies, and the abstract's headline coverage statistics are inconsistent with the results section. With substantial qualification and corrected reporting, the work can make a valuable pilot contribution to the reporting-guidelines literature, but as written it does not support the 'implicit community standards' conclusion.
major comments (5)
- The quantitative headline is inconsistent with the results. The Abstract and Discussion state that 'the guidelines covered 45-65% of our recommendations' and 'addressed 60-80% of the guidelines' items.' §3.3 reports, for our 49 items: 26 (53.1%) present in GLOBAL, 16 (32.7%) in PRIBA, and 18 (36.7%) in BIBLIO. Conversely, §3.3.2 reports that 16/25 PRIBA items (64.0%), 24/29 GLOBAL items (82.8%), and 14/20 BIBLIO items (70.0%) are addressed by our recommendations. Thus the correct ranges are 33-53% and 64-83%, not 45-65% and 60-80%. The Conclusion's 'up to two-thirds' also overstates a maximum of 53.1%. These numbers are central to the paper's comparative claim and must be corrected and harmonized throughout.
- The conclusion that the 49 recommendations 'represent the implicit standards of the community for reporting bibliometric studies' is not supported by the sample. As the authors acknowledge in §4.1, 'essentially only two review policies are in effect': all LIS articles come from Quantitative Science Studies, all conference papers from STI 2023, and only five articles from non-LIS journals. All sampled papers were accepted, and open peer review itself is not the norm. The data can support claims about the implicit standards of reviewers at these specific venues under these specific review policies; they cannot, without additional evidence, support a community-wide claim. Either the central claim must be reframed as venue-specific/exploratory, or representativeness evidence must be provided.
- Three authors participated in the development of GLOBAL, one of the three guidelines benchmarked against the inductively derived recommendations. The competing-interests statement says 'The authors have no competing interests' despite this disclosed involvement. Because the comparison in §2.4 relies on subjective judgments about whether items are 'similar, if not identical, or partial aspects' of each other, the GLOBAL comparison is at particular risk of unintended bias. The manuscript should either have the item matching conducted blindly or by independent coders, include a reliability assessment, and explicitly discuss this conflict and its possible influence on the reported overlap rates.
- The derivation of recommendations from reviewer comments assumes that reviewer comments are a valid proxy for community reporting standards. This assumption is weakened by the non-systematic nature of the comments: reviewers did not use a common checklist, the review forms differed by venue, and all manuscripts in the sample were ultimately accepted, with 80.5% of conference papers accepted. Consequently, the prevalence figures in Table 1 (e.g., 77.7% negative comments) and the resulting recommendations reflect these specific review contexts and exclude rejection-oriented comment types. The limitations paragraph in §4.1 acknowledges several of these points, but the framing of the recommendations as community standards in §5 does not carry the necessary caveats.
- No inter-coder reliability or sensitivity analysis is reported for the three-stage thematic coding or for the guideline-item matching. Since the central overlap percentages depend entirely on these classifications, a reliability check (e.g., a second coder coding a subsample, or a strict versus lenient matching threshold) would materially strengthen the main comparative claim. Without it, the quantitative precision of numbers such as 53.1% versus 32.7% is difficult to interpret.
minor comments (5)
- In the discussion of specificity, the text refers to 'our item 2' when describing the recommendation about describing bibliographic databases. Table 3 shows that item 2 concerns reviewing and reporting sufficient literature; the intended reference appears to be item 10.
- The row for 'Conflict of interest' under Declarations reads '1 01 100.0'; '01' should be '0.1'.
- The phrase 'The recent surge in bibliometric studies published' is grammatically awkward; consider 'The recent surge in published bibliometric studies'.
- Because comments could be assigned to more than one broad category, the percentages in Table 1 sum to more than 100%. The text notes that the total count exceeds 968, but a table note explaining that row percentages are computed over the total number of categorizations rather than the number of comments would help readers.
- The competing-interests statement should be reworded to acknowledge the GLOBAL development involvement as a competing interest rather than stating 'no competing interests' while simultaneously noting that involvement.
Circularity Check
No circular derivation: recommendations are built inductively from peer reviews; GLOBAL/PRIBA/BIBLIO are external benchmarks.
full rationale
The derivation chain is self-contained: 968 reviewer comments were inductively coded (Section 2.2) into 11 categories, 68 sub-categories, and then 49 recommendations; the three guidelines enter only afterward as external comparators (Sections 2.4, 3.3). No guideline item is used to construct a recommendation, and no quantitative result is a fitted parameter renamed as a prediction. The only self-citation (Ng et al. 2025, co-authored by D. Stephen) describes GLOBAL's development and is background, not load-bearing; the disclosed GLOBAL involvement of three authors is a conflict-of-interest/bias risk in the qualitative matching, not a definitional loop. The conclusion that the recommendations 'represent the implicit standards of the community' restates the Section 1.1 interpretive assumption and is weakened by the sample's two review policies (Section 4.1), but this is a generalizability/overclaim issue rather than a circular reduction. The abstract's coverage range (45-65%) conflicts with the reported 33-53% figures, an accuracy error, not circularity.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Peer reviewer comments on manuscripts are a valid reflection of implicit reporting standards of the scientific community.
- domain assumption The qualitative categorization and matching of comments to guideline items is reliable despite no inter-coder reliability assessment.
- domain assumption The three existing guidelines (GLOBAL, PRIBA, BIBLIO) are appropriate benchmarks for bibliometric reporting.
Cite this review
Pith. "Pith review of Implicit reporting standards in bibliometric research: what can reviewers' comments tell us about reporting completeness?." pith.science (2026). https://pith.science/paper/I6YTR7XH
@misc{pith2026250816276,
author = {Pith},
title = {Pith review of: Implicit reporting standards in bibliometric research: what can reviewers' comments tell us about reporting completeness?},
year = {2026},
howpublished = {\url{https://pith.science/paper/I6YTR7XH}},
note = {Machine review of arXiv:2508.16276}
}
read the original abstract
The recent surge in bibliometric studies published has been accompanied by increasing diversity in the completeness of reporting these studies' details, affecting reliability, reproducibility, and robustness. Our study systematises the reporting of bibliometric research using open peer reviews. We examined 182 peer reviews of 85 bibliometric studies published in library and information science (LIS) journals and conference proceedings, and non-LIS journals. We extracted 968 reviewer comments and inductively classified them into 11 broad thematic categories and 68 sub-categories, determining that reviewers largely focus on the completeness and clarity of reporting data, methods, and results. We subsequently derived 49 recommendations for the details authors should report and compared them with the GLOBAL, PRIBA, and BIBLIO reporting guidelines to identify (dis)similarities in content. Our recommendations addressed 60-80% of the guidelines' items, while the guidelines covered 45-65% of our recommendations. Our recommendations provided greater range and specificity, but did not incorporate the functions of guidelines beyond addressing academic content. We argue that peer reviews provide valuable information for the development of future guidelines. Further, our recommendations can be read as the implicit community standards for reporting bibliometric studies and could be used by authors to aid complete and accurate reporting of their manuscripts.
Figures
Reference graph
Works this paper leans on
-
[1]
Boyack, K. W., Klavans, R., & Smith, C. (2022). Raising the bar for bibliometric analysis. In N. Robinson-Garcia, D. Torres-Salinas, & W. Arroyo-Machado (Eds.), 26th International Conference on Science and Technology Indicators, STI
work page 2022
-
[12]
DOI: 10.1186/s13643-023-02410-2. Moher, D., Schulz, K. F., Simera, I., & Altman, D. G. (2010). Guidance for developers of health research reporting guidelines. PLOS Medicine, 7(2), e1000217. DOI: 10.1371/journal.pmed.1000217. Ng, J. Y., Haustein, S., Ebrahimzadeh, S., Chen, C., Sabe, M., Solmi, M., & Moher, D. (2023). Guidance List for reporting bibliomet...
-
[1686]
https://doi.org/10.21105/joss.01686
-
[2022]
sti22143. DOI: 10.5281/zenodo.6975632. Cabezas-Clavijo, A., Milanés-Guisado, Y., Alba-Ruiz, R., & Delgado-Vázquez, A. M. (2023). The need to develop tailored tools for improving the quality of thematic bibliometric analyses: Evidence from papers published in Sustainability and Scientometrics. Journal of Data and Information Science, 8(4), 10–35. DOI: 10.2...
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.