REVIEW 2 major objections 5 minor 17 references
Criteria for assessing grant applications: A systematic review
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Grant peer review can be described by 15 evaluation criteria applied to 30 evaluated entities, grouped into aims, means, and outcomes.
desk verdict Useful framework and honest synthesis, but 'criteria peers use' overstates the evidence—several included studies developed criteria rather than observed them. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing tool is a four-component analytical framework that splits any evaluative act into an evaluated entity, an evaluation criterion, a frame of reference, and an assigned value. This lets the authors parse criteria that studies report at different levels of abstraction and systematically relate criteria to entities. The synthesis then uses conceptual counting and community detection on a bipartite network of 30 entities and 15 criteria; the detected communities, merged into five, ground the aims/means/outcomes conceptualization.
What would settle it
Conduct an inductive content analysis of a large corpus of actual review reports from a context the synthesis underrepresented, such as fellowship programs, interdisciplinary funding panels, or non-Western agencies. If coders need criteria beyond the 15 (for example, strategic importance, environmental sustainability, or return on investment) or evaluated entities beyond the 30, the claim that these categories describe grant peer review criteria would be shown to be incomplete.
Extended reading notes
Core claim
The central discovery is that grant peer review, though usually described as a chaotic plurality of criteria, has a describable underlying structure. The synthesis claims that peers evaluate the proposed project's aims and expected outcomes in terms of originality and relevance; the research process in terms of quality, appropriateness, rigor, coherence/justification, clarity, and completeness; the resources needed to carry out the project in terms of feasibility; and the applicant both as an instrument for implementation and, separately, as a person evaluated along motivation, traits, and diversity. The paper therefore argues that non-epistemic criteria are not merely noise or bias but are constitutive parts of peer review as practiced, so grant review as actually conducted conflicts with the fairness doctrine and the ideal of impartiality.
Load-bearing premise
The central assumption is that the 12 included studies, mostly from US and European medical and health sciences, individual written review, and project funding, are representative enough that the 15 criteria and 30 entities capture the main targets of grant peer review; the authors themselves acknowledge this limitation.
Editorial extensions
If this is right
- Funding agencies can compare their prescribed review forms against an empirically grounded repertoire of 15 criteria and 30 entities rather than relying on unstated conventions.
- Research on grant peer review can move from studying only reliability, fairness, and predictive validity to studying content validity: whether the criteria used are the right ones.
- Because applicant-focused criteria such as motivation, traits, and diversity are part of actual peer review, normative models for peer review must move beyond the fairness doctrine and the ideal of impartiality and decide which non-epistemic values are legitimate.
- Future empirical work should examine fellowship and career-development programs, where the applicant may be the primary evaluated entity rather than a supporting resource.
- The weak overlap among included studies indicates that current evidence is fragmented; the entity/criterion framework gives future studies a common language for comparison.
Reading between the lines
- The explicit mapping of applicant criteria onto the review process offers a plausible mechanism for documented demographic biases: if reviewers assess traits, diversity, and motivation, social identity can enter funding decisions through criteria that look individual rather than institutional. This connection is our inference, not the paper's claim.
- The proposed 15-criteria/30-entity map can be turned into a testable instrument: code a fresh corpus of review reports from under-represented contexts, such as interdisciplinary grants, fellowship interviews, or non-Western funders, and check whether new criteria or entities emerge.
- The entity/criterion distinction may transfer to other evaluative settings, including journal peer review and research assessment exercises; if it holds there, it would suggest a general logic of academic evaluation rather than a grant-specific one.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a systematic review of empirical-inductive studies on criteria used in grant peer review. It introduces an analytical framework that separates evaluated entities from evaluation criteria, based on Scriven's Logic of Evaluation and Goertz's concept structure. The review includes 12 studies, coded with a transparent protocol: dual screening, dual coding with Krippendorff's alpha (0.78 for entities, 0.69 for criteria), conceptual counting, Jaccard similarity, and bipartite network analysis. The authors identify 15 evaluation criteria and 30 evaluated entities, group them via network communities into aims, means, and outcomes, and compare the resulting 'criteria formula' with research-quality literature and funding-agency guidelines. They then discuss the findings against the fairness doctrine and the ideal of impartiality, concluding that grant peer review may be considered unfair and biased because peers also use non-epistemic criteria. The paper explicitly acknowledges limitations: few studies, overrepresentation of medical/health sciences, USA/Europe, individual review, and project funding, and the possibility of additional criteria and entities.
Significance. If the synthesis is valid, the paper fills a genuine gap: it offers a systematic, empirically grounded map of grant-review criteria and their associated evaluated entities, plus a reusable framework (entity vs. criterion vs. frame of reference) that can structure future research and practical review design. The methodological transparency is a strength: the search strings are given in full, inclusion/exclusion criteria are explicit, reliability is quantified, and the coding data are presented in detailed tables. The network-based conceptualization into aims, means, and outcomes is a useful synthesis that goes beyond a simple list. The comparison with research-quality frameworks and funder guidelines also provides a bridge to adjacent literatures. The main significance would rest on the descriptive claim that these are criteria peers actually use, which is exactly the point that needs careful scrutiny.
major comments (2)
- [Inclusion criteria; Qualitative synthesis] The abstract and discussion repeatedly state that the review identifies 'the criteria peers use' and concludes that 'peers assess proposals, as our synthesis has shown, also in terms of non-epistemic criteria' (Discussion). However, inclusion criterion 1 explicitly admits studies that 'developed peer review criteria for grant proposals' as well as studies that 'established reasons used by peers', and Table 1 shows that 7 of the 12 included studies are coded as 'Improving grant peer review'. Studies such as Schmitt et al. (2015), Lahtinen et al. (2005), and Whaley et al. (2006) develop criteria for a funding program or rating form, i.e., they are normative-empirical or prescriptive in purpose, not descriptions of actual peer practice. The synthesis and the network analysis pool these studies with descriptive ones (e.g., Lamont 2009, Reinhart 2010, Pier et al. 2018) without weighting or separating them. As a result, the claim that peers use criteria such as motivation, traits, and diversity is not directly supported by the aggregate evidence: motivation and traits appear in only two studies (Lamont; van Arensbergen et al. 2014b; see Table 3 and Supplementary Part D), and diversity in three, with several of these contributions potentially coming from improvement-focused studies that deliberately include policy goals. This is a construct-validity problem, not merely a coverage problem, because the descriptive claim underlies the fairness-doctrine discussion and the 'criteria formula'. I request a sensitivity analysis or separate reporting for descriptive vs. prescriptive/improvement studies, and a rewording of claims that imply all included evidence describes actual peer behavior.
- [Results; Discussion and conclusion] The 'personal qualities' cluster (motivation, traits, diversity) is given a prominent place in the conceptualization (Figure 4) and in the fairness-doctrine conclusion, but the evidence base for this cluster is very thin. Only two studies contribute motivation and traits, and only three contribute diversity (Table 3, Supplementary Part D). The network description itself states that the red community containing these criteria is 'only weakly connected to the other communities' (Section 'Results'). Yet the Discussion elevates this cluster to one of the four components of the 'criteria formula' and uses it to argue that grant peer review deviates from the fairness doctrine and the ideal of impartiality. Given the small number of studies and the likely mix of descriptive and prescriptive sources, the conceptualization should either be restricted to the well-supported core (research quality, description quality, feasibility) or explicitly labeled as a provisional hypothesis about an understudied aspect, rather than a synthesized finding about peer practice. This is load-bearing because the normative conclusion about unfairness depends on the reliability of this cluster.
minor comments (5)
- [Searching and screening] The text states 'A total of 3,558 records were identified' in the citation-based search, but the flow diagram (Figure 1) and the subsequent arithmetic (3,541 discarded plus 47 screened) indicate 3,588; please correct the number in the text.
- [Network analysis] The sentence 'Using the the DIRTLPAwb+ algorithm' contains a duplicated 'the'; please fix.
- [References] The reference for Goertz (2006) has a typo: 'Princetion University Press' should be 'Princeton University Press'.
- [Network analysis; Results] The decision to merge community 3 with community 4 because community 3 could not be interpreted is a substantive analytical choice that directly shapes the conceptualization in Figure 4. Please report the unmerged community structure in the supplementary materials for transparency.
- [Inclusion criteria; Notes] The main text classifies studies as 'improving' or 'understanding' (Table 1) and only the supplementary materials (Part A) introduce the descriptive-inductive vs. normative-prescriptive distinction. Since the paper's central claim is about criteria that peers use, the main text should explicitly address the extent to which included studies are descriptive rather than prescriptive, and should clarify that 'developed' in inclusion criterion 1 covers both empirical derivation of criteria and consensus-based development for a funding program.
Circularity Check
No significant circularity; the synthesis is derived from external studies and an external analytical framework.
full rationale
The paper's central contribution is a qualitative content analysis and network synthesis of 12 external studies that inductively examined grant review criteria. The entity/criterion framework is imported from Scriven (1980) and Goertz (2006), not defined in terms of the paper's own conclusions. The 15 criteria and 30 entities are coded from the included studies with reported inter-coder reliability (Krippendorff's alpha 0.78 for entities, 0.69 for criteria), and the network communities are detected with the external DIRTLPAwb+ algorithm. No equation or construction makes an output quantity equal to an input by definition. The only self-citations (Hug et al. 2013, Ochsner et al. 2013) are background or exclusion references and do not carry the central claim. The paper's candid limitations about overrepresentation of medical/health fields and Western regions affect generalizability, not circularity. The pooling of improvement-oriented and understanding-oriented studies is a construct-validity concern but does not reduce the synthesis to its inputs by construction. Therefore no circular step is identified.
Assumptions & free parameters
assumptions (4)
- domain assumption Scriven's Logic of Evaluation and Goertz's concept structure provide a valid basis for decomposing any evaluation criterion into evaluated entity, evaluation criterion, frame of reference, and assigned value.
- domain assumption The 12 included studies accurately report the criteria peers actually use in grant review.
- domain assumption The included 12 studies are sufficiently representative to support a general conceptualization of grant review criteria.
- standard math Krippendorff's alpha thresholds are an acceptable basis for drawing tentative conclusions from the coding.
Cite this review
Pith. "Pith review of Criteria for assessing grant applications: A systematic review." pith.science (2026). https://pith.science/paper/TCRFORDA
@misc{pith2026190802598,
author = {Pith},
title = {Pith review of: Criteria for assessing grant applications: A systematic review},
year = {2026},
howpublished = {\url{https://pith.science/paper/TCRFORDA}},
note = {Machine review of arXiv:1908.02598}
}
read the original abstract
Criteria are an essential component of any procedure for assessing merit. Yet, little is known about the criteria peers use in assessing grant applications. In this systematic review we therefore identify and synthesize studies that examine grant peer review criteria in an empirical and inductive manner. To facilitate the synthesis, we introduce a framework that classifies what is generally referred to as 'criterion' into an evaluated entity (i.e. the object of evaluation) and an evaluation criterion (i.e. the dimension along which an entity is evaluated). In total, this synthesis includes 12 studies. Two-thirds of these studies examine criteria in the medical and health sciences, while studies in other fields are scarce. Few studies compare criteria across different fields, and none focus on criteria for interdisciplinary research. We conducted a qualitative content analysis of the 12 studies and thereby identified 15 evaluation criteria and 30 evaluated entities as well as the relations between them. Based on a network analysis, we propose a conceptualization that groups the identified evaluation criteria and evaluated entities into aims, means, and outcomes. We compare our results to criteria found in studies on research quality and guidelines of funding agencies. Since peer review is often approached from a normative perspective, we discuss our findings in relation to two normative positions, the fairness doctrine and the ideal of impartiality. Our findings suggest that future studies on criteria in grant peer review should focus on the applicant, include data from non-Western countries, and examine fields other than the medical and health sciences.
Reference graph
Works this paper leans on
-
[1]
Data given as number and percentage of total studies included (N = 12)
Characteristics of studies that inductively examined grant review criteria. Data given as number and percentage of total studies included (N = 12). Characteristic Summary data Publication year First study 1990 Latest study 2018 Mean 2006 Median 2007.5 Fields in which criteria were studied1 Natural sciences 2 (17%) Engineering and technology 2 (17%) Medica...
work page 1990
-
[4]
Conceptualization of evaluation criteria and evaluated entities used in grant peer review. Evaluated entities identified in the qualitative content analysis are displayed in boxes and regular type. Evaluation criteria are linked to evaluated entities by grey lines. MeansProject resources AimsTopicResearch question Research processCurrent stateTheoryApproa...
work page 2019
-
[8]
The coders then compared and discussed their codes and agreed on common codes
In particular, coders went through the data line by line and generated a new code each time data could not be subsumed under existing codes. The coders then compared and discussed their codes and agreed on common codes. They worked through the first part of the data again and then jointly revised and finalized the codes. Residual codes (e.g. ‘other entity...
work page 2004
-
[11]
For example, Moed (2005) und Harnad (2008,
Generating evidence on the content validity of peer review can be a goal in itself, but content validation could also be a way to overcome the circularity inherent to the validation strategies discussed and employed in research on peer review and bibliometrics. For example, Moed (2005) und Harnad (2008,
work page 2005
-
[13]
Cronin and Sugimoto 2015, Harnad 2008, Marsh et al
but neither peer review ratings nor bibliometric indicators fulfill this requirement (e.g. Cronin and Sugimoto 2015, Harnad 2008, Marsh et al. 2008). From our point of view, a possible solution could be to generate evidence on the content validity of peer review and thus validate peer ratings (for content validation, see Haynes et al. 1995)
work page 2015
-
[15]
We understand a value as ‘something that is desirable or worthy of pursuit’ (Elliott 2017, p. 11). For example, the evaluation criterion ‘originality’ is an epistemic value while ‘extra-academic relevance’ is a social value. 31 References Abdoul H, Perrey C, Amiel P, Tubach F, Gottot S, Durand-Zaleski I and Alberti C (2012) Peer review of grant applicatio...
2012
-
[219]
and ‘to determine how much “Weight of Evidence” should be given to the findings of a research study’ (Gough 2007, p. 214). In the present study, we did not conduct such an appraisal for two reasons. First, we have applied narrow inclusion criteria and we therefore consider all included studies as useful. Second, giving some studies more weight than others...
work page 2007
-
[252]
articulated the ‘fairness doctrine’ which holds that access to journal space and federal funds has to be ‘judged on the merit of one’s ideas, not on the basis of academic rank, sex, place of work, publication record, and so on’. The fairness doctrine resembles the ideal of impartiality, which implicitly underlies quantitative research on bias in peer revi...
work page 2013
Show all 17 references
-
[279]
quality, poor – good, weak – strong)
Quality Criterion evaluates an entity in terms of general quality (incl. quality, poor – good, weak – strong). Examples: ‘methodological quality’, ‘weak dissemination plan’ 10 (3%) 10 (4%) Originality Criterion evaluates the originality of an entity. Evaluations of originality...
1996
-
[639]
Most studies analyzed criteria of individual reviews (n = 7), while only few studies focused on criteria used in panels (n = 3)
than the sample size of the studies that elicited data from scholars (unit: persons, mean = 48, median = 48.5, minimum = 12, maximum = 81). Most studies analyzed criteria of individual reviews (n = 7), while only few studies focused on criteria used in panels (n = 3). One stud...
2007
-
[2005]
research proposal*
or the cognitive distance between reviewers and applicants (van den Besselaar and Sandström 2017). Bias and fairness factors are often discussed in literature reviews (e.g. Guthrie et al. 2018, Lee et al. 2013), however, they are not referred to as criteria or contextualized a...
2009
-
[2006]
Prescriptive criteria of funding agencies, as summarized by Abdoul et al
that are (re)shaped by the actual assessment in which they are enacted (Kaltenbrunner and de Rijcke 2019). Prescriptive criteria of funding agencies, as summarized by Abdoul et al. (2012), Berning et al. (2015), Falk-Krzesinski and Tobin (2015), and Langfeldt and Scordato (201...
2012
-
[2009]
the criterion variable)
discuss the validation of bibliometric indicators by correlating them with peer review ratings (i.e. the criterion variable). Some studies also proceed conversely and seek to validate peer judgments with bibliometric indicators (i.e. the criterion variable in this case). These...
2014
-
[2014]
and the Handbook of Test Development (Lane et al. 2016). It casts peer review as an instrument or test that has to be evaluated with respect to efficiency, reliability, fairness, and (predictive) validity
2016
-
[2015]
American Journal of Physical Medicine & Rehabilitation 70(1): 161–164
Thomas J P and Lawrence T S (1991) Common deficiencies of NIDRR research applications. American Journal of Physical Medicine & Rehabilitation 70(1): 161–164. Thorngate W, Dawes R M and Foddy M (2009) Judging merit. Taylor & Francis: New York. Tong A, Flemming K, McInnes E, Oli...
1991
-
[2017]
Drawing on Douglas (2016) and Elliott (2017), the following questions may guide the development of new normative models for peer review
could prove to be particularly fruitful for this purpose as it started from the value-free ideal, which is similar to the fairness doctrine and impartiality ideal, and has advanced to acknowledging and including non-epistemic values. Drawing on Douglas (2016) and Elliott (2017...
2016
-
[2019]
Alpha was calculated for each of these variables and the coefficients of the 15 criteria variables were then averaged to provide a single reliability index for the coded criteria
to assess the inter-coder reliability of the two coders with regard to the evaluation criteria (15 variables) and the evaluated entities (one variable). Alpha was calculated for each of these variables and the coefficients of the 15 criteria variables were then averaged to pro...
1977
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.