{"id":"923d24ad-6bf8-462d-a2a9-1cc217a99ed5","arxiv_id":"1908.02598","paper_version":3,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The synthesis of 12 inductive studies yields 15 evaluation criteria and 30 evaluated entities used in grant peer review, grouped into aims, means, and outcomes.","lead":"This systematic review synthesized 12 studies on how grant reviewers actually judge proposals and organized the results into a framework separating what is evaluated from the criteria used. It provides an evidence-based map of review criteria that funders, applicants, and researchers can use to study or improve grant peer review.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthesis pools descriptive and prescriptive studies; the claim that peers use non-epistemic criteria (motivation, traits, diversity) may rest on funder-designed criteria, not observed peer practice.","rationale":"The reader's weakest assumption was representativeness of the 12-study corpus across fields and regions, and I agree that is a limitation acknowledged by the authors. However, I identify a distinct, more fundamental concern: the synthesis conflates descriptive studies of peer practice with prescriptive studies that develop criteria for funders. The paper's own study-characteristics coding distinguishes 'improving' from 'understanding,' but the synthesis treats all 12 studies as equivalent evidence for what peers use. This matters because the headline finding about non-epistemic criteria (motivation, traits, diversity) and the fairness critique depend on the descriptive interpretation. A re-analysis by study purpose is a concrete, feasible check that would settle the issue. If the concern lands, the paper remains useful as a map and framework, but the central descriptive claim and the normative conclusions would need substantial qualification. Hence I recommend CONDITIONAL rather than a full rejection, since the issue is addressable by re-analysis or reframing.","tokens_in":23597,"tokens_out":7046,"duration_ms":75170,"concrete_test":"Re-run the qualitative content analysis, Jaccard overlap, and bipartite network analysis separately for the descriptive subset (purpose = understanding, data from actual reviews or interviews about actual practice) versus the prescriptive/improvement subset (purpose = improving, Delphi or rating-form development). If the 'personal qualities' community (motivation, traits, diversity) and the non-epistemic criteria disappear or lose coherence in the descriptive subset, then the claim that peers use these criteria is unsupported; the manuscript should then be revised to present the results as a mixed synthesis of actual and recommended criteria, or to restrict the 'criteria used by peers' claims to the descriptive subset.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the 15 criteria and 30 entities describe what peers use when assessing grant applications. But the inclusion criteria explicitly admit studies that 'developed peer review criteria,' and 7 of the 12 included studies are coded as 'improving grant peer review,' several of which set up criteria for a funding program (e.g., Schmitt et al. 2015 Delphi, Lahtinen et al. 2005 criteria development, Whaley et al. 2006 rating-form development). These are prescriptive or normative-empirical studies, not descriptions of actual peer behavior. The synthesis and network analysis pool them with descriptive studies (e.g., Lamont 2009, Reinhart 2010, Pier et al. 2018) without weighting or separating them. Consequently, the abstract's statement that 'peers use' criteria such as motivation, traits, and diversity is not directly supported by the aggregate evidence: motivation and traits appear in only two studies (Lamont; van Arensbergen), and diversity may be driven by improvement-focused studies that deliberately include policy goals. The fairness-doctrine discussion rests on this descriptive claim, so the conflation is load-bearing. This is a construct-validity problem, not merely a coverage problem: it affects what the synthesis actually demonstrates about peer review practice.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic review of empirical-inductive studies on criteria used in grant peer review. It introduces an analytical framework that separates evaluated entities from evaluation criteria, based on Scriven's Logic of Evaluation and Goertz's concept structure. The review includes 12 studies, coded with a transparent protocol: dual screening, dual coding with Krippendorff's alpha (0.78 for entities, 0.69 for criteria), conceptual counting, Jaccard similarity, and bipartite network analysis. The authors identify 15 evaluation criteria and 30 evaluated entities, group them via network communities into aims, means, and outcomes, and compare the resulting 'criteria formula' with research-quality literature and funding-agency guidelines. They then discuss the findings against the fairness doctrine and the ideal of impartiality, concluding that grant peer review may be considered unfair and biased because peers also use non-epistemic criteria. The paper explicitly acknowledges limitations: few studies, overrepresentation of medical/health sciences, USA/Europe, individual review, and project funding, and the possibility of additional criteria and entities.","tokens_in":23743,"tokens_out":7448,"duration_ms":72341,"significance":"If the synthesis is valid, the paper fills a genuine gap: it offers a systematic, empirically grounded map of grant-review criteria and their associated evaluated entities, plus a reusable framework (entity vs. criterion vs. frame of reference) that can structure future research and practical review design. The methodological transparency is a strength: the search strings are given in full, inclusion/exclusion criteria are explicit, reliability is quantified, and the coding data are presented in detailed tables. The network-based conceptualization into aims, means, and outcomes is a useful synthesis that goes beyond a simple list. The comparison with research-quality frameworks and funder guidelines also provides a bridge to adjacent literatures. The main significance would rest on the descriptive claim that these are criteria peers actually use, which is exactly the point that needs careful scrutiny.","major_comments":[{"comment":"The abstract and discussion repeatedly state that the review identifies 'the criteria peers use' and concludes that 'peers assess proposals, as our synthesis has shown, also in terms of non-epistemic criteria' (Discussion). However, inclusion criterion 1 explicitly admits studies that 'developed peer review criteria for grant proposals' as well as studies that 'established reasons used by peers', and Table 1 shows that 7 of the 12 included studies are coded as 'Improving grant peer review'. Studies such as Schmitt et al. (2015), Lahtinen et al. (2005), and Whaley et al. (2006) develop criteria for a funding program or rating form, i.e., they are normative-empirical or prescriptive in purpose, not descriptions of actual peer practice. The synthesis and the network analysis pool these studies with descriptive ones (e.g., Lamont 2009, Reinhart 2010, Pier et al. 2018) without weighting or separating them. As a result, the claim that peers use criteria such as motivation, traits, and diversity is not directly supported by the aggregate evidence: motivation and traits appear in only two studies (Lamont; van Arensbergen et al. 2014b; see Table 3 and Supplementary Part D), and diversity in three, with several of these contributions potentially coming from improvement-focused studies that deliberately include policy goals. This is a construct-validity problem, not merely a coverage problem, because the descriptive claim underlies the fairness-doctrine discussion and the 'criteria formula'. I request a sensitivity analysis or separate reporting for descriptive vs. prescriptive/improvement studies, and a rewording of claims that imply all included evidence describes actual peer behavior.","section":"Inclusion criteria; Qualitative synthesis"},{"comment":"The 'personal qualities' cluster (motivation, traits, diversity) is given a prominent place in the conceptualization (Figure 4) and in the fairness-doctrine conclusion, but the evidence base for this cluster is very thin. Only two studies contribute motivation and traits, and only three contribute diversity (Table 3, Supplementary Part D). The network description itself states that the red community containing these criteria is 'only weakly connected to the other communities' (Section 'Results'). Yet the Discussion elevates this cluster to one of the four components of the 'criteria formula' and uses it to argue that grant peer review deviates from the fairness doctrine and the ideal of impartiality. Given the small number of studies and the likely mix of descriptive and prescriptive sources, the conceptualization should either be restricted to the well-supported core (research quality, description quality, feasibility) or explicitly labeled as a provisional hypothesis about an understudied aspect, rather than a synthesized finding about peer practice. This is load-bearing because the normative conclusion about unfairness depends on the reliability of this cluster.","section":"Results; Discussion and conclusion"}],"minor_comments":[{"comment":"The text states 'A total of 3,558 records were identified' in the citation-based search, but the flow diagram (Figure 1) and the subsequent arithmetic (3,541 discarded plus 47 screened) indicate 3,588; please correct the number in the text.","section":"Searching and screening"},{"comment":"The sentence 'Using the the DIRTLPAwb+ algorithm' contains a duplicated 'the'; please fix.","section":"Network analysis"},{"comment":"The reference for Goertz (2006) has a typo: 'Princetion University Press' should be 'Princeton University Press'.","section":"References"},{"comment":"The decision to merge community 3 with community 4 because community 3 could not be interpreted is a substantive analytical choice that directly shapes the conceptualization in Figure 4. Please report the unmerged community structure in the supplementary materials for transparency.","section":"Network analysis; Results"},{"comment":"The main text classifies studies as 'improving' or 'understanding' (Table 1) and only the supplementary materials (Part A) introduce the descriptive-inductive vs. normative-prescriptive distinction. Since the paper's central claim is about criteria that peers use, the main text should explicitly address the extent to which included studies are descriptive rather than prescriptive, and should clarify that 'developed' in inclusion criterion 1 covers both empirical derivation of criteria and consensus-based development for a funding program.","section":"Inclusion criteria; Notes"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid methodological core and the framework is a useful contribution. The main concern is not the quality of the coding or the transparency, but the construct-validity gap between the evidence base (which mixes descriptive and prescriptive studies) and the claims about criteria that peers actually use. With a careful re-analysis that separates study purposes, or with appropriately qualified wording, the paper would be a strong contribution. I would advise the editor to require the authors to address this distinction before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. It is a genuinely useful systematic review, not a breakthrough, and the main claim in the abstract is a bit too confident. What is new: this is the first systematic synthesis of inductively derived grant review criteria, and the entity/criterion split plus the aims-means-outcomes structure gives the field a common vocabulary. The method is respectable: explicit inclusion criteria, documented search strings, multi-round screening, dual coding with Krippendorff's alpha (0.78 for entities, 0.69 for criteria), conceptual counting, and network analysis. The authors also do the honest thing in the limitations: small corpus, health sciences/US-Europe overrepresentation, English-based search, no systematic grey literature.\n\nThe soft spot is real and load-bearing. Their inclusion criterion admits studies that 'developed peer review criteria' as well as studies that established reasons peers actually used. Seven of the twelve studies are coded as 'improving' grant peer review. Some set up criteria for a funding program (Schmitt's Delphi, Whaley's rating form, Lahtinen's quality criteria). Those are normative or normative-empirical studies, not descriptions of observed peer practice. The synthesis and network analysis pool them with descriptive studies like Lamont, Reinhart, and Pier. So the abstract's claim that peers use criteria such as motivation, traits, and diversity is not directly supported by the aggregate evidence. Motivation and traits appear in only two studies; diversity appears in three and may be driven by improvement-oriented designs. The fairness doctrine discussion rests on this descriptive claim, so this is not a minor caveat. The fix is easy: report the descriptive and prescriptive subsets separately, or at least weight the claims to match the evidence.\n\nMinor issues: the Jaccard overlap of 0.39 is weak, and the authors decline quality appraisal, which they justify but which limits weight-of-evidence reasoning. These are minor relative to the conflation.\n\nWho is this for: anyone working on peer review, funders designing review forms, and applicants wanting to know what reviewers look at. It deserves a serious referee. My recommendation: send it to review, with a required revision that distinguishes observed criteria from developed/deemed-appropriate criteria and softens the 'peers use' language accordingly. The framework and map are worth keeping.","headline":"Useful framework and honest synthesis, but 'criteria peers use' overstates the evidence—several included studies developed criteria rather than observed them.","tokens_in":24337,"tokens_out":3521,"would_cite":true,"duration_ms":36967,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Grant peer review can be described by 15 evaluation criteria applied to 30 evaluated entities, grouped into aims, means, and outcomes.","keywords":["grant peer review","grant funding","evaluation criteria","evaluated entities","systematic review","qualitative synthesis","fairness doctrine","impartiality"],"falsifier":"Conduct an inductive content analysis of a large corpus of actual review reports from a context the synthesis underrepresented, such as fellowship programs, interdisciplinary funding panels, or non-Western agencies. If coders need criteria beyond the 15 (for example, strategic importance, environmental sustainability, or return on investment) or evaluated entities beyond the 30, the claim that these categories describe grant peer review criteria would be shown to be incomplete.","tokens_in":23312,"feed_emoji":"🔬","tokens_out":5467,"duration_ms":53316,"temperature":0.7,"pith_summary":"This systematic review identifies and synthesizes all empirically grounded, inductive studies of the criteria peers actually use when judging grant applications. It introduces a vocabulary that separates the object being judged (the evaluated entity) from the dimension along which it is judged (the evaluation criterion), and uses it to extract 15 criteria and 30 entities from 12 included studies. A network analysis groups these into a conceptual map in which projects are assessed for aims, means, and outcomes, while applicants are assessed both as project resources and as persons (motivation, traits, diversity). The paper shows that actual peer review includes non-epistemic applicant-centered criteria, which puts it in tension with the fairness doctrine and the ideal of impartiality.","feed_headline":"15 criteria and 30 targets map grant peer review","feed_subtitle":"A synthesis of 12 studies shows reviewers weigh motivation, traits, and diversity—not just scientific merit.","key_machinery":"The load-bearing tool is a four-component analytical framework that splits any evaluative act into an evaluated entity, an evaluation criterion, a frame of reference, and an assigned value. This lets the authors parse criteria that studies report at different levels of abstraction and systematically relate criteria to entities. The synthesis then uses conceptual counting and community detection on a bipartite network of 30 entities and 15 criteria; the detected communities, merged into five, ground the aims/means/outcomes conceptualization.","core_discovery":"The central discovery is that grant peer review, though usually described as a chaotic plurality of criteria, has a describable underlying structure. The synthesis claims that peers evaluate the proposed project's aims and expected outcomes in terms of originality and relevance; the research process in terms of quality, appropriateness, rigor, coherence/justification, clarity, and completeness; the resources needed to carry out the project in terms of feasibility; and the applicant both as an instrument for implementation and, separately, as a person evaluated along motivation, traits, and diversity. The paper therefore argues that non-epistemic criteria are not merely noise or bias but are constitutive parts of peer review as practiced, so grant review as actually conducted conflicts with the fairness doctrine and the ideal of impartiality.","pith_inferences":["The explicit mapping of applicant criteria onto the review process offers a plausible mechanism for documented demographic biases: if reviewers assess traits, diversity, and motivation, social identity can enter funding decisions through criteria that look individual rather than institutional. This connection is our inference, not the paper's claim.","The proposed 15-criteria/30-entity map can be turned into a testable instrument: code a fresh corpus of review reports from under-represented contexts, such as interdisciplinary grants, fellowship interviews, or non-Western funders, and check whether new criteria or entities emerge.","The entity/criterion distinction may transfer to other evaluative settings, including journal peer review and research assessment exercises; if it holds there, it would suggest a general logic of academic evaluation rather than a grant-specific one."],"forward_implications":["Funding agencies can compare their prescribed review forms against an empirically grounded repertoire of 15 criteria and 30 entities rather than relying on unstated conventions.","Research on grant peer review can move from studying only reliability, fairness, and predictive validity to studying content validity: whether the criteria used are the right ones.","Because applicant-focused criteria such as motivation, traits, and diversity are part of actual peer review, normative models for peer review must move beyond the fairness doctrine and the ideal of impartiality and decide which non-epistemic values are legitimate.","Future empirical work should examine fellowship and career-development programs, where the applicant may be the primary evaluated entity rather than a supporting resource.","The weak overlap among included studies indicates that current evidence is fragmented; the entity/criterion framework gives future studies a common language for comparison."],"supporting_citations":[{"why":"Supplies the qualitative study of French reviewer practices that is one of the 12 synthesized datasets and grounds criteria such as rigor and appropriateness.","marker":"Abdoul et al. 2012"},{"why":"Contributes the humanities and social-science panel criteria, especially originality, that anchor the aims/outcomes community.","marker":"Guetzkow et al. 2004"},{"why":"Provides the interview-based account of panel judgment and is the source of the motivation, traits, and diversity criteria in the applicant-person community.","marker":"Lamont 2009"},{"why":"The earliest inductive study; its review arguments from the German Research Foundation are part of the pool that shapes the project-process criteria.","marker":"Hartmann and Neidhardt 1990"},{"why":"Content analysis of external reviews that supplies many of the project-entity and criterion connections used in the network.","marker":"Reinhart 2010"},{"why":"Large corpus of NIH critiques that contributes completeness, clarity, and feasibility relations and the applicant-as-resource coding.","marker":"Pier et al. 2018"},{"why":"One of only two studies reporting motivation and traits, and thus load-bearing for the applicant-person community.","marker":"van Arensbergen et al. 2014b"},{"why":"Delphi-based criteria for the German Innovation Fund that supply extra-academic relevance and feasibility connections.","marker":"Schmitt et al. 2015"}],"fun_headline_variants":["15 criteria, 30 targets: grant review's hidden structure","Grant reviews weigh merit and more: 15 criteria found","Systematic review: 15 criteria shape grant peer review","Grant peer review: not just merit, but fairness clash","Study maps 15 grant review criteria, 30 evaluated entities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the 12 included studies, mostly from US and European medical and health sciences, individual written review, and project funding, are representative enough that the 15 criteria and 30 entities capture the main targets of grant peer review; the authors themselves acknowledge this limitation.","fun_headline_variants_meta":{"raw":{"variants":["15 criteria, 30 targets: grant review's hidden structure","Grant reviews weigh merit and more: 15 criteria found","Systematic review: 15 criteria shape grant peer review","Grant peer review: not just merit, but fairness clash","Study maps 15 grant review criteria, 30 evaluated entities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000758,"raw_usage":{"total_tokens":3370,"prompt_tokens":952,"completion_tokens":2418,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":2335}},"tokens_in":568,"tokens_out":2418,"duration_ms":17256,"temperature":1.0,"reasoning_tokens":2335,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:14:56.281018+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct an inductive content analysis of a large corpus of actual review reports from a context the synthesis underrepresented, such as fellowship programs, interdisciplinary funding panels, or non-Western agencies. If coders need criteria beyond the 15 (for example, strategic importance, environmental sustainability, or return on investment) or evaluated entities beyond the 30, the claim that these categories describe grant peer review criteria would be shown to be incomplete.","supporting_citations":[],"review_version":1}