{"id":"69f839e0-10bf-4e34-9303-ed814e7ae31b","arxiv_id":"2505.06801","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new composite framework ranks Web3 grant maturity, placing Arbitrum LTIPP and Optimism in the developmental stage, Mantle in the foundational stage, and Taiko in the experimental stage.","lead":"This paper builds a maturity scorecard for Web3 grant programs and applies it to four Ethereum layer-two programs, finding Arbitrum's LTIPP and Optimism's Mission Rounds more mature than Mantle and Taiko. It offers grant operators a benchmarking tool and informs the design of a new grant management platform called AUTHBOND.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Min-max normalization against the six sampled programs makes the maturity-stage thresholds sample-dependent; Taiko's 'experimental' label may be a sample artifact.","rationale":"The paper is transparent about its methodology and provides a reproducible composite calculation, which is a strength. However, the maturity stages are the main output, and their validity depends on the normalization producing an absolute scale. The reader's weakest assumption identifies exactly this issue. My analysis confirms that the min-max procedure uses the sample range, so the 'quartiles from zero to one' are not quartiles of an external distribution but arbitrary equal bins on a sample-dependent scale. This is load-bearing because the headline findings (Taiko experimental, ARB LTIPP developmental) would change if the sample composition changed. The missing-data issue in Table V is secondary but reinforces the concern: scoring absent indicators as zero can artificially lower Taiko's score, possibly pushing it below the 0.25 threshold. The Mantle stage contradiction (experimental in the abstract and grouped text, foundational in §IV-B.4) is a reporting error that further weakens confidence but does not determine the framework's validity. The proposed leave-one-out test would settle whether the classifications are robust. Given these issues, I do not disagree with the reader's CONDITIONAL verdict; the paper needs reanalysis with fixed anchors or sensitivity analysis before the maturity stages can be trusted.","tokens_in":13840,"tokens_out":5050,"duration_ms":49388,"concrete_test":"Perform a leave-one-out renormalization: for each of the six programs, recompute the min and max of the additive GMF scores (or underlying rubric scores) using only the other five programs, then re-normalize all six using those anchors. If Taiko's normalized score rises above 0.25 or ARB LTIPP's falls below 0.75 in any leave-one-out run, the stage thresholds are sample-dependent and the classification is not robust. Alternatively, set fixed anchors by adding synthetic programs with theoretical minimum (0) and maximum (1) indicator values and re-run the normalization; if any stage assignment changes, the quartile boundaries lack absolute meaning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ARB LTIPP is 'developmental' and Taiko is 'experimental' rests on stage boundaries defined as quartiles of a min-max normalized scale computed from the six observations. Section III.B states that the data was 'normalised through the min-max normalisation function' and that 'the maturity stages in Table I were defined as quartiles from zero to one.' Because the min and max are taken from the sample, adding or removing a program changes all normalized scores. There is no external anchor tying the 0.25/0.5/0.75 thresholds to an absolute maturity scale. If a higher-maturity program were included, Taiko's 0.2334 could exceed 0.25; if a lower one were included, ARB LTIPP's 0.6755 could fall below 0.75. The paper does not justify that these six observations span the full maturity range, so the stage classifications are relative to the sample. This is compounded by Table V showing Taiko with zeros in GOV, EFI, and TAC, which may represent missing data rather than true zero maturity, further depressing its score.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the Grant Maturity Framework (GMF), a composite-indicator model for evaluating the maturity of Web3 grant programs across six rubric categories (FAO, PSO, GOV, EFI, TAC, COM). The framework is constructed through a Delphi-based rubric scoring exercise and applied to six observations drawn from four Ethereum L2 grant programs: Arbitrum STIP/Backfund, STIP Bridge, LTIPP, Optimism Mission Rounds, Mantle, and Taiko. The authors compute normalized composite scores, classify programs into four maturity stages (experimental, foundational, developmental, advanced), and use the results to inform features of a Web3 grant platform prototype. The central claim is that the GMF provides a systematic, replicable measure of grant program maturity and that the application shows ARB LTIPP and Optimism as more mature, while Taiko and Mantle are early-stage.","tokens_in":14043,"tokens_out":4639,"duration_ms":44562,"significance":"The paper addresses a genuinely underexplored topic—Web3 grant program evaluation—and proposes a structured, mixed-methods framework with a practical application to platform design. The explicit adaptation of the World Bank GTMI approach and the transparent use of equal weights are strengths. If the framework's methodological issues were resolved, the GMF could serve as a useful benchmarking and self-assessment tool for Web3 grant operators. However, as it stands, the sample-relative normalization and missing-data treatment undermine the validity of the absolute maturity classifications, and the internal contradiction in the reported results further reduces confidence in the findings.","major_comments":[{"comment":"The claim that the GMF measures an absolute 'maturity' is undermined by the construction: indicators are min-max normalized using the six programs in the sample, and the maturity stages are defined as quartiles of this normalized composite. Because the min and max are sample-dependent, the boundaries 0.25, 0.5, and 0.75 are relative to the composition of the analyzed set. Adding a program with higher scores than ARB LTIPP would shift the max and could move ARB LTIPP below the 0.75 boundary, while adding a lower-scoring program could move Taiko above 0.25. The paper provides no external anchor or justification that the six observations span the maturity range, so the absolute labels 'experimental' and 'developmental' are not supported.","section":"Section III.B, Table III"},{"comment":"There is a direct internal contradiction in the reported results. The introductory paragraph of Section IV-B states that Taiko and Mantle 'scored lower ... which places them in the experimental stage,' but Table IV lists Mantle's normalized composite as 0.2729, which falls in the foundational stage (0.25–0.5), and subsection IV-B.4 explicitly classifies Mantle as 'foundational.' The paper must resolve this inconsistency before the results can be considered reliable.","section":"Section IV-B, Table IV"},{"comment":"The treatment of missing or undocumented indicators as zero scores is a load-bearing methodological choice. Taiko receives zeros in GOV, EFI, and TAC, which appear to reflect the absence of publicly documented governance, impact, and transparency practices rather than verified nonexistence. Min-max normalization then sets Taiko's raw score to the minimum, artificially depressing its composite. Without a clear rule distinguishing zero maturity from missing data and without a sensitivity analysis, the rankings and stage assignments in Table IV are not robust.","section":"Section III.B, Table V"},{"comment":"The paper's replicability claim is not met. It states that 'public data was used to allow for the replicability of the construction of the framework,' but the manuscript does not provide the full list of 46 indicators, the subset of 40 used in the index, the raw data values per program, or the rubric scores from the Delphi study. An independent researcher cannot reproduce the normalized rubric scores in Table V or the composite scores in Table IV. The authors should include a data appendix or supplementary material with the indicator definitions, data, and calculation steps.","section":"Section III.B"}],"minor_comments":[{"comment":"The phrase 'little is known on about their effectiveness' contains a typo; it should read 'little is known about their effectiveness.'","section":"Abstract"},{"comment":"The header 'F AO' should be 'FAO' (remove the extra space).","section":"Table V"},{"comment":"The subsection title 'Taiko Labs Grants' is inconsistent with the program name 'Taiko's Incentivisation Grant Program' used elsewhere in Section IV-B; please unify the terminology.","section":"Section IV-B.3"},{"comment":"Several references lack complete metadata, including page numbers or DOIs (e.g., [3], [45], [58]); please provide full citation details.","section":"References"},{"comment":"Figures 1 and 2 are referenced in the text but do not appear in the manuscript; if this is the full submission, the figures need to be included, or the references should be removed.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of the journal, but the methodological issues are substantial and currently undermine the central claims. The sample-relative normalization could be reframed as a relative benchmarking tool, but the absolute maturity-stage language is not justified. I also note the small sample size and the fact that some Delphi participants are affiliated with the analyzed programs; the authors should at least acknowledge this potential conflict in the final version. I recommend asking for a major revision that addresses the normalization, missing-data handling, internal contradiction, and data availability before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper proposes a Grant Maturity Framework (GMF) for Web3 grant programs and applies it to six observations from four L2 ecosystems. The application is new, the formulas are transparent, and the authors are honest about the framework's limits. But the maturity stages are defined as quartiles of a min-max normalized composite computed from these six programs, so labels like \"experimental\" and \"developmental\" are sample-relative, not absolute. Add a higher-scoring program and Taiko's 0.2334 could cross into \"foundational\"; add a lower one and ARB LTIPP could drop below \"developmental\". The paper never justifies that these six programs span the relevant range. That is the load-bearing flaw.\n\nThe paper does some things well. It adapts the World Bank GTMI to an understudied area, uses a sensible set of rubric categories, and makes the equal-weighting assumption explicit. The link to the AUTHBOND platform gives the work a practical feedback loop, and the authors acknowledge the GMF is a non-exhaustive first step. The Delphi-based indicator selection is standard practice and not itself a problem.\n\nThe soft spots are real but not fatal to the core idea. First, the zero scores for Taiko in GOV, EFI, and TAC likely represent missing documentation rather than measured zero maturity; scoring absence as zero depresses the score without discussion. Second, the dataset is not shipped, so the replicability claim rests on a promise. Third, there is a direct internal contradiction: Section IV-B says Mantle's 27.29% places it in the experimental stage, but the thresholds in Table III and subsection IV-B.4 place it in foundational. That needs fixing. The circularity concern about Delphi experts informing both indicator selection and qualitative scoring is worth noting but secondary; the quantitative indicators come from public program data, so the feedback is not total.\n\nFor a niche audience of Web3 grant operators and DAO tooling designers, this is a reasonable benchmarking starting point. The ordering of the programs (ARB and Optimism ahead of Mantle and Taiko) is plausible and probably robust; the stage labels are fragile. I would send this to a serious referee—it deserves engagement, not a desk reject—but the authors need to publish the data, run a sensitivity analysis, anchor the thresholds to something external, and correct the Mantle classification before I'd trust the stages. I likely wouldn't cite it in current form, but I'd revisit after revision.","headline":"Useful niche framework but the stage labels are sample-relative and the Mantle classification contradicts itself; worth a revision cycle, not a desk reject.","tokens_in":14570,"tokens_out":2057,"would_cite":false,"duration_ms":22581,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Grant Maturity Framework scores Web3 grant programs on a composite of 40 indicators and classifies them into four maturity stages, with Arbitrum's LTIPP ranking highest and Taiko lowest.","keywords":["maturity model","Web3 governance","grant programs","decentralised autonomous organisations","crypto-economic systems","mixed-methods","composite index","Ethereum layer 2"],"falsifier":"Recompute the GMF for the same six programs plus a larger sample of, say, twenty additional Web3 grant programs using the same 40 indicators and the same min-max normalisation, and test whether Taiko's normalised score remains below 0.25 and Mantle's above it; if adding programs flips a program's stage classification, the thresholds are sample-relative rather than absolute.","tokens_in":13636,"feed_emoji":"📊","tokens_out":7996,"duration_ms":69265,"temperature":0.7,"pith_summary":"The paper introduces the Grant Maturity Framework (GMF), a mixed-methods index that measures the maturity of Web3 grant programs across six rubric clusters: focus and objectives, program structure, governance, effectiveness and impact, transparency, and community engagement. The framework combines expert rubric scores with 40 quantitative indicators, normalises them, and aggregates them into a composite score between 0 and 1, with quartile thresholds corresponding to experimental, foundational, developmental, and advanced stages. Applying the GMF to six Ethereum layer-two grant programs, the paper finds that Arbitrum's Long-Term Incentive Pilot Program (LTIPP) scores highest at 67.55% and Taiko's program lowest at 23.34%, with the other programs spread across the foundational and developmental stages. The paper also reports that the diagnostic results directly informed the design of a Web3 grant platform prototype, including milestone-based funding and a streamlined proposal submission flow. For grant operators, the value is a replicable benchmarking tool that turns qualitative program reputation into a structured, comparable maturity profile.","feed_headline":"ARB LTIPP tops new Web3 grant maturity ranking","feed_subtitle":"A 40-indicator framework scores Ethereum L2 grant programs on structure, governance, impact, and community.","key_machinery":"The central object is the Grant Maturity Framework (GMF) itself: a composite index built from six rubric clusters—Focus Areas and Objectives, Program Structure, Governance, Effectiveness and Impact, Transparency, and Community Engagement—each populated by indicators from public program data and expert scores. The carrying mechanism is a two-level equal-weight aggregation: first, the normalised indicators within each rubric are averaged into a composite rubric score; then the six rubric scores are normalised and averaged into a single GMF score in [0,1]. Min-max normalisation renders different indicator units comparable, and predefined quartile cutoffs translate the numeric score into a maturity stage (experimental, foundational, developmental, advanced). This design turns qualitative judgments about governance and community engagement into an auditable number, and the per-rubric breakdown tells operators which dimension is dragging their maturity down.","core_discovery":"The central claim is that Web3 grant program effectiveness can be measured through the concept of maturity, and that the Grant Maturity Framework (GMF) provides a systematic, replicable way to do so. Concretely, the GMF assigns each program a composite score by averaging normalised rubric scores across six dimensions, where each rubric score is itself an equal-weight average of normalised indicators drawn from public program documentation and expert scoring. Applied to the six observed programs, the framework yields a clear ordering: ARB LTIPP at 0.6755 and Optimism Mission Rounds at 0.6105 fall in the developmental stage; ARB STIP and STIP Bridge at 0.4349 and 0.5251 fall in the foundational stage; Mantle at 0.2729 sits just above the experimental/foundational threshold; and Taiko at 0.2334 falls in the experimental stage. The paper further claims that these scores are actionable: low impact and transparency scores across programs motivated concrete platform features such as milestone-based payment releases and a one-step submission flow. The GMF is intended as a baseline and benchmark for program operators, not a final verdict on any individual grant's quality.","pith_inferences":["The authors do not claim, but the equal-weighting choice means the maturity scores are a modelling decision; allowing rubric-specific weights could shift the stage boundaries for Mantle and Taiko.","The GMF's monotone 'more process and transparency is better' logic may not generalise to deliberately lightweight grant experiments; a maturity score could be paired with a separate measure of context-fit.","A straightforward reliability check, not run in the paper, would be to have two independent expert panels score the same six programs; if the assigned stages differ, the rubric's subjectivity is the binding constraint.","Because min-max normalisation is relative to the six observed programs, the stage labels ('experimental' etc.) are best read as rankings within this sample rather than absolute certificates of maturity."],"forward_implications":["Web3 grant operators can use the GMF as a diagnostic benchmark, seeing which of the six rubric dimensions is dragging their program's maturity down.","Programs near a stage boundary, such as Mantle at 0.2729, can identify the specific indicators that would lift them into the next quartile.","The GMF's structure extends to grant programs outside Ethereum L2s and to retroactive funding, as long as the normalisation set and thresholds are recalibrated.","The platform prototype described in the paper shows that low scores in impact and transparency translate into concrete design features, such as milestone-based payment releases and a one-step submission flow."],"supporting_citations":[{"why":"Supplies the underlying government digital maturity index approach that the GMF adapts to Web3 grant programs.","marker":"[5]"},{"why":"Provides the explicit definition of maturity as a dynamic state or degree of perfection that the framework operationalises.","marker":"[31]"},{"why":"Grounds the four-stage model in the idea of discrete evolution stages along a typical maturity path.","marker":"[33]"},{"why":"Provides the composite-indicator construction guidelines the GMF follows for normalisation and aggregation.","marker":"[48]"},{"why":"Introduces the rubric scoring framework whose clusters and indicator set feed into the GMF.","marker":"[37]"},{"why":"Supplies the Delphi expert-scoring method used to develop the rubric and select indicators.","marker":"[28]"}],"fun_headline_variants":["ARB LTIPP tops new Web3 grant maturity ranking","GMF scores grant programs: ARB LTIPP leads, Taiko lags","Web3 grants maturity: ARB LTIPP and Optimism ahead of rivals","40-indicator GMF ranks ARB LTIPP first, Taiko last","Grant Maturity Framework: ARB LTIPP top, Optimism second"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's maturity classifications assume that scoring a grant program as zero on indicators it does not document, and min-max normalising each indicator against just the six analysed programs, yields an absolute scale on which the quartile stage thresholds are meaningful for the whole population of Web3 grant programs.","fun_headline_variants_meta":{"raw":{"variants":["ARB LTIPP tops new Web3 grant maturity ranking","GMF scores grant programs: ARB LTIPP leads, Taiko lags","Web3 grants maturity: ARB LTIPP and Optimism ahead of rivals","40-indicator GMF ranks ARB LTIPP first, Taiko last","Grant Maturity Framework: ARB LTIPP top, Optimism second"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1475,"prompt_tokens":1006,"completion_tokens":469,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":372}},"tokens_in":622,"tokens_out":469,"duration_ms":4671,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:32:46.488672+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the GMF for the same six programs plus a larger sample of, say, twenty additional Web3 grant programs using the same 40 indicators and the same min-max normalisation, and test whether Taiko's normalised score remains below 0.25 and Mantle's above it; if adding programs flips a program's stage classification, the thresholds are sample-relative rather than absolute.","supporting_citations":[{"cited_title":"Assembling a population health management maturity index using a Delphi method,","cited_arxiv_id":null,"evidence_quote":"Supplies the Delphi expert-scoring method used to develop the rubric and select indicators."},{"cited_title":"Dener, H","cited_arxiv_id":null,"evidence_quote":"Supplies the underlying government digital maturity index approach that the GMF adapts to Web3 grant programs."},{"cited_title":"Organizational maturity models: the leading research fields and opportunities for further studies,","cited_arxiv_id":null,"evidence_quote":"Provides the explicit definition of maturity as a dynamic state or degree of perfection that the framework operationalises."},{"cited_title":"Risk management maturity model,","cited_arxiv_id":null,"evidence_quote":"Grounds the four-stage model in the idea of discrete evolution stages along a typical maturity path."},{"cited_title":"Cryptocurrency Prices, Charts And Market Capitaliza- tions,","cited_arxiv_id":null,"evidence_quote":"Provides the composite-indicator construction guidelines the GMF follows for normalisation and aggregation."},{"cited_title":"Evaluating Progress in Web3 Grants: Introducing the Grant Maturity Index","cited_arxiv_id":"2410.19828","evidence_quote":"Introduces the rubric scoring framework whose clusters and indicator set feed into the GMF."}],"review_version":1}