{"id":"274009fd-b5d1-4a3a-a9b6-22b726d02ecc","arxiv_id":"1908.01344","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Across three citizen science platforms, most volunteers explore multiple projects but regularly engage in only one or two, a small set of projects draws most contributions, and volunteers inherited from other projects are more engaged than those recruited externally.","lead":"Citizen science volunteers who start in one project often try several projects but regularly stick with just a few, and volunteers drawn from other projects on the same platform tend to contribute more than those recruited from outside. The paper develops a metrics-and-inspection method for measuring such cross-project engagement and uses it on three real platforms to guide platform design.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Balance-in-computing is confounded by selection: 'inherited' volunteers are prior returnees by definition, so the comparison against all outside recruits does not show that cross-project features add value.","rationale":"The reader's weakest assumption targets the attribution of the recruiting project. That is a legitimate measurement-error concern: if volunteers browse but do not contribute before their first task, or arrive via a recommendation that is not the first project they click, the 'first task = recruiter' rule misclassifies. However, even if the true recruiter were known perfectly, the comparison in Section 4.2's balance-in-computing remains confounded because the operational definition of 'inherited' requires prior successful participation in another project. Inherited volunteers are necessarily a subset of volunteers who have already demonstrated enough engagement to return to the platform and contribute elsewhere; outside-recruited volunteers include every first-time contributor, most of whom are classified 'platform transient' in Table 2. Thus the positive balance in computing is predicted under a null model of no cross-project benefit, solely from selection on prior retention. The paper's Section 4.3 discussion attributes the result to intrinsic motivation ('volunteering behaviour rather than helping behaviour') but does not control for the composition difference. A within-volunteer comparison would distinguish the selection explanation from the value-added explanation. This concern is load-bearing because the design guideline to encourage cross-project engagement (Section 4.3, guideline 1) rests on the claim that inherited volunteers are more engaged. I therefore recommend keeping the CONDITIONAL verdict (UNCHANGED), with the requested revision being the addition of a within-volunteer or matching analysis, or a softening of the causal interpretation. I also note the definitional slip in Section 4.2, where 'multi-project regular' is stated identically to 'multi-project explorer' ('executed tasks on at least two different projects'), contradicting the claim that the classes are mutually exclusive; this should be corrected but is secondary to the central claim.","tokens_in":13507,"tokens_out":9561,"duration_ms":89713,"concrete_test":"Recompute the balance-in-computing analysis using a within-volunteer design: for every volunteer who contributed to at least two distinct projects on Crowdcrafting, Socientize, and GeoTag-X, calculate the number of tasks done in their first project (where they are classified 'recruited outside') and the number done in their second project (where they are classified 'inherited'), and test the paired difference across all such volunteers with a Wilcoxon signed-rank test per platform. If the median difference is not significantly positive (or is negative), the project-level inherited advantage in Figure 3 is a selection artifact rather than evidence that cross-project recruitment increases engagement. If the within-volunteer difference is positive and robust, the central claim survives this concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 defines a volunteer as 'recruited' by the project in which she/he performed the first task after registering, and as 'inherited' by any project in which she/he performs tasks after having performed tasks in another project. The balance-in-computing result (Section 4.2, Figure 3) then compares, per project, the average number of tasks of inherited vs outside-recruited volunteers, and the paper concludes that 'inherited volunteers perform more tasks than the recruited ones,' using this to recommend cross-project engagement (Section 4.3). The load-bearing problem is not only misattribution (the reader's point) but an inherent selection confound: a volunteer cannot be 'inherited' into any project without having already returned to the platform and contributed to a prior project. Inherited volunteers are, by construction, survivors of the initial churn that the paper itself documents (93% platform transients on Crowdcrafting, 67-73% on the others, Table 2). Outside-recruited volunteers are the entire inflow, including the large majority who never return. Any positive average task count difference is therefore expected even if the platform's cross-project features contribute nothing: the inherited group is a self-selected subset of returnees. The paper's limitations (Section 4.3) acknowledge task-complexity and generalizability issues but do not address this selection effect. Without a within-volunteer or matched comparison, the positive balance in computing does not identify a causal benefit of cross-project recruitment, and the design recommendation to encourage cross-project engagement is not supported by this particular result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Goal-Question-Metric (GQM) framework plus the Semiotic Inspection Method (SIM) to characterize volunteers' cross-project task-execution behavior on multi-project citizen science platforms. It applies the quantitative metrics to task-execution logs from Crowdcrafting, Socientize, and GeoTag-X and the qualitative inspection to Crowdcrafting, Zooniverse, and CitSci.org. The main reported findings are that only a minority of volunteers explore multiple projects, very few become regular multi-project contributors, attention and recruitment are highly concentrated in a small number of projects, and volunteers inherited from other projects on the platform tend to perform more tasks than volunteers recruited from outside. The paper also identifies three classes of cross-project interface signs and derives design recommendations, most notably that platforms should offer personalized, explainable project recommendations.","tokens_in":13725,"tokens_out":4257,"duration_ms":48237,"significance":"If the quantitative claims were causally supported, the paper would be a useful empirical contribution to citizen-science HCI: it addresses a real gap regarding cross-project dynamics, its metrics are defined transparently from log data, and its use of actual task-execution data from three PyBossa-based platforms is a strength. The SIM-based inspection is a credible qualitative complement, and the resulting design recommendations are actionable. However, the headline claim that inherited volunteers are more engaged than outside-recruited volunteers is currently supported only by a confounded comparison, which substantially lowers the significance of the paper's central contribution unless the analysis is redone or the claim is reframed.","major_comments":[{"comment":"The 'balance in computing' comparison is confounded by selection. A volunteer is classified as 'inherited' only after having performed tasks in another project on the platform, whereas outside-recruited volunteers are all volunteers whose first task was in the project, including the 67%–93% platform transients documented in Table 2. Inherited volunteers are therefore, by construction, a self-selected subset of returning volunteers, and a positive mean difference in task counts is expected even if cross-project features add no value. The conclusion in Section 4.3 that 'inherited volunteers perform more tasks than the recruited ones' and the design recommendation to invest in cross-project engagement rest directly on this comparison. The paper should compare inherited volunteers with outside-recruited volunteers who have comparable platform tenure or activity (for example, via matching on number of active days or number of prior projects), or perform a within-volunteer comparison of task counts in the first versus subsequent projects, or explicitly reframe the result as a descriptive statement about group composition rather than as evidence for the value of cross-project engagement.","section":"Section 4.2, Figure 3, Section 4.3"},{"comment":"The recruitment attribution assumption is untestable with the available data. The paper estimates that each volunteer was recruited by the project in which she/he performed the first task after registering, but volunteers may browse the platform, see recommendations, or perform tasks in a project that is not the one that actually brought them to the platform. Because the balance-in-recruitment and balance-in-computing metrics depend entirely on this binary classification, the potential direction and magnitude of the resulting misclassification bias should be discussed, and the conclusions should be explicitly conditioned on the assumption. At minimum, the paper should consistently describe the comparison as being between first-project contributors and later-project contributors rather than between 'outside-recruited' and 'inherited' volunteers.","section":"Section 4.1"},{"comment":"The definition of the 'multi-project regular' class is internally inconsistent. The text states that this class consists of volunteers who 'executed tasks on least two different projects', which is the same condition used for the 'multi-project explorer' class, yet the two classes are said to be mutually exclusive and have different reported percentages. If 'regular' is intended to require a minimum number of active days per project, as suggested by the engagement-rate metric in Section 3.2, that condition should be stated explicitly and applied consistently. This distinction matters because Figure 2 and the claim that multi-project regulars exhibit longer relative activity duration rely on it.","section":"Section 4.2, Table 2"}],"minor_comments":[{"comment":"The balance-in-recruitment and balance-in-computing formulas use min(n,u) and min(t,m) in the denominator, which are undefined when either value is zero. Please specify how projects with no inherited or no outside-recruited volunteers are handled in the analysis.","section":"Section 3.2"},{"comment":"There are several typographical and formatting errors, including 'on least two different projects', 'ananalysis', 'Crowcrafting' (in the Figure 3 caption), and 'e.g. of for example' in Section 4.3. Please proofread the text and correct these issues.","section":"Section 4.2"},{"comment":"The caption refers to the balance in recruitment as the 'right' plot and the balance in computing as the 'left' plot, while the surrounding text discusses them in the reverse order. Please verify that the panels and captions correspond to the intended metrics.","section":"Figure 3"},{"comment":"The Gini coefficient values have inconsistent spacing (for example, '0 .95' in the Crowdcrafting row). Please standardize the formatting.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The reader's selection-confound concern is real and lands on the paper's central claim. A modest revision cannot fix this; the authors need to either add a matched or within-volunteer comparison, which is feasible with the logs they already have, or substantially soften the framing from 'inherited volunteers are more valuable' to 'first-project contributors and later-project contributors differ in task counts.' The descriptive parts of the paper, particularly the Gini inequalities and the SIM-based sign taxonomy, are solid and could survive such a revision. I rate my confidence as moderate because the necessary data for a matched analysis may not all be present in the publicly available APIs, but the authors are in the best position to determine that."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a straightforward empirical study of cross-project volunteer behavior on three PyBossa-based citizen science platforms. What's new is the application: GQM-style metrics plus semiotic inspection are brought to the multi-project case, giving platform managers concrete numbers on exploration, regular engagement, project inequality, and recruitment balance. The descriptive findings—few multi-project regulars, high Gini coefficients, most volunteers transient—are plausible and reproduced across three real task-execution logs. The authors are also transparent about scope and about the fact that they are using historical platform data, not a controlled experiment.\n\nThe soft spot is in the headline claim that inherited volunteers perform more tasks than outside-recruited volunteers. The reader flagged the attribution assumption: the project of a volunteer's first task is treated as the recruiter. The stress-test note goes further and is correct: even if you fixed the misattribution, the comparison is confounded by selection. A volunteer can only be 'inherited' by a new project if she first returned to the platform after her initial contribution. The outside-recruited group includes the entire inflow, including the large majority who never return (93% transients on Crowdcrafting, 67–73% on the others). So a positive balance in computing is expected even if cross-project features contribute nothing. The design recommendation to encourage cross-project engagement rests on this result, and that recommendation is not supported by the analysis as presented. A within-volunteer or matched comparison would be needed.\n\nMinor issues: the definition of 'multi-project regular' is muddled—at one point it reads like the same as 'multi-project explorer'—and the Gini values are reported without intervals, so treat them as descriptive. The semiotic inspection is a legitimate but modest addition.\n\nVerdict: a useful descriptive case study, worth knowing for anyone working on citizen science platforms. The causal-sounding takeaway about inherited volunteers should be soft-pedaled. I'd send it to peer review—it deserves referee time—but I would not let the balance-in-computing result stand as evidence for cross-project features.","headline":"A useful descriptive case study of cross-project engagement on citizen science platforms, but its central claim about inherited volunteers' higher engagement is confounded by selection and should not drive design recommendations.","tokens_in":14300,"tokens_out":2469,"would_cite":false,"duration_ms":24322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On multi-project citizen science platforms, volunteers who already contribute to one project tend to do more tasks in a new project than volunteers recruited from outside, even though outside recruitment delivers more people.","keywords":["citizen science","multi-project platforms","volunteer engagement","cross-project participation","human computation","GQM","semiotic inspection","Gini coefficient"],"falsifier":"Recompute the balance-in-computing metric on a platform that logs referral sources or first page views before the first task. If volunteers who verifiably arrived from outside the platform perform as many or more tasks per person than volunteers who verifiably came from other projects, the paper's central claim fails.","tokens_in":13248,"feed_emoji":"🔬","tokens_out":11921,"duration_ms":106025,"temperature":0.7,"pith_summary":"This paper claims that cross-project movement is the central mechanism behind volunteer productivity on multi-project citizen science platforms. Analysing task-execution logs from three platforms, it finds that most volunteers sample several projects but regularly contribute to only a few, that a small set of projects absorbs most volunteer attention, and that volunteers inherited from other projects on the same platform perform more tasks on average than volunteers recruited from outside. The authors read this as evidence that platforms should cultivate movement between projects, through personalised and explainable project recommendations, rather than treating recruitment as the main growth lever. They also contribute a Goal-Question-Metric (GQM) metric suite and a Semiotic Inspection protocol that platform managers can reuse to measure their own cross-project dynamics. The numerical results are presented as case studies, not as a cross-platform comparison.","feed_headline":"Inherited volunteers outwork outside recruits in citizen science","feed_subtitle":"On multi-project platforms, volunteers drawn from other projects do more tasks per person than outside recruits.","key_machinery":"The argument turns on a metric the paper calls balance in computing, $(t - m)/\\min(t, m)$, where $t$ is the average number of tasks performed by volunteers inherited from other projects on the platform and $m$ is the average for volunteers recruited from outside; a positive value means inherited volunteers do more work per person. This metric is the direct evidence for the paper's main claim about the value of cross-project engagement. It sits inside a Goal-Question-Metric (GQM) framework that also defines the exploration rate ($p/a$, projects tried over projects available), the engagement rate ($g/a$, projects contributed to on at least two days over projects available), relative activity duration, balance in recruitment, and two Gini inequality coefficients. The qualitative half uses the Semiotic Inspection Method, a protocol for reading the designer-to-user messages encoded in interface signs, to classify the cross-project features platforms actually communicate.","core_discovery":"On the paper's own terms, the central discovery is a consistent asymmetry in how volunteers reach new projects. On Crowdcrafting, Socientize, and GeoTag-X, most projects recruit more volunteers from outside the platform than from its other projects, yet the volunteers inherited from other projects perform more tasks per person; the paper's 'balance in computing' is positive for most projects. At platform level, between 13 and 26 per cent of volunteers explore multiple projects, at most 6 per cent become regular contributors to multiple projects, and a few flagship projects capture most recruitment and task execution, with Gini coefficients between 0.47 and 0.95. The interface inspection adds a design finding: the platforms offer project search, featured-project lists, and lists of the volunteer's own projects, but no personalised or explainable recommendation of which new project to try next.","pith_inferences":["Our inference: the balance-in-computing ratio is scale-free but denominator-sensitive; on a project with very few inherited volunteers, a single highly active volunteer can make the metric look strongly positive, so the aggregate result should be read alongside counts, not just ratios.","Our inference: the paper's design recommendation assumes that multi-project participation causes longer retention, but the data are correlational; a randomised trial that prompts some explorers with personalised recommendations and not others would test whether the recommendations themselves extend engagement.","Our inference: the extreme Gini inequality suggests an early-visibility advantage; if that dynamic holds, recommendation algorithms that favour new or small projects could be a testable way to flatten the distribution of attention."],"forward_implications":["Platform managers can treat cross-project movement as a retention lever: volunteers who regularly engage with multiple projects show longer relative activity duration than those who stay in one project, so encouraging exploration may keep people on the platform.","New projects should expect recruitment campaigns to bring many one-time helpers and should not mistake headcount for engagement; the smaller group of inherited volunteers is the more productive segment per person.","A concrete design gap is the absence of personalised, explainable project recommendations; filling it could raise the small fraction of volunteers who become multi-project regulars.","The Gini coefficients give platform managers a monitoring instrument: when a few projects capture nearly all recruitment and task execution, the platform may need to intervene to keep less visible projects viable."],"supporting_citations":[{"why":"Supplies the Goal-Question-Metric approach that structures the paper's quantitative characterisation.","marker":"[30]"},{"why":"Defines the engagement-profile view and the relative activity duration metric used to compare single- and multi-project volunteers.","marker":"[19]"},{"why":"Provides the operational definition of engagement (point, sustained, disengagement, re-engagement) that grounds the regular-versus-transient classification.","marker":"[17]"},{"why":"Documents how broadcast recruitment campaigns attract many curious, short-term volunteers, which the paper uses to explain why outside recruits do fewer tasks per person.","marker":"[25]"},{"why":"Defines the multi-project platform ecosystem (scientists, volunteers, platform) that the metrics are built to measure.","marker":"[21]"},{"why":"Introduces the Semiotic Inspection Method used to assess whether platforms communicate cross-project features to volunteers.","marker":"[7]"}],"fun_headline_variants":["Volunteers from other projects outwork outside recruits","Cross-project citizen science recruits outperform external ones","Citizen science: internal recruits do more tasks per person","Multi-project platforms lack personalized project suggestions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a volunteer was recruited by the first project they contributed to after registering; if volunteers browse or are directed to other projects before that first contribution, the split between 'inherited' and 'recruited outside' is misclassified, and the higher productivity of inherited volunteers could be an artifact of that misclassification.","fun_headline_variants_meta":{"raw":{"variants":["Volunteers from other projects outwork outside recruits","Cross-project citizen science recruits outperform external ones","Citizen science: internal recruits do more tasks per person","Multi-project platforms lack personalized project suggestions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000844,"raw_usage":{"total_tokens":3682,"prompt_tokens":961,"completion_tokens":2721,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":2671}},"tokens_in":577,"tokens_out":2721,"duration_ms":23201,"temperature":1.0,"reasoning_tokens":2671,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:15:16.621887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the balance-in-computing metric on a platform that logs referral sources or first page views before the first task. If volunteers who verifiably arrived from outside the platform perform as many or more tasks per person than volunteers who verifiably came from other projects, the paper's central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Goal-Question-Metric approach that structures the paper's quantitative characterisation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the engagement-profile view and the relative activity duration metric used to compare single- and multi-project volunteers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents how broadcast recruitment campaigns attract many curious, short-term volunteers, which the paper uses to explain why outside recruits do fewer tasks per person."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the multi-project platform ecosystem (scientists, volunteers, platform) that the metrics are built to measure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Semiotic Inspection Method used to assess whether platforms communicate cross-project features to volunteers."}],"review_version":1}