{"id":"be88e539-4ce2-4c96-99d6-f55d07d7a811","arxiv_id":"1908.08196","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Elite open source developers spend most of their effort on communication, organization, and support rather than coding, shift further away from coding as projects grow, and their non-coding effort is negatively correlated with the project's commit volume and bug outcomes.","lead":"This study maps the public GitHub activity of elite developers in 20 large open source projects into four activity types and finds that coding is a minority of their effort, that they shift toward communication and support as projects grow, and that more non-coding effort correlates with lower productivity and quality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Omitted time fixed effects in RQ3 panel regressions may confound the headline negative associations with project lifecycle trends.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the panel regressions omit time fixed effects even though the paper's own diagnostics show they are significant for the two central outcome models. My review of the full text confirms this. Section 3.6.3 presents equations (2)–(5) with project-specific fixed effects only, and Section 4.3.3 reports F-tests showing significant time effects for New Commit and New Bug models, yet the reported models in Tables 5 and 6 do not include time dummies. Given the RQ2 finding that effort shares trend over time, this is a textbook omitted-variable threat to the headline negative associations. The concern is concrete, testable, and directly affects the main empirical claim. It does not by itself invalidate the descriptive contributions (RQ1 and RQ2), and the paper is transparent about the correlational nature of RQ3, so a conditional verdict remains appropriate. No other issue outweighs this one: the elite/organizational circularity mainly affects RQ1 and the organizational share (a small component), while the time-trend confound cuts at the core of the negative-association finding. The proposed robustness test would settle the matter empirically.","tokens_in":27335,"tokens_out":3569,"duration_ms":37624,"concrete_test":"Re-estimate Model P1 (New Commits) and Model Q1 (New Bugs) from Tables 5 and 6, adding 37 month dummies (time fixed effects) to the existing LSDV project-specific fixed-effects specification. Compare the coefficients, standard errors, and p-values for S−Com, S−Org, and S−Sup with those reported. Also run a joint F-test on the month dummies. If the negative coefficients in P1 and the positive coefficients in Q1 remain significant and materially unchanged, the headline association is robust to common time shocks; if they shrink or lose significance, the omitted time trend is the likely driver and the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ3 claim rests on coefficients from panel regressions with project-specific fixed effects only (Section 3.6.3, Eqs. 2–5). Section 4.3.3 reports that time-specific effects are significant for New Commits (Model P1: F(38,662)=1.59, p=0.02) and New Bugs (Model Q1: F(38,662)=3.29, p<0.001). Yet Tables 5 and 6 contain no month dummies. RQ2 shows that elite effort shares trend over time: typical activity declines (~−1.63%/month) while communicative and supportive shares rise. If project maturation also reduces commit counts and changes bug-reporting rates, the omitted common time trend is correlated with both the regressors and the outcomes. The reported negative coefficients on S−Com and S−Sup in Model P1 and positive coefficients on S−Org and S−Sup in Model Q1 could therefore be artifacts of lifecycle trends rather than effort-allocation effects. This is the most load-bearing concern because the paper's own diagnostics flag the omitted time dimension, and the headline conclusion of negative associations with non-technical effort depends on those coefficients.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an empirical study of elite developers (developers with repository write permission) in 20 large open-source GitHub projects, using GHArchive and GitHub API event data from 2015 to 2018. The authors map raw events into four activity categories (communicative, organizational, supportive, typical) following Sonnentag's taxonomy, identify elite developers through a write-permission inference mechanism with a 90-day inactivity window, and analyze (RQ1) the distribution of elite activity across categories, (RQ2) the evolution of per-developer effort shares over project months using growth rates and ANOVA, and (RQ3) the association between elite effort shares and project outcomes (new commits, bug cycle time, new bugs, bug fix rate) using project-specific fixed-effects panel regressions. The headline findings are that elites dominate most activity categories (93% of organizational, 67% of typical), that their typical activities decline over time while communicative and supportive activities rise, and that non-technical effort shares are negatively associated with productivity and quality, with a positive association between supportive effort and bug fix rate.","tokens_in":27620,"tokens_out":5698,"duration_ms":54079,"significance":"If the results hold, the paper would provide a rare holistic view of elite developer work portfolios and useful evidence for effort-allocation guidance and automation tooling. The paper's strengths include a publicly released cleaned dataset, a structured card-sorting procedure with acceptable inter-rater agreement (kappa = 0.77), and the use of panel econometrics with fixed-effects and explicit diagnostics. However, the contribution is currently weakened by two load-bearing issues: the RQ1 organizational-activity share is partly definitional because elite status is inferred from write-permission-requiring events that largely coincide with the organizational category, and the RQ3 panel regressions omit time fixed effects despite the paper's own diagnostics showing significant time effects. The descriptive RQ1 and RQ2 findings are plausible and align with prior literature, but the central correlational claims of RQ3 require re-estimation before they can be accepted.","major_comments":[{"comment":"The claim that elite developers perform 93% of organizational activities (Table 4) is partly definitional: elite status is inferred from any action requiring write permission (Section 3.5), and most organizational event types in the taxonomy (assigning issues or PRs, requesting reviews, member/team events) require write permission by GitHubs design. The paper acknowledges this in Section 4.1: 'according to our definitions, most organizational events automatically require the write permission.' Consequently, the high elite share in the organizational category is a near-tautology and does not constitute independent evidence about elite developers' activity portfolio. Please report the fraction of organizational events that require write permission by construction, re-estimate RQ1 using an elite identification independent of the same event types (e.g., team membership or a contribution-threshold definition), and provide a robustness check that excludes the definitional overlap.","section":"Section 4.1 (Table 4) and Section 3.5"},{"comment":"The headline negative associations in Models P1 and Q1 may be artifacts of an omitted time trend. The panel regressions include only project-specific fixed effects; month dummies are not included, yet Section 4.3.3 reports that time-fixed effects are significant for the new-commit model (F(38,662)=1.59, p=0.02) and the new-bug model (F(38,662)=3.29, p<0.001). Because RQ2 establishes that elite effort shares trend over time (typical activities decline at -1.63% per month while communicative and supportive shares rise), and because project maturation plausibly affects commit counts and bug reports, the omitted common time trend is correlated with both the regressors and the outcomes. The statement that time effects are 'small' (adjusted R²=0.01) does not address omitted-variable bias, since even small confounders can substantially bias coefficients when they correlate with the regressors. Please add month fixed effects or a smooth time trend to Eqs. (2)-(5), report the resulting coefficients for S-Com, S-Org, and S-Sup, and conduct a formal test of whether the coefficients are stable. Without this re-analysis, the RQ3 conclusions are not supported.","section":"Section 3.6.3, Eqs. (2)-(5), and Section 4.3.3 (Tables 5 and 6)"}],"minor_comments":[{"comment":"The sentence 'For Model P2, where the bug cycle time ... the time-fixed effects model is not significant (F(38,662) = 1.80, p < 0.01)' is self-contradictory because p < 0.01 is significant; also, the later paragraph about the bug fix rate refers to 'Model P2' but should refer to Model Q2.","section":"Section 4.3.3"},{"comment":"The project name 'splitebrowser' should be 'sqlitebrowser'.","section":"Table 4"},{"comment":"The opening sentence refers to 'the two project productivity indicators' for the quality models; it should say 'project quality indicators'.","section":"Section 4.3.2"},{"comment":"The figure places PullRequestReviewComment and PullRequestReviewEvent under 'Typical', while Section 3.3.4 states that typical activities are 'counted as submitted commits and pull requests' and 'we only include commit activity under this category'; please reconcile these statements and clarify whether pull request review events are part of the typical category in the analysis.","section":"Figure 2 and Section 3.3.4"},{"comment":"The abstract's phrase 'technical contributions (e.g., coding) accounting for a small proportion only' is ambiguous because Table 4 shows elites perform 67% of typical activities; clarify that 'small proportion' refers to the share of elites' own effort distribution, not the project-wide share.","section":"Section 1 (Abstract)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an empirical software engineering journal. The main risk is the omitted time trend in RQ3; the authors have the data to re-run the models with month dummies, and the verdict should hinge on whether the coefficients survive. The RQ1 organizational-share result is unfalsifiable as presented and should be reframed with an independent elite identification. No concerns about data provenance or authorship conduct; the dataset availability is a positive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper gives the first comprehensive categorized view of what elite GitHub developers actually do, and the descriptive half is worth reading. Mapping 35 raw event types onto Sonnentag's four activity categories is a real contribution, and the card-sorting procedure (kappa 0.77) is reasonable. The longitudinal finding that coding effort per elite declines while communicative and supportive effort rises is clear, and the public data link is a plus.\n\nThe soft spots are real but not fatal. The elite definition (anyone performing a write-permission-requiring action) overlaps heavily with the organizational category, so the 93% organizational share is partly definitional — the authors admit this in Section 4.1. That inflates one descriptive statistic but does not break RQ1/RQ2. The bigger issue is RQ3. The panel regressions use project fixed effects only, and the paper's own diagnostics show time fixed effects are significant for new commits (p=0.02) and new bugs (p<0.001). Since elite effort shares trend over time and project outcomes also trend with maturity, the headline negative coefficients in Models P1 and Q1 could be contaminated by an omitted common time path. The paper argues the time effects are small (adjusted R2=0.01), which is a legitimate counter, but they never rerun the models with month dummies, so we cannot see how much the coefficients move. That is a genuine weakness, not just a hypothetical.\n\nWho gets value: software engineering researchers and OSS community managers who want a panoramic, quantified description of elite effort allocation. The taxonomy and dataset are the lasting contributions; the correlational claims should be treated as suggestive. Needs a serious referee, but the referee should ask for a robustness check with time fixed effects and for a cleaner separation of the elite-ship definition from the activity category definitions.","headline":"A fresh descriptive map of elite OSS developers' effort, with solid RQ1/RQ2 findings and a softer RQ3 that needs time fixed effects before the negative correlations are taken at face value.","tokens_in":28037,"tokens_out":2797,"would_cite":true,"duration_ms":31055,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using fine-grained event data from 20 large open source projects, this paper claims that elite developers' effort shifts from coding toward communication and support as projects mature, and that higher shares of such non-technical effort…","keywords":["elite developers","open source software","developer activity taxonomy","GitHub event data","effort allocation","panel regression","project productivity","software quality"],"falsifier":"Re-estimate Models P1 and Q1 with month fixed effects added alongside project fixed effects; if the negative coefficients on communicative and supportive shares and the positive coefficients on organizational and supportive shares disappear or change sign, the claimed associations are artifacts of project maturation rather than evidence about effort allocation.","tokens_in":27097,"feed_emoji":"🧑💻","tokens_out":10362,"duration_ms":88625,"temperature":0.7,"pith_summary":"Open source projects are run by a small set of elite developers, but little systematic evidence shows what those elites actually do with their time. This paper uses roughly 900,000 event records from 20 large GitHub projects to map every elite action into four categories — communicative, organizational, supportive, and typical (coding) — and track how those shares move over time. It reports three results: elite activity portfolios are dominated by non-coding work; as projects grow, elites shift further toward communication and support and away from coding; and higher shares of non-technical effort line up with lower productivity and quality in the same months, except that supportive effort is positively associated with the bug fix rate. The point of the study is to make effort allocation visible so that projects can decide how to support, automate, or redistribute the administrative burden that falls on elite developers.","feed_headline":"Elite open-source maintainers spend most effort beyond coding","feed_subtitle":"In 20 large projects, maintainers shift from coding to communication and support as the project grows, and commits drop.","key_machinery":"The load-bearing machinery is a four-category taxonomy of developer activity — communicative, organizational, supportive, and typical — adapted from prior field studies of software professionals and mapped onto 35 raw GitHub event types by closed card sorting, with 0.77 kappa agreement among the sorters. Elite status is identified dynamically: anyone observed performing an action that requires repository write permission is tagged as elite for a rolling 90-day window. The outcome analysis then runs LSDV (least-squares dummy variable) project fixed-effects panel regressions of four project-level indicators (new commits, bug cycle time, new bugs, bug fix rate) on the monthly shares of the three non-coding categories, with the coding share omitted because the four shares sum to one.","core_discovery":"The paper's central claim is that elite developers' work is not mainly coding: across 20 large projects, typical (code-writing) events are a small part of an elite developer's monthly activity, while communicative, organizational, and supportive events dominate, and their share grows as the project ages. Using project fixed-effects panel regressions on 720 project-months, the paper finds that when elites devote a larger share of effort to communicative or supportive activities, the project's new commit count in that month is lower; more organizational and supportive effort is associated with more newly reported bugs; yet more supportive effort is also associated with a higher bug fix rate. The authors read these results as showing that elite time and attention are finite resources, so non-technical duties crowd out technical contribution, while some supportive work genuinely helps the defect-removal process, and they are careful to label the findings as correlations rather than established causes.","pith_inferences":["Because the four effort shares sum to one, the regressions describe relative allocation only: a month with fewer commits mechanically has a higher non-coding share even if elites' absolute non-coding effort did not change. Re-running the analysis with per-elite activity counts, rather than shares, would separate 'communication crowds out coding' from 'coding fell for other reasons.'","The paper's own time-fixed-effect tests are significant for the new-commit and new-bug models, so an omitted project-lifecycle trend is a live rival explanation; estimating the same models with month dummies would show whether the negative coefficients survive.","A direct testable extension would follow a single project through an exogenous shock, such as a spike in issue inflow or a core maintainer going on leave, and compare months with high versus low elite communication after matching on bug volume.","The event stream covers only public platform activity; if elites move coordination to private channels as projects mature, the reported growth in communicative effort may be underestimated."],"forward_implications":["If the associations are real, project managers should expect commit throughput to fall as a maintainer takes on more communication and support work, and should plan staffing accordingly.","The positive link between supportive effort and bug fix rate suggests that labeling, documentation, branch management, and similar maintenance work is not pure overhead; it may be what keeps the defect-removal pipeline moving.","Company-sponsored projects show stronger associations in this study, so corporate governance practices that add communication overhead may be the first place to look for savings.","Automating routine organizational and supportive tasks (issue assignment, labeling, triage, release steps) would, under the paper's interpretation, free elite time for coding without giving up the quality benefits of support work.","Decentralizing administrative privileges could both reduce elite burden and give non-elite contributors more involvement, addressing the concentration of authority the paper documents."],"supporting_citations":[{"why":"It supplies the four-category activity taxonomy (communicative, organizational, supportive, typical) that the study maps GitHub events onto.","marker":"[72]"},{"why":"It supplies the write-permission-based method for identifying elite developers with a rolling 90-day eligibility window.","marker":"[38]"},{"why":"It supplies the productivity and quality outcome indicators and the keyword method for classifying bug issues.","marker":"[82]"},{"why":"It supplies the panel-data fixed-effects (LSDV) econometric framework used for all RQ3 regressions.","marker":"[91]"},{"why":"It establishes the empirical baseline that a small group of developers accounts for most contributions, which motivates the elite-developer focus.","marker":"[55]"},{"why":"It provides the chance-corrected agreement measure (kappa) used to validate the event-to-category card sorting.","marker":"[14]"}],"fun_headline_variants":["Elite devs: coding is only a fraction of their work","Maintainers' real job: talking, not typing","Open-source elite: communication over code","Coding takes a backseat for elite open-source devs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The panel regressions assume that, once fixed project differences are removed, the month-to-month variation in elite effort shares is not confounded by an underlying project-lifecycle trend; the paper's own time-fixed-effect tests are significant for the new-commit and new-bug models.","fun_headline_variants_meta":{"raw":{"variants":["Elite devs: coding is only a fraction of their work","Maintainers' real job: talking, not typing","Open-source elite: communication over code","Coding takes a backseat for elite open-source devs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000312,"raw_usage":{"total_tokens":1767,"prompt_tokens":929,"completion_tokens":838,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":772}},"tokens_in":545,"tokens_out":838,"duration_ms":7059,"temperature":1.0,"reasoning_tokens":772,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:46:24.684762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate Models P1 and Q1 with month fixed effects added alongside project fixed effects; if the negative coefficients on communicative and supportive shares and the positive coefficients on organizational and supportive shares disappear or change sign, the claimed associations are artifacts of project maturation rather than evidence about effort allocation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the four-category activity taxonomy (communicative, organizational, supportive, typical) that the study maps GitHub events onto."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the write-permission-based method for identifying elite developers with a rolling 90-day eligibility window."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the productivity and quality outcome indicators and the keyword method for classifying bug issues."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the panel-data fixed-effects (LSDV) econometric framework used for all RQ3 regressions."}],"review_version":1}