Pith. sign in

REVIEW 2 major objections 5 minor 93 references

Unveiling Elite Developers' Activities in Open Source Projects

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Using fine-grained event data from 20 large open source projects, this paper claims that elite developers' effort shifts from coding toward communication and support as projects mature, and that higher shares of such non-technical effort…

desk verdict A fresh descriptive map of elite OSS developers' effort, with solid RQ1/RQ2 findings and a softer RQ3 that needs time fixed effects before the negative correlations are taken at face value. read the letter →

arxiv 1908.08196 v2 pith:A5NNIMV6 submitted 2019-08-22 cs.SE

classification cs.SE
keywords elitedevelopersopensourcesoftwaredeveloperactivitytaxonomyGitHubeventdataeffortallocationpanelregressionprojectproductivityquality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Open source projects are run by a small set of elite developers, but little systematic evidence shows what those elites actually do with their time. This paper uses roughly 900,000 event records from 20 large GitHub projects to map every elite action into four categories — communicative, organizational, supportive, and typical (coding) — and track how those shares move over time. It reports three results: elite activity portfolios are dominated by non-coding work; as projects grow, elites shift further toward communication and support and away from coding; and higher shares of non-technical effort line up with lower productivity and quality in the same months, except that supportive effort is positively associated with the bug fix rate. The point of the study is to make effort allocation visible so that projects can decide how to support, automate, or redistribute the administrative burden that falls on elite developers.

What carries the argument

The load-bearing machinery is a four-category taxonomy of developer activity — communicative, organizational, supportive, and typical — adapted from prior field studies of software professionals and mapped onto 35 raw GitHub event types by closed card sorting, with 0.77 kappa agreement among the sorters. Elite status is identified dynamically: anyone observed performing an action that requires repository write permission is tagged as elite for a rolling 90-day window. The outcome analysis then runs LSDV (least-squares dummy variable) project fixed-effects panel regressions of four project-level indicators (new commits, bug cycle time, new bugs, bug fix rate) on the monthly shares of the three non-coding categories, with the coding share omitted because the four shares sum to one.

What would settle it

Re-estimate Models P1 and Q1 with month fixed effects added alongside project fixed effects; if the negative coefficients on communicative and supportive shares and the positive coefficients on organizational and supportive shares disappear or change sign, the claimed associations are artifacts of project maturation rather than evidence about effort allocation.

Watch

Extended reading notes

Core claim

The paper's central claim is that elite developers' work is not mainly coding: across 20 large projects, typical (code-writing) events are a small part of an elite developer's monthly activity, while communicative, organizational, and supportive events dominate, and their share grows as the project ages. Using project fixed-effects panel regressions on 720 project-months, the paper finds that when elites devote a larger share of effort to communicative or supportive activities, the project's new commit count in that month is lower; more organizational and supportive effort is associated with more newly reported bugs; yet more supportive effort is also associated with a higher bug fix rate. The authors read these results as showing that elite time and attention are finite resources, so non-technical duties crowd out technical contribution, while some supportive work genuinely helps the defect-removal process, and they are careful to label the findings as correlations rather than established causes.

Load-bearing premise

The panel regressions assume that, once fixed project differences are removed, the month-to-month variation in elite effort shares is not confounded by an underlying project-lifecycle trend; the paper's own time-fixed-effect tests are significant for the new-commit and new-bug models.

Editorial extensions

If this is right

  • If the associations are real, project managers should expect commit throughput to fall as a maintainer takes on more communication and support work, and should plan staffing accordingly.
  • The positive link between supportive effort and bug fix rate suggests that labeling, documentation, branch management, and similar maintenance work is not pure overhead; it may be what keeps the defect-removal pipeline moving.
  • Company-sponsored projects show stronger associations in this study, so corporate governance practices that add communication overhead may be the first place to look for savings.
  • Automating routine organizational and supportive tasks (issue assignment, labeling, triage, release steps) would, under the paper's interpretation, free elite time for coding without giving up the quality benefits of support work.
  • Decentralizing administrative privileges could both reduce elite burden and give non-elite contributors more involvement, addressing the concentration of authority the paper documents.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the four effort shares sum to one, the regressions describe relative allocation only: a month with fewer commits mechanically has a higher non-coding share even if elites' absolute non-coding effort did not change. Re-running the analysis with per-elite activity counts, rather than shares, would separate 'communication crowds out coding' from 'coding fell for other reasons.'
  • The paper's own time-fixed-effect tests are significant for the new-commit and new-bug models, so an omitted project-lifecycle trend is a live rival explanation; estimating the same models with month dummies would show whether the negative coefficients survive.
  • A direct testable extension would follow a single project through an exogenous shock, such as a spike in issue inflow or a core maintainer going on leave, and compare months with high versus low elite communication after matching on bug volume.
  • The event stream covers only public platform activity; if elites move coordination to private channels as projects mature, the reported growth in communicative effort may be underestimated.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper presents an empirical study of elite developers (developers with repository write permission) in 20 large open-source GitHub projects, using GHArchive and GitHub API event data from 2015 to 2018. The authors map raw events into four activity categories (communicative, organizational, supportive, typical) following Sonnentag's taxonomy, identify elite developers through a write-permission inference mechanism with a 90-day inactivity window, and analyze (RQ1) the distribution of elite activity across categories, (RQ2) the evolution of per-developer effort shares over project months using growth rates and ANOVA, and (RQ3) the association between elite effort shares and project outcomes (new commits, bug cycle time, new bugs, bug fix rate) using project-specific fixed-effects panel regressions. The headline findings are that elites dominate most activity categories (93% of organizational, 67% of typical), that their typical activities decline over time while communicative and supportive activities rise, and that non-technical effort shares are negatively associated with productivity and quality, with a positive association between supportive effort and bug fix rate.

Significance. If the results hold, the paper would provide a rare holistic view of elite developer work portfolios and useful evidence for effort-allocation guidance and automation tooling. The paper's strengths include a publicly released cleaned dataset, a structured card-sorting procedure with acceptable inter-rater agreement (kappa = 0.77), and the use of panel econometrics with fixed-effects and explicit diagnostics. However, the contribution is currently weakened by two load-bearing issues: the RQ1 organizational-activity share is partly definitional because elite status is inferred from write-permission-requiring events that largely coincide with the organizational category, and the RQ3 panel regressions omit time fixed effects despite the paper's own diagnostics showing significant time effects. The descriptive RQ1 and RQ2 findings are plausible and align with prior literature, but the central correlational claims of RQ3 require re-estimation before they can be accepted.

major comments (2)
  1. [Section 4.1 (Table 4) and Section 3.5] The claim that elite developers perform 93% of organizational activities (Table 4) is partly definitional: elite status is inferred from any action requiring write permission (Section 3.5), and most organizational event types in the taxonomy (assigning issues or PRs, requesting reviews, member/team events) require write permission by GitHubs design. The paper acknowledges this in Section 4.1: 'according to our definitions, most organizational events automatically require the write permission.' Consequently, the high elite share in the organizational category is a near-tautology and does not constitute independent evidence about elite developers' activity portfolio. Please report the fraction of organizational events that require write permission by construction, re-estimate RQ1 using an elite identification independent of the same event types (e.g., team membership or a contribution-threshold definition), and provide a robustness check that excludes the definitional overlap.
  2. [Section 3.6.3, Eqs. (2)-(5), and Section 4.3.3 (Tables 5 and 6)] The headline negative associations in Models P1 and Q1 may be artifacts of an omitted time trend. The panel regressions include only project-specific fixed effects; month dummies are not included, yet Section 4.3.3 reports that time-fixed effects are significant for the new-commit model (F(38,662)=1.59, p=0.02) and the new-bug model (F(38,662)=3.29, p<0.001). Because RQ2 establishes that elite effort shares trend over time (typical activities decline at -1.63% per month while communicative and supportive shares rise), and because project maturation plausibly affects commit counts and bug reports, the omitted common time trend is correlated with both the regressors and the outcomes. The statement that time effects are 'small' (adjusted R²=0.01) does not address omitted-variable bias, since even small confounders can substantially bias coefficients when they correlate with the regressors. Please add month fixed effects or a smooth time trend to Eqs. (2)-(5), report the resulting coefficients for S-Com, S-Org, and S-Sup, and conduct a formal test of whether the coefficients are stable. Without this re-analysis, the RQ3 conclusions are not supported.
minor comments (5)
  1. [Section 4.3.3] The sentence 'For Model P2, where the bug cycle time ... the time-fixed effects model is not significant (F(38,662) = 1.80, p < 0.01)' is self-contradictory because p < 0.01 is significant; also, the later paragraph about the bug fix rate refers to 'Model P2' but should refer to Model Q2.
  2. [Table 4] The project name 'splitebrowser' should be 'sqlitebrowser'.
  3. [Section 4.3.2] The opening sentence refers to 'the two project productivity indicators' for the quality models; it should say 'project quality indicators'.
  4. [Figure 2 and Section 3.3.4] The figure places PullRequestReviewComment and PullRequestReviewEvent under 'Typical', while Section 3.3.4 states that typical activities are 'counted as submitted commits and pull requests' and 'we only include commit activity under this category'; please reconcile these statements and clarify whether pull request review events are part of the typical category in the analysis.
  5. [Section 1 (Abstract)] The abstract's phrase 'technical contributions (e.g., coding) accounting for a small proportion only' is ambiguous because Table 4 shows elites perform 67% of typical activities; clarify that 'small proportion' refers to the share of elites' own effort distribution, not the project-wide share.

Circularity Check

2 steps flagged · score 6.0 of 10

Two construction overlaps make the elite share of organizational events and the supportive-effort/bug-fix-rate association partly tautological; the main negative RQ3 correlations remain empirical.

  1. self definitional [Section 3.5 (elite identification); Section 4.1 and Table 4 (RQ1 organizational share)]
    "When a developer in the repository performs a task that requires the write permission, we tag this developer with 'elite-ship' of the repository. ... Besides organizational events (according to our definitions, most organizational events automatically require the write permission), elite developers perform over 60% of supportive activities and even created 34% of communicative activities."

    The elite set is defined by performing write-permission-requiring actions, and the organizational category (Fig. 2) consists largely of such actions (member management, assigning issues/PRs, managing reviewers). Consequently the 93% mean elite share of organizational events in Table 4 is entailed by the identification rule: by construction, non-elites cannot perform most organizational events. The paper presents this as an empirical finding about elite responsibility, but it is a property of the operationalization.

  2. self definitional [Section 3.4 (BF R definition); Fig. 2 taxonomy (Supportive/Issue Event); Table 6 Model Q2]
    "Supportive ... Issue and PR ... IssueDuplicateEvent, IssueRenameEvent, Merge Event, Issue(Un)LockEvent, Issue Event (Fig. 2); BF Rim = No. of Fixed Bugs / No. of Found Bugs, for project i in month m (Section 3.4)."

    Closing a bug report is an IssueEvent, so it is counted in the supportive share S-Sup; the same closure also increments the numerator of the dependent variable BF R (fixed bugs). In Model Q2, the positive coefficient of S-Sup on BF R (1.12***) is therefore partly an identity: one action is recorded on both sides of the regression. The paper's interpretation that supportive effort improves the defect-removal process is contaminated by this definitional overlap.

full rationale

The paper's central negative RQ3 findings for New Commits, Bug Cycle Time, and New Bugs come from panel regressions whose coefficients are not forced by definition; those results are independent empirical correlations, so the paper is not wholly circular. However, two load-bearing results do reduce by construction. First, the RQ1 claim that elites perform 93% of organizational events follows from defining elites as actors who perform write-permission-requiring actions while defining most organizational events as write-permission-requiring; the paper itself concedes this. Second, the Bug Fix Rate outcome overlaps with the supportive activity category through IssueEvent/closing events, making the positive S-Sup coefficient in Model Q2 partly an accounting identity. The omitted time fixed effects concern noted in Section 4.3.3 is a validity threat, not a circularity, and no self-citation chain is load-bearing here; the score reflects these two definitional overlaps while recognizing that the main negative association claims retain independent empirical content.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The central empirical claims rest on a set of domain assumptions about GitHub event coverage, the transferability of Sonnentag's taxonomy, the permission-based identification of elites, and the regression specifications. There are no invented entities and only two hand-chosen thresholds (90-day elite window and bug keywords). The most consequential assumption is that the elite definition and the organizational category do not collide; the paper itself notes that they overlap.

free parameters (2)
  • Elite-ship inactivity window = 90 days
    Hand-chosen threshold used to renew or expire elite status; authors state 30 days was too short and 90 days avoids rush decisions. Changes the set of developers counted as elite.
  • Bug keyword list = defect, error, bug, issue, mistake, incorrect, fault, flaw
    Hand-chosen keyword set (adapted from Vasilescu et al. 2015) used to classify issues as bugs from title or tags; no per-project validation.
assumptions (6)
  • domain assumption GitHub public event data captures the relevant universe of elite developer activities across all four categories.
    Section 3.2 and Section 5.5 acknowledge that private channels (email, IRC, instant messages) are excluded, yet the analysis treats GitHub events as the complete activity record.
  • domain assumption Sonnentag's four-category taxonomy (communicative, organizational, supportive, typical) transfers from in-house software developers to open source developers.
    Section 3.3 adopts and modifies the taxonomy with card sorting (kappa=0.77); the validity of the category semantics across OSS settings is assumed.
  • domain assumption Performing any write-permission-requiring event marks a developer as elite, and a 90-day activity window captures current elite status.
    Section 3.5: elite identity is inferred from observed privileged events, not from repository permission lists, which GitHub does not expose.
  • domain assumption Fixed-effects panel regressions without time fixed effects are sufficient to identify associations between effort shares and outcomes.
    Section 3.6.3 and Section 4.3.3: time effects are significant for the new-commit and new-bug models but are omitted from reported models; the authors argue the effects are small, which is a modeling choice rather than a proven fact.
  • domain assumption Issues whose title or tags contain one of eight keywords are bugs.
    Section 3.4: keyword matching adapted from Vasilescu et al. is used without project-specific validation.
  • standard math LSDV fixed-effects estimation and ANOVA are appropriate for the panel structure.
    Section 3.6.3: standard econometric assumptions are invoked; diagnostics are run but not reported in full.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unveiling Elite Developers' Activities in Open Source Projects." pith.science (2026). https://pith.science/paper/A5NNIMV6

@misc{pith2026190808196,
  author       = {Pith},
  title        = {Pith review of: Unveiling Elite Developers' Activities in Open Source Projects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A5NNIMV6}},
  note         = {Machine review of arXiv:1908.08196}
}
read the original abstract

Open-source developers, particularly the elite developers, maintain a diverse portfolio of contributing activities. They do not only commit source code but also spend a significant amount of effort on other communicative, organizational, and supportive activities. However, almost all prior research focuses on a limited number of specific activities and fails to analyze elite developers' activities in a comprehensive way. To bridge this gap, we conduct an empirical study with fine-grained event data from 20 large open-source projects hosted on GitHub. Thus, we investigate elite developers' contributing activities and their impacts on project outcomes. Our analyses reveal three key findings: (1) they participate in a variety of activities while technical contributions (e.g., coding) accounting for a small proportion only; (2) they tend to put more effort into supportive and communicative activities and less effort into coding as the project grows; (3) their participation in non-technical activities is negatively associated with the project's outcomes in term of productivity and software quality. These results provide a panoramic view of elite developers' activities and can inform an individual's decision making about effort allocation, thus leading to finer project outcomes. The results also provide implications for supporting these elite developers.

Figures

Figures reproduced from arXiv: 1908.08196 by the authors.

Figure 1
Figure 1. Data collection and cleanup process. and author more easily, and derive necessary metrics for the later data analysis on the project productivity, we also download the commit logs of all sampled projects. In total, we have collected 1.81 GB data of issue events and commit logs. Finally, we use Python scripts to merge event data based on event ID and commit SHA, and clean the redundant data that were recorded on both… view at source ↗
Figure 2
Figure 2. The taxonomy of GitHub event types. The definition of each raw event can be found in the official GitHub Events API documentation page: https://developer.github.com/v3/activity/events/. the setting of distributed software development where open-source project usually employs, each project applies various communication channels including mailing list, instant message, and online discussion board [7]. Moreover, some p… view at source ↗
Figure 3
Figure 3. The distributions of elite devel￾opers’ activity shares in each activity category over 20 projects. In addition to elites’ code submission, we also found em￾pirical evidence that elite developers are also “responsible” for most other types of events. Besides organizational events (according to our definitions, most organizational events au￾tomatically require the write permission), elite developers perform over 60% … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Average monthly activities comparisons between elite and non-elite developers. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Trends of individual elite developer’s activities in the four activity categories of the [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Changes of time-related effects. Elite developers’ effort distributions have significant correlations with project outcomes. (1) Project Productivity: (a) Efforts on communicative and supportive activities are negatively correlated with the project productivity in term…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 62 canonical work pages

  1. [1]

    Mark Aberdour. 2007. Achieving quality in open-source software. IEEE Software 24, 1 (2007), 58–64

  2. [2]

    Gonzalez-Barahona

    Juan Jose Amor, Gregorio Robles, and Jesus M. Gonzalez-Barahona. 2006. Effort Estimation by Characterizing Developer Activity. In Proceedings of the 2006 International Workshop on Economics Driven Software Engineering Research (EDSER ’06). ACM, New York, NY, USA, 3–6. https://doi.org/10.1145/1139113.1139116

  3. [3]

    John Anvik, Lyndon Hiew, and Gail C. Murphy. 2006. Who Should Fix This Bug?. InProceedings of the 28th International Conference on Software Engineering (ICSE ’06) . ACM, New York, NY, USA, 361–370. https://doi.org/10.1145/1134285. 1134336

  4. [4]

    Blake Ashforth. 2000. Role transitions in organizational life: An identity-based perspective . Routledge

  5. [5]

    Sogol Balali, Igor Steinmacher, Umayal Annamalai, Anita Sarma, and Marco Aurelio Gerosa. 2018. Newcomers’ Barriers. . . Is That All? An Analysis of Mentors’ and Newcomers’ Barriers in OSS Projects. Computer Supported Cooperative Work (CSCW) 27, 3 (01 Dec 2018), 679–714. https://doi.org/10.1007/s10606-018-9310-8

  6. [6]

    Sebastian Baltes and Stephan Diehl. 2018. Towards a theory of software development expertise. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ACM, 187–200

  7. [7]

    Christian Bird. 2011. Sociotechnical coordination and collaboration in open source software. In 2011 27th IEEE International Conference on Software Maintenance (ICSM) . IEEE, 568–573

  8. [8]

    Christian Bird, Nachiappan Nagappan, Brendan Murphy, Harald Gall, and Premkumar Devanbu. 2011. Don’t touch my code!: examining the effects of ownership on software quality. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering . ACM, 4–14

Show all 93 references
  1. [9]

    Christian Bird, David Pattison, Raissa D’Souza, Vladimir Filkov, and Premkumar Devanbu. 2008. Latent Social Structure in Open Source Projects. In Proceedings of the 16th ACM SIGSOFT International Symposium on Foundations of Software Engineering (SIGSOFT ’08/FSE-16) . ACM, New ...

  2. [10]

    Christian Bird, Peter C Rigby, Earl T Barr, David J Hamilton, Daniel M German, and Prem Devanbu. 2009. The promises and perils of mining git. In 2009 6th IEEE International Working Conference on Mining Software Repositories . IEEE, 1–10

  3. [11]

    Tegawendé F Bissyandé, David Lo, Lingxiao Jiang, Laurent Réveillere, Jacques Klein, and Yves Le Traon. 2013. Got issues? who cares about it? a large scale investigation of issue trackers from github. In 2013 IEEE 24th international symposium on software reliability engineering...

  4. [12]

    Kenneth S Bordens and Bruce B Abbott. 2002. Research Design and Methods: A Process Approach . McGraw-Hill

  5. [13]

    danah boyd and Kate Crawford. 2012. Critical questions for big data: Provocations for a cultural, technological, and scholarly phenomenon. Information, Communication & Society 15 (01 2012), 662–679

  6. [14]

    Robert L Brennan and Dale J Prediger. 1981. Coefficient kappa: Some uses, misuses, and alternatives. Educational and psychological measurement 41, 3 (1981), 687–699

  7. [15]

    Felix C Brodbeck. 1994. Software-Entwicklung: Ein Tätigkeitsspektrum mit vielfältigen Kommunikations-und Lernan- forderungen. na

  8. [16]

    Gerardo Canfora, Massimiliano Di Penta, Rocco Oliveto, and Sebastiano Panichella. 2012. Who is Going to Mentor Newcomers in Open Source Projects?. In Proceedings of the ACM SIGSOFT 20th International Symposium on the Foundations of Software Engineering (FSE ’12) . ACM, New Yor...

  9. [17]

    Oscar Chaparro, Jing Lu, Fiorella Zampetti, Laura Moreno, Massimiliano Di Penta, Andrian Marcus, Gabriele Bavota, and Vincent Ng. 2017. Detecting Missing Information in Bug Descriptions. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering (ESEC...

  10. [18]

    Benjamin Collier, Moira Burke, Niki Kittur, and Robert E Kraut. [n.d.]. Promoting Good Management: Governance, Promotion, and Leadership in Open Collaboration Communities.. In Proceedings of the 2010 International Conference on Information Systems (ICIS’10). 220

  11. [19]

    Josh Cowls and Ralph Schroeder. 2015. Causation, Correlation, and Big Data in Social Science Research. Policy & Internet 7 (08 2015), n/a–n/a. https://doi.org/10.1002/poi3.100

  12. [20]

    Yves Croissant and Giovanni Millo. 2008. Panel Data Econometrics in R: The plm Package. Journal of Statistical Software 27, 2 (2008), 1–43. https://doi.org/10.18637/jss.v027.i02

  13. [21]

    Kevin Crowston, Hala Annabi, James Howison, and Chengetai Masango. 2004. Effective Work Practices for Software En- gineering: Free/Libre Open Source Software Development. InProceedings of the 2004 ACM Workshop on Interdisciplinary Software Engineering Research (WISER ’04) . AC...

  14. [22]

    Kevin Crowston and James Howison. 2005. The social structure of free and open source software development. First Monday 10, 2 (2005)

  15. [23]

    Kevin Crowston and James Howison. 2006. Assessing the health of open source communities. Computer 39, 5 (2006), 89–91

  16. [24]

    Kevin Crowston, Kangning Wei, James Howison, and Andrea Wiggins. 2008. Free/Libre Open-source Software Development: What We Know and What We Do Not Know. ACM Comput. Surv. 44, 2, Article 7 (March 2008), 35 pages. https://doi.org/10.1145/2089125.2089127

  17. [25]

    Kevin Crowston, Kangning Wei, Qing Li, and James Howison. 2006. Core and periphery in free/libre and open source software team communications. In Proceedings of the 39th Annual Hawaii International Conference on System Sciences (HICSS’06), Vol. 6. IEEE, 118a–118a

  18. [26]

    Daniel Alencar da Costa, Uirá Kulesza, Eduardo Aranha, and Roberta Coelho. 2014. Unveiling Developers Contributions Behind Code Commits: An Exploratory Study. InProceedings of the 29th Annual ACM Symposium on Applied Computing (SAC ’14). ACM, New York, NY, USA, 1152–1157. http...

  19. [27]

    Laura Dabbish, Colleen Stuart, Jason Tsay, and Jim Herbsleb. 2012. Social coding in GitHub: transparency and collaboration in an open software repository. In Proceedings of the ACM 2012 conference on computer supported cooperative work. ACM, 1277–1286

  20. [28]

    Barthélémy Dagenais, Harold Ossher, Rachel K. E. Bellamy, Martin P. Robillard, and Jacqueline P. de Vries. 2010. Moving into a New Software Project Landscape. In Proceedings of the 32Nd ACM/IEEE International Conference on Software Engineering - Volume 1 (ICSE ’10) . ACM, New ...

  21. [29]

    Martyn Denscombe. 2014. The Good Research Guide: For Small-scale Social Research Projects . McGraw-Hill Education (UK)

  22. [30]

    Luis Felipe Dias, Igor Steinmacher, and Gustavo Pinto. 2018. Who drives company-owned OSS projects: internal or external members? Journal of the Brazilian Computer Society 24, 1 (2018), 16

  23. [31]

    Dino Distefano, Manuel Fähndrich, Francesco Logozzo, and Peter W. O’Hearn. 2019. Scaling Static Analyses at Facebook. Commun. ACM 62, 8 (July 2019), 62–70. https://doi.org/10.1145/3338112

  24. [32]

    Nicolas Ducheneaut. 2005. Socialization in an open source software community: A socio-technical analysis. Computer Supported Cooperative Work (CSCW) 14, 4 (2005), 323–368. ACM Trans. Softw. Eng. Methodol., Vol. 28, No. 1, Article 1. Publication date: January 2019. Unveiling El...

  25. [33]

    Liran Einav and Jonathan Levin. 2014. Economics in the age of big data. Science 346 (11 2014), 1243089. https: //doi.org/10.1126/science.1243089

  26. [34]

    Kristin E Flegal and Michael C Anderson. 2008. Overthinking skilled motor performance: Or why those who teach canâĂŹt do. Psychonomic Bulletin & Review 15, 5 (2008), 927–932

  27. [35]

    Matt Germonprez, Julie E Kendall, Kenneth E Kendall, Lars Mathiassen, Brett Young, and Brian Warner. 2016. A theory of responsive design: A field study of corporate engagement with open source communities. Information Systems Research 28, 1 (2016), 64–83

  28. [36]

    Georgios Gousios, Margaret-Anne Storey, and Alberto Bacchelli. 2016. Work practices and challenges in pull-based development: the contributor’s perspective. In Proceedings of the 38th IEEE/ACM International Conference on Software Engineering (ICSE ’16) . IEEE, 285–296

  29. [37]

    Philip J Guo, Thomas Zimmermann, Nachiappan Nagappan, and Brendan Murphy. 2011. Not my bug! and other reasons for software bug report reassignments. In Proceedings of the ACM 2011 conference on Computer supported cooperative work. ACM, 395–404

  30. [38]

    Marvin Hanisch, Carolin Haeussler, Stefan Berreiter, and Sven Apel. 2018. Developers’ Progression from Periphery to Core in the Linux Kernel Development Project. In Academy of Management Proceedings , Vol. 2018. Academy of Management Briarcliff Manor, NY 10510, 14263

  31. [39]

    James Howison and Kevin Crowston. 2014. Collaboration through open superposition: a theory of the open source way. Management Information Systems Quarterly 38, 1 (2014), 29–50

  32. [40]

    Federico Iannacci. 2005. Coordination processes in open source software development: The Linux case study.Emergence: Complexity & Organization 7, 2 (2005)

  33. [41]

    Chris Jensen and Walt Scacchi. 2007. Role Migration and Advancement Processes in OSSD Projects: A Comparative Case Study. In Proceedings of the 29th International Conference on Software Engineering (ICSE ’07) . IEEE Computer Society, Washington, DC, USA, 364–374. https://doi.o...

  34. [42]

    Corey Jergensen, Anita Sarma, and Patrick Wagstrom. 2011. The onion patch: migration in open source ecosystems. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering (FSE’11). ACM, 70–80

  35. [43]

    Mitchell Joblin, Sven Apel, Claus Hunsen, and Wolfgang Mauerer. 2017. Classifying Developers into Core and Peripheral: An Empirical Study on Count and Network Metrics. In Proceedings of the 39th International Conference on Software Engineering (ICSE ’17) . IEEE Press, Piscataw...

  36. [44]

    Eirini Kalliamvakou, Georgios Gousios, Kelly Blincoe, Leif Singer, Daniel M German, and Daniela Damian. 2014. The promises and perils of mining GitHub. In Proceedings of the 11th working conference on mining software repositories (MSR). ACM, 92–101

  37. [45]

    Foutse Khomh, Tejinder Dhaliwal, Ying Zou, and Bram Adams. 2012. Do faster releases improve software quality?: an empirical case study of Mozilla Firefox. In Proceedings of the 9th IEEE Working Conference on Mining Software Repositories. IEEE Press, 179–188

  38. [46]

    Dongsun Kim, Yida Tao, Sunghun Kim, and Andreas Zeller. 2013. Where should we fix this bug? a two-phase recommendation model. IEEE Transactions on Software Engineering 39, 11 (2013), 1597–1610

  39. [47]

    Sunghun Kim and E James Whitehead Jr. 2006. How long did it take to fix bugs?. InProceedings of the 2006 international workshop on Mining software repositories . ACM, 173–174

  40. [48]

    Chakravanti Rajagopalachari Kothari. 2004. Research Methodology: Methods and Techniques . New Age International

  41. [49]

    LaToza, Gina Venolia, and Robert DeLine

    Thomas D. LaToza, Gina Venolia, and Robert DeLine. 2006. Maintaining Mental Models: A Study of Developer Work Habits. In Proceedings of the 28th International Conference on Software Engineering (ICSE ’06) . ACM, New York, NY, USA, 492–501. https://doi.org/10.1145/1134285.1134355

  42. [50]

    Josh Lerner and Jean Tirole. 2002. Some simple economics of open source. The journal of industrial economics 50, 2 (2002), 197–234

  43. [51]

    Ytzhak Levendel. 1990. Reliability analysis of large software systems: Defect data modeling. IEEE Transactions on Software Engineering 16, 2 (1990), 141–152

  44. [52]

    Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A Diversity-Promoting Objective Function for Neural Conversation Models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...

  45. [53]

    Bin Lin, Gregorio Robles, and Alexander Serebrenik. 2017. Developer turnover in global, industrial open source projects: Insights from applying survival analysis. In 2017 IEEE 12th International Conference on Global Software Engineering (ICGSE). IEEE, 66–75

  46. [54]

    DV Luciv, DV Koznov, George A Chernishev, Andrey N Terekhov, K Yu Romanovsky, and DA Grigoriev. 2018. Detecting near duplicates in software documentation. Programming and Computer Software 44, 5 (2018), 335–343

  47. [55]

    Fielding, and James D

    Audris Mockus, Roy T. Fielding, and James D. Herbsleb. 2002. Two Case Studies of Open Source Software Development: Apache and Mozilla. ACM Trans. Softw. Eng. Methodol. 11, 3 (July 2002), 309–346. https://doi.org/10.1145/567793.567795 ACM Trans. Softw. Eng. Methodol., Vol. 28, ...

  48. [56]

    Nigel Nicholson. 1984. A theory of work role transitions. Administrative science quarterly (1984), 172–191

  49. [57]

    Siobhan O’Mahony and Fabrizio Ferraro. 2007. The emergence of governance in an open source community. Academy of Management Journal 50, 5 (2007), 1079–1106

  50. [58]

    Brian T Pentland and Martha S Feldman. 2005. Organizational routines as a unit of analysis. Industrial and Corporate Change 14, 5 (2005), 793–815

  51. [59]

    Huilian Sophie Qiu, Alexander Nolte, Anita Brown, A Serebrenik, and Bogdan Vasilescu. 2018. Going Farther Together: The Impact of Social Capital on Sustained Participation in Open Source. In International Conference on Software Engineering. IEEE Computer Society

  52. [60]

    R Development Core Team. 2008. R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria. http://www.R-project.org ISBN 3-900051-07-0

  53. [61]

    Foyzur Rahman and Premkumar Devanbu. 2011. Ownership, experience and defects: a fine-grained study of authorship. In Proceedings of the 33rd International Conference on Software Engineering . ACM, 491–500

  54. [62]

    Baishakhi Ray, Daryl Posnett, Vladimir Filkov, and Premkumar Devanbu. 2014. A large scale study of programming languages and code quality in github. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering. ACM, 155–165

  55. [63]

    Eric Raymond. 1999. The cathedral and the bazaar. Knowledge, Technology & Policy 12, 3 (1999), 23–49

  56. [64]

    Rigby, Daniel M

    Peter C. Rigby, Daniel M. German, and Margaret-Anne Storey. 2008. Open Source Software Peer Review Practices: A Case Study of the Apache Server. In Proceedings of the 30th International Conference on Software Engineering (ICSE ’08) . ACM, New York, NY, USA, 541–550. https://do...

  57. [65]

    Jeffrey A Roberts, Il-Horn Hann, and Sandra A Slaughter. 2006. Understanding the motivations, participation, and performance of open source software developers: A longitudinal study of the Apache projects. Management science 52, 7 (2006), 984–999

  58. [66]

    Bertil Rolandsson, Magnus Bergquist, and Jan Ljungberg. 2011. Open source in the firm: Opening up professional practices of software development. Research Policy 40, 4 (2011), 576–587

  59. [67]

    Mike Savage and Roger Burrows. 2007. The Coming Crisis of Empirical Sociology. Sociology 41 (10 2007). https: //doi.org/10.1177/0038038507080443

  60. [68]

    Walt Scacchi. 2007. Free/Open Source Software Development. InProceedings of the the 6th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineering (ESEC- FSE ’07). ACM, New York, NY, USA, 459–468. http...

  61. [69]

    Mario Schaarschmidt, Gianfranco Walsh, and Harald FO von Kortzfleisch. 2015. How do firms influence open source software communities? A framework and empirical analysis of different governance modes. Information and Organization 25, 2 (2015), 99–114

  62. [70]

    Sonali K Shah. 2006. Motivation, governance, and the viability of hybrid forms in open source software development. Management science 52, 7 (2006), 1000–1014

  63. [71]

    Emad Shihab, Akinori Ihara, Yasutaka Kamei, Walid M Ibrahim, Masao Ohira, Bram Adams, Ahmed E Hassan, and Ken-ichi Matsumoto. 2013. Studying re-opened bugs in open source software. Empirical Software Engineering 18, 5 (2013), 1005–1042

  64. [72]

    Sabine Sonnentag. 1995. Excellent software professionals: Experience, work activities, and perception by peers. Behaviour & Information Technology 14, 5 (1995), 289–299

  65. [73]

    Sabine Sonnentag. 1998. Expertise in professional software design: A process study. Journal of applied psychology 83, 5 (1998), 703

  66. [74]

    Igor Steinmacher, Tayana Conte, Marco Aurélio Gerosa, and David Redmiles. 2015. Social barriers faced by newcomers placing their first contribution in open source software projects. InProceedings of the 18th ACM conference on Computer supported cooperative work & social comput...

  67. [75]

    Margaret-Anne Storey. 2019. Publish or Perish: Questioning the Impact of Our Research on the Software Developer. In Proceedings of the 41st International Conference on Software Engineering: Companion Proceedings (ICSE ’19) . IEEE Press, Piscataway, NJ, USA, 2–2. https://doi.or...

  68. [76]

    Fei Liu Jeffrey Flanigan Sam Thomson and Norman Sadeh Noah A Smith. 2015. Toward Abstractive Summarization Using Semantic Representations. , 1077-1086 pages

  69. [77]

    Marat Valiev, Bogdan Vasilescu, and James Herbsleb. 2018. Ecosystem-level determinants of sustained activity in open-source projects: a case study of the PyPI ecosystem. InProceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium ...

  70. [78]

    Perry van Wesel, Bin Lin, Gregorio Robles, and Alexander Serebrenik. 2017. Reviewing career paths of the openstack developers. In Proceedings of the 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME ’17) . IEEE, 544–548

  71. [79]

    Bogdan Vasilescu, Kelly Blincoe, Qi Xuan, Casey Casalnuovo, Daniela Damian, Premkumar Devanbu, and Vladimir Filkov. 2016. The sky is not the limit: multitasking across GitHub projects. In 2016 IEEE/ACM 38th International ACM Trans. Softw. Eng. Methodol., Vol. 28, No. 1, Articl...

  72. [80]

    Bogdan Vasilescu, Vladimir Filkov, and Alexander Serebrenik. 2013. Stackoverflow and github: Associations between software development and crowdsourced knowledge. In 2013 International Conference on Social Computing . IEEE, 188–195

  73. [81]

    Bogdan Vasilescu, Daryl Posnett, Baishakhi Ray, Mark GJ van den Brand, Alexander Serebrenik, Premkumar Devanbu, and Vladimir Filkov. 2015. Gender and tenure diversity in GitHub teams. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (CHI ’...

  74. [82]

    Bogdan Vasilescu, Yue Yu, Huaimin Wang, Premkumar Devanbu, and Vladimir Filkov. 2015. Quality and productivity outcomes relating to continuous integration in GitHub. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering (FSE ’15) . ACM, 805–816

  75. [83]

    Georg Von Krogh and Eric Von Hippel. 2006. The promise of research on open source software. Management science 52, 7 (2006), 975–983

  76. [84]

    Patrick Wagstrom. 2009. Vertical interaction in open software engineering communities . PhD dissertation. Carnegie Mellon University

  77. [85]

    Patrick Wagstrom, Corey Jergensen, and Anita Sarma. 2012. Roles in a networked software development ecosystem: A case study in GitHub. (2012)

  78. [86]

    Wasserstein and Nicole A

    Ronald L. Wasserstein and Nicole A. Lazar. 2016. The ASA’s Statement on p-Values: Context, Process, and Purpose. The American Statistician 70, 2 (2016), 129–133

  79. [87]

    Weiss, R

    C. Weiss, R. Premraj, T. Zimmermann, and A. Zeller. 2007. How Long Will It Take to Fix This Bug?. InFourth International Workshop on Mining Software Repositories (MSR’07:ICSE Workshops 2007) . 1–1. https://doi.org/10.1109/MSR.2007.13

  80. [88]

    E. F. Weller. 2000. Practical applications of statistical process control [in software development projects].IEEE Software 17, 3 (May 2000), 48–55. https://doi.org/10.1109/52.896249

  81. [89]

    Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015. Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processin...

  82. [90]

    Terry Winograd, Fernando Flores, and Fernando F Flores. 1986. Understanding computers and cognition: A new foundation for design. Intellect Books

  83. [91]

    Jeffrey M Wooldridge. 2015. Introductory Econometrics: A Modern Approach . Nelson Education

  84. [92]

    Judy L Wynekoop and Diane B Walz. 2000. Investigating traits of top performing software developers. Information Technology & People 13, 3 (2000), 186–195

  85. [93]

    Daniel Bärl Torsten Zesch and Iryna Gurevych. 2012. Text Reuse Detection Using a Composition of Text Similarity Measures. In Proceedings of the 24th International Conference on Computational Linguistics (COLING’212), Vol. 1. Citeseer, 167–184. ACM Trans. Softw. Eng. Methodol.,...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.