REVIEW 4 major objections 6 minor 44 references
Experimentation in Gaming: an Adoption Guide
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Experimentation can be systematically built into every stage of game development and marketing, and this guide maps when, where, and how.
desk verdict A clear, practitioner-focused adoption guide whose concrete pricing heuristics in Section 5 are the weakest link: stated as universal rules but backed only by experience and self-citation, in tension with the paper's own context caveats. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the lifecycle-by-marketing-mix matrix: four development and publishing stages (planning, development, pre-launch and soft launch, live operation) crossed with the four Ps (product, place, price, promotion), with each cell populated by a task such as marketability testing, first-time-user-experience design, economy balancing, LiveOps events, ad monetization, or price targeting. The matrix carries the argument by forcing a studio to decide, for every task, what kind of experiment applies, what sample size and randomization are feasible, and which team owns the learning. The second named mechanism is the Engagement Engineering framework, which draws on self-determination theory's three basic psychological needs—competence, relatedness, and autonomy—to generate engagement-driving innovations and to serve as an autonomy filter that screens out monetization designs, such as predatory lootboxes or mis-targeted offers, that trade long-term engagement for short-term revenue.
What would settle it
A cross-studio audit of free-to-play monetization experiments would settle the central heuristics: if bundles priced outside the 30–70% discount range, or country-level prices varying by more than a factor of two, consistently match or beat the recommended ranges on revenue and retention across a large, diverse sample of games, then the generalizability claim fails. The transferability heuristic could be tested directly by running the same experiment in a small game and a large game in the same portfolio and checking whether effect directions and sizes survive the change in audience.
Extended reading notes
Core claim
The central claim is that a game studio can treat experimentation as an end-to-end discipline by locating every key decision in a lifecycle-by-marketing-mix matrix and matching each cell to the right experimental mode. Before launch, experimentation is mostly qualitative—small samples, no randomization, expert and player feedback on prototypes, first-time-user experience, and balance—culminating in soft-launch tests that inform launch-or-kill decisions. After launch, it becomes quantitative at scale: A/B tests on LiveOps events, reward systems, ad placements, price personalization, and user acquisition, with long-term effects assessed through extended measurement windows, holdout groups, and predictive modeling. The author further claims that clear ownership structures make this work—product management and marketing as business owners, analytics as methods owner, with game designers looped in—and that monetization experiments can start from concrete experience-based ranges, such as bundle discounts of 30–70% and a maximum factor-of-two price corridor at country level. Underlying it all is the claim that experimentation and the art of game design are compatible: innovations should be designed to satisfy players' needs for competence, relatedness, and autonomy, and filtered to avoid designs that exploit vulnerable players.
Load-bearing premise
The guide's practical value rests on the assumption that the author's experience-based heuristics—the 30–70% bundle-discount range, the factor-of-two price corridor, and the transferability of learnings from small games to large ones—generalize across studios, business models, and player populations; if they do not, the concrete recommendations in the monetization and scaling sections would mislead rather than guide.
Editorial extensions
If this is right
- A studio can build a single experimentation roadmap spanning pre- and post-launch, with product, marketing, and analytics teams aligned on who owns each use case.
- Pre-launch effort concentrates on a few high-impact assumptions—drastically different first-time-user experiences, balancing scenarios, and personalization strategies—rather than precise effect sizes, which are deferred until sample sizes grow.
- LiveOps teams should test one thing at a time, keep a knowledge repository and learning agenda across event cycles, and calibrate observational models against high-validity A/B results instead of trying to measure everything every week.
- Monetization experiments can adopt concrete starting points—bundle discounts of 30–70% and a factor-of-two country-level price corridor—and refine them step by step with targeting experiments such as skimming, device-based, and recency-frequency-monetary-value personalization.
- Innovation that respects player autonomy, using the competence–relatedness–autonomy filter, is expected to produce higher long-term engagement and lower community backlash than short-term revenue tactics like exploitative lootboxes.
Reading between the lines
- The matrix and ownership model could transfer to other interactive entertainment and live-service verticals, such as social platforms or streaming services, where engaged communities, heterogeneity, and continuous content releases create the same experimental tensions.
- The experience-based monetization ranges read as testable hypotheses: a systematic search over a wider discount band or a larger price corridor, across multiple studios and genres, would reveal whether the recommended starting points are near-optimal or merely safe.
- The paper's gestures toward generative AI could be extended into a concrete workflow: using LLM-simulated player segments to pre-screen prototype and LiveOps hypotheses cheaply before committing scarce live-audience sample size.
- The autonomy filter could be operationalized as a measurable guardrail—tracking spending concentration, repeat-loss chasing, and offer-misfit rates as early warning indicators that a monetization experiment is eroding long-term engagement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a practitioner-oriented adoption guide for experimentation (A/B testing and related methods) across the game development lifecycle. It proposes a matrix of four development/publishing stages and the 4Ps of the marketing mix, then discusses pre-launch testing, soft launch, live operations, team ownership, pricing and personalization, ethical considerations, and an 'Engagement Engineering' framework for innovation. The central claim is that experimentation can be systematically embedded into game development and marketing decisions to improve engagement, retention, and monetization. The paper is explicitly a draft for feedback and relies heavily on the author's prior work and practitioner sources rather than on new empirical evidence.
Significance. If taken as a practical guide, the paper addresses a real gap: few sources map experimentation onto the full game lifecycle and the marketing mix in one place. Its proposed ownership matrix, emphasis on knowledge repositories, calibration of observational models with experimental results, and attention to LiveOps challenges are useful and actionable for industry readers. The paper also correctly identifies important gaming-specific issues such as community backlash, network interference, player heterogeneity, and novelty effects, and it recommends sensible tools such as switchback experiments, holdouts, and segment analysis. However, the scientific support for the most concrete prescriptions is thin: the numeric pricing and discount heuristics are unsupported, the innovation framework rests on a non-peer-reviewed source, and no empirical validation or comparison data is provided. The paper is best read as codified practitioner experience, and its significance for a research journal depends on whether the venue explicitly welcomes such bridge pieces.
major comments (4)
- [Section 5, pricing recommendations] The central prescriptive claims in Section 5 — a country-level price corridor of at most a factor of two, bundle discounts from 30% to a maximum 70%, and price points from $1.99 to $99.99 ($199.99 for long-tail games) — are presented as general guidance without any supporting dataset, meta-analysis, or comparative evidence. The cited references ([3], [19], [24]) are case studies in specific freemium mobile contexts, several by the author, and do not establish generalizable bounds. This conflicts with the paper's own caveats in Section 3 that publishing model, platform, audience size, and monetization design materially change experimentation, and with Section 6.3, which notes that player heterogeneity can reverse effects. The reader is never told over which game types, business models, or player populations these numbers apply. As written, the guide could mislead studios with different price sensitivity or monetization structures. The authors should either provide empirical evidence for these ranges or explicitly reframe them as experience-based, context-dependent starting hypotheses with clear boundary conditions.
- [Section 7, Engagement Engineering framework] The 'Engagement Engineering' framework [34] is a central pillar of the innovation guidance, yet it is cited to a non-peer-reviewed blog post by the author. The framework asserts that sustainable long-term engagement is achieved primarily through intrinsic motivation and that monetization should be optimized after engagement is established; this causal and normative ordering is not defended. Section 7's recommendations about dynamic difficulty adaptation, matchmaking, offer personalization, and lootboxes derive from this framework, so the load-bearing justification is thin. The authors should clarify the evidentiary status of the framework, provide any validation or independent empirical support, and be explicit when a recommendation follows from a design philosophy rather than from demonstrated causal evidence.
- [Sections 2, 3, and 5, conceptual matrix] The manuscript repeatedly refers to a 'matrix' as the backbone of its structure ('The following matrix will be the backbone for our further discussion') and also references a 'stylized representation of ownership structures' in Section 5, but no figures, tables, or visual diagrams appear in the submitted text. Since the paper's organizational and conceptual contribution is this matrix, its absence is a substantive gap: the reader cannot evaluate the proposed structure. The authors should include the missing figures/tables or, if the figures are only missing from the arXiv rendering, state this clearly and ensure the matrix is legible in the submitted version.
- [General, statistical rigor] The introduction states that a successful experimentation program needs 'centrally ensured rigor in experiment design and analysis,' but the article never discusses basic statistical requirements such as sample-size planning, minimum detectable effects, multiple-testing corrections, guardrail metrics, or pre-registration of hypotheses. Given that the guide aims to be comprehensive and that Sections 4 and 6 repeatedly emphasize reliable decision support, this omission weakens the practical value of the adoption guidance. At minimum, the authors should point readers to resources on these topics or qualify the guide's scope as non-technical.
minor comments (6)
- [Abstract] The phrase 'provides practical guidance to game makers how to adopt experimentation' should be 'provides practical guidance to game makers on how to adopt experimentation.'
- [Header and references] The running header contains the typo 'A PRIL' and reference [10] misspells 'Quantitative' as 'Quantative' in the journal name.
- [Section 6.2] The sentence 'e.g., via surrogates like this study in a freemium app context [28]' is awkwardly phrased; it should specify that [28] is a study on targeting for long-term outcomes and explain how it illustrates the surrogate approach.
- [References] Reference [32] is a 2016 working paper with no subsequent publication listed; if a peer-reviewed version exists, it should be cited instead, and if not, the current status should be made clear.
- [Section 8] The closing sentence 'Experimentation is here to stay and its future in gaming truly exciting' is informal and grammatically incomplete; it should be revised for a professional register.
- [Section 3] The reliance on a 2023 blog post [7] for the game development stages is appropriate for a practitioner context, but the simplification to four stages should be justified more explicitly, since the choice of stages directly shapes the matrix and therefore the rest of the paper.
Circularity Check
No significant circularity: the article is an experience-based adoption guide whose heuristics are not derived from fitted inputs or self-cited theorems.
full rationale
The paper is a practitioner-oriented guide, not a formal derivation. It makes no quantitative predictions and contains no equation chain in which an output is defined in terms of the very result it claims to produce. The self-citations (e.g., [3], [10], [24], [34]) are used as pointers to prior empirical studies or to the author's own Engagement Engineering framework, but the framework's psychological basis is explicitly attributed to an external source, self-determination theory [35]. None of the central prescriptions, such as the 30-70% bundle discount range or the factor-of-two price corridor, is claimed to follow from those citations; they are presented as experience-based heuristics. That they may be insufficiently supported or overgeneralized is a correctness or calibration concern, not circularity. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in through citation. The guide is self-contained in the sense that its claims stand or fall on the plausibility and applicability of its heuristics, not on a circular reduction to its own inputs.
Assumptions & free parameters
free parameters (3)
- Bundle discount range =
30% to 70%
- Bundle price points =
$1.99 to $199.99
- Price personalization corridor =
factor of two
assumptions (3)
- ad hoc to paper The four-stage game development process and 4Ps marketing mix provide an adequate organizing structure for experimentation decisions.
- domain assumption Strong long-term engagement is best achieved through intrinsic player motivation (competence, relatedness, autonomy).
- domain assumption Experimentation results generalize across games when learnings are transferred from small games to large games.
Cite this review
Pith. "Pith review of Experimentation in Gaming: an Adoption Guide." pith.science (2026). https://pith.science/paper/JM3TNSAA
@misc{pith2026250413840,
author = {Pith},
title = {Pith review of: Experimentation in Gaming: an Adoption Guide},
year = {2026},
howpublished = {\url{https://pith.science/paper/JM3TNSAA}},
note = {Machine review of arXiv:2504.13840}
}
read the original abstract
Experimentation is a cornerstone of successful game development and live operations, enabling teams to optimize player engagement, retention, and monetization. This article provides a comprehensive guide to implementing experimentation in gaming, structured around the game development lifecycle and the marketing mix. From pre-launch concept testing and prototyping to post-launch personalization and LiveOps, experimentation plays a pivotal role in driving innovation and adapting game experiences to diverse player preferences. Gaming presents unique challenges, such as highly engaged communities, complex interactive systems, and highly heterogeneous and evolving player behaviors, which require tailored approaches to experimentation. The article emphasizes the importance of collaborative frameworks across product, marketing, and analytics teams and provides practical guidance to game makers how to adopt experimentation successfully. It also addresses ethical considerations like fairness and player autonomy.
Reference graph
Works this paper leans on
-
[3]
Pricing Starter Packs: When Conversion is King, We May Price Too Low
Julian Runge. Pricing Starter Packs: When Conversion is King, We May Price Too Low. Deconstructor of Fun, 2024
work page 2024
-
[19]
Why Free-to-Play Apps Can Ignore the Old Rules About Cutting Prices
Sachin Waikar. Why Free-to-Play Apps Can Ignore the Old Rules About Cutting Prices. Stanford Insights, 2022
work page 2022
-
[24]
Julian Runge, Anders Drachen, and William Grosso. Exploratory Bandit Experiments with “Starter Packs” in a Free-to-Play Mobile Game. 2024 IEEE Conference on Games , 2024
work page 2024
-
[34]
Getting ML-Based App Personalization Right: The Engagement Engineering Framework
Julian Runge. Getting ML-Based App Personalization Right: The Engagement Engineering Framework. Mobile Dev Memo, 2023
work page 2023
-
[1]
Want Your Company to Get Better at Experimentation? Harvard Business Review, 2025
Iavor Bojinov, David Holtz, Ramesh Johari, Sven Schmit, and Martin Tingley. Want Your Company to Get Better at Experimentation? Harvard Business Review, 2025
work page 2025
-
[2]
Marketers Underuse Ad Experiments
Julian Runge. Marketers Underuse Ad Experiments. That’s a Big Mistake. Harvard Business Review, 2020
work page 2020
-
[4]
Algorithmic Assortative Matching on a Digital Social Medium
Kristian López Vargas, Julian Runge, and Ruizhi Zhang. Algorithmic Assortative Matching on a Digital Social Medium. Information Systems Research 33(4):1138-1156, 2022
work page 2022
-
[5]
List of Metrics you Need to Know to Effectively Execute LiveOps Strategies
Angad Singh. List of Metrics you Need to Know to Effectively Execute LiveOps Strategies. Segwise, 2024
work page 2024
Show all 44 references
-
[6]
How We Boosted App Revenue by 10% with Real-time Personalization
Julian Runge. How We Boosted App Revenue by 10% with Real-time Personalization. Medium, 2019
2019
-
[7]
Stages of Game Development | Your Guide On Game Development Process
Egor Piskunov. Stages of Game Development | Your Guide On Game Development Process. iLogos, 2023
2023
-
[8]
Neil H. Borden. The Concept of the Marketing Mix. Journal of Advertising Research, 1964
1964
-
[9]
Personalized Game Design for Improved User Retention and Monetization in Freemium Mobile Games
Eva Ascarza, Oded Netzer, and Julian Runge. Personalized Game Design for Improved User Retention and Monetization in Freemium Mobile Games. International Journal of Research in Marketing, forthcoming , 2025
2025
-
[10]
freemium
Julian Runge, Jonathan Levav, and Harikesh S. Nair. Price Promotions and "freemium" app monetization. Quantative Marketing Economics 20(2):101-139 , 2022
2022
-
[11]
Pre-Release Experimentation in Indie Game Development: An Interview Survey
Johan Linåker, Elizabeth Bjarnason, and Fabian Fagerholm. Pre-Release Experimentation in Indie Game Development: An Interview Survey. arXiv preprint arXiv:2411.17183, 2024
2024 arXiv
-
[12]
Experimentation in Early-Stage Video Game Startups: Practices and Challenges
Henry Edison, Jorge Melegati, and Elizabeth Bjarnason. Experimentation in Early-Stage Video Game Startups: Practices and Challenges. International Conference on Software Business , 2024
2024
-
[13]
Using LLMs for Market Research
James Brand, Ayelet Israeli, and Donald Ngwe. Using LLMs for Market Research. Harvard Business School Working Paper, 2024
2024
-
[14]
The Power of AI in Game Development (with Christoffer Holmgard)
Eric Seufert. The Power of AI in Game Development (with Christoffer Holmgard). Mobile Dev Memo Podcast, Season 4 Episode 7 , 2024. https://open.spotify.com/episode/6Yfba1FLuDekQkdfIjWc3O
2024
-
[15]
It’s time to close the experimentation gap in advertising: Confronting myths surrounding ad testing
Colin Campbell, Julian Runge, Kenneth Bates, Stacey Haefele, and Neeraj Jayaraman. It’s time to close the experimentation gap in advertising: Confronting myths surrounding ad testing. Business Horizons 65(4):437-446, 2022
2022
-
[16]
The Next Chapter of Supercell
Ilkka Paananen. The Next Chapter of Supercell. Supercell Blog – Ilkka’s Long Texts, 2023
2023
-
[17]
Combating Misinformation in Business Analytics: Experiment, Calibrate, Validate
Julian Runge and William Grosso. Combating Misinformation in Business Analytics: Experiment, Calibrate, Validate. Game Data Pros Blog, 2024
2024
-
[18]
A New Gold Standard for Digital Ad Measurement? Harvard Business Review, 2023
Julian Runge, Harpeet Patter, and Igor Skokan. A New Gold Standard for Digital Ad Measurement? Harvard Business Review, 2023
2023
-
[20]
Results of a massive experiment on virtual currency endowments and money demand
Nenad Živi´c, Igor Andjelkovi´c, Tolga Özden, Milovan Deki´c, and Edward Castronova. Results of a massive experiment on virtual currency endowments and money demand. PLOS one, 2017
2017
-
[21]
5 Ways to Ensure Data Quality When Running Experiments
Mikkel Dengsøe. 5 Ways to Ensure Data Quality When Running Experiments. Eppo Outperform, 2023
2023
-
[22]
What to Look For in an Experimentation Platform (Beyond Features)
Ryan Lucht. What to Look For in an Experimentation Platform (Beyond Features). Eppo Outperform, 2024
2024
-
[23]
How to Hackathon: Where Experimentation Mindset Meets Rapid Action
Aaron Silverman. How to Hackathon: Where Experimentation Mindset Meets Rapid Action. Eppo Outperform, 2024
2024
-
[25]
Churn Prediction for High-Value Players in Casual Social Games
Julian Runge, Peng Gao, Florent Garcin, and Boi Faltings. Churn Prediction for High-Value Players in Casual Social Games. IEEE Conference on Computational Intelligence and Games , 2014. 9 EXPERIMENTATION IN GAMING : AN ADOPTION GUIDE - A PRIL 22, 2025
2014
-
[26]
A Brief History of Revenue Optimization: Airline Fares and Price/Time-to-Flight Curves
William Grosso. A Brief History of Revenue Optimization: Airline Fares and Price/Time-to-Flight Curves. Game Data Pros Blog, 2024
2024
-
[27]
Customer Lifetime Value Prediction in Non-Contractual Freemium Settings: Chasing High-Value Users Using Deep Neural Networks and SMOTE
Rafet Sifa, Julian Runge, Christian Bauckhage, and Daniel Klapper. Customer Lifetime Value Prediction in Non-Contractual Freemium Settings: Chasing High-Value Users Using Deep Neural Networks and SMOTE. 51st Hawaii International Conference on System Sciences , 2018
2018
-
[28]
Targeting for Long-Term Outcomes.Management Science 70(6):3841-3855, 2023
Jeremy Yang, Dean Eckles, Paramveer Dhillon, and Sinan Aral. Targeting for Long-Term Outcomes.Management Science 70(6):3841-3855, 2023
2023
-
[29]
Treatment Effect Estimation Amidst Dynamic Network Interference in Online Gaming Experiments
Yu Zhu, Yang Su, Richard Li, and Zhenyu Zhao. Treatment Effect Estimation Amidst Dynamic Network Interference in Online Gaming Experiments. arXiv:2402.05336v1, 2024
2024 arXiv
-
[30]
Bandit vs
Sven Schmit. Bandit vs. Experiment Testing: Which Is Right for You? Eppo Outperform, 2023
2023
-
[31]
Optimizing Across the Free-to-Play Marketing Mix with Bandit Algorithms
Julian Runge and Mark Rieke. Optimizing Across the Free-to-Play Marketing Mix with Bandit Algorithms. Game Data Pros Blog, 2024
2024
-
[32]
Freemium Pricing: Evidence from a Large- scale Field Experiment
Julian Runge, Stefan Wagner, Joerg Claussen, and Daniel Klapper. Freemium Pricing: Evidence from a Large- scale Field Experiment. Humboldt University Berlin, School of Business and Economics, Institute of Marketing Working Paper, 2016
2016
-
[33]
Virtual Economies: Design and Analysis
Vili Lehdonvirta and Edward Castronova. Virtual Economies: Design and Analysis. The MIT Press, 2014
2014
-
[35]
Ryan and Edward L
Richard M. Ryan and Edward L. Deci. Self-determination theory: Basic psychological needs in motivation, development, and wellness. Guilford Press, 2017
2017
-
[36]
Dynamic Difficulty Adjustment (DDA) in Computer Games: A Review
Mohammad Zohaib. Dynamic Difficulty Adjustment (DDA) in Computer Games: A Review. Advances in Human-Computer Interaction, 2018
2018
-
[37]
Applications of Advanced Analytics to the Promotion of Freemium Goods
Julian Runge. Applications of Advanced Analytics to the Promotion of Freemium Goods. Doctoral dissertation at Humboldt University Berlin , 2020
2020
-
[38]
A Comparison of Methods for Player Clustering via Behavioral Telemetry
Anders Drachen, Rafet Sifa, Christian Thurau, and Christian Bauckhage. A Comparison of Methods for Player Clustering via Behavioral Telemetry. arXiv:1407.3950, 2014
2014 arXiv
-
[39]
Selling Bonus Actions in Video Games
Lifei Sheng, Xuying Zhao, and Christopher Thomas Ryan. Selling Bonus Actions in Video Games. Management Science, 2024
2024
-
[40]
Free-to-Play: Making Money From Games You Give Away (English Edition)
Will Luton. Free-to-Play: Making Money From Games You Give Away (English Edition). New Riders, 2013
2013
-
[41]
Optimal world design in video games.Working Paper, 2023
Yifu Li, Christopher Thomas Ryan, Lifei Sheng, and Benny Wong. Optimal world design in video games.Working Paper, 2023
2023
-
[42]
Optimal sequencing in single-player games
Yifu Li, Christopher Thomas Ryan, and Lifei Sheng. Optimal sequencing in single-player games. Management Science 69(10):6057-6075, 2023
2023
-
[43]
Personalized content, engagement, and monetization in a mobile puzzle game
Louis-Daniel Pape, Christian Helmers, Alessandro Iaria, Stefan Wagner, and Julian Runge. Personalized content, engagement, and monetization in a mobile puzzle game. International Journal of Industrial Organization , 2024
2024
-
[44]
How to Use Games to Build Relationships with Your Customers
Julian Runge and Joost van Dreunen. How to Use Games to Build Relationships with Your Customers. Harvard Business Review, 2024. 10
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.