{"id":"df50ce27-3ae6-4d25-9448-8c7dd803b1cb","arxiv_id":"2411.17183","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"An interview survey of ten indie game studios synthesizes a five-part framework for continuous experimentation before a game's release, chiefly through qualitative playtesting data.","lead":"Interviews with ten indie game developers show that pre-release experimentation centers on goals, design, experiment objects, sampling, and execution, and relies mostly on qualitative feedback like observations and playtests. The paper offers a framework that small game studios and other software teams can use when they have few users and limited data before launch.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The framework rests on one self-report per company from founder-level informants; the five parts may reflect a single role's perspective rather than company practice, and the interview guide is not shown, so instrument-driven categories cannot be excluded.","rationale":"The reader's weakest assumption is that one representative per company provides an accurate account of that company's experimentation practice. My concern is the same load-bearing point, sharpened: the informants are predominantly founders and leads, and the paper does not triangulate with other team members or with artifacts such as playtest recordings or design logs. This is not an attack on the authors' integrity; it is a standard validity threat for qualitative surveys. The paper's own limitation statement says generalization should be anecdotal, which is honest. However, the framework claims to describe company practice, and the evidence base is a single self-report per company. The proposed concrete test—re-interviewing a second, non-lead member of a subset of companies—directly targets whether the five parts are stable across roles and whether the framework is instrument-driven. This check is feasible with the existing sample and would either strengthen or weaken the central claim. Since the reader already reached CONDITIONAL based on the same fundamental issue, no verdict adjustment is needed; the condition is simply to validate with additional data. The paper is otherwise methodologically careful: dual coding, member-checking, audit trail, and an openly available codebook all support the trustworthiness of the analysis; the concern is about the inferential step from interviews to company-level practice, not about the conduct of the study.","tokens_in":20,"tokens_out":5308,"duration_ms":115860,"concrete_test":"Select 3–5 of the original ten companies and conduct a second interview with a different team member (e.g., a programmer, artist, or QA tester) using the same semi-structured guide, asking for concrete recent examples. Independently code the new interviews without reference to the published codebook, then compare whether the same five-part framework emerges. If the second informant's account does not reproduce all five parts, or introduces new key parts, the single-informant design is insufficient and the framework is likely role- or instrument-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that pre-release indie-game experimentation is organized around the five key parts of the framework, and the evidence consists of 10 interviews, one per company, conducted online for about 30 minutes each (Section 3). The interviewees are overwhelmingly founders, CEOs, or lead developers (Table 1), i.e., the people who would design experiments, not necessarily those who execute or experience them. For team-based practices, a single retrospective self-report cannot validate company-level behaviour; the paper asserts 'One representative per indie company was deemed sufficient' (Section 3) without any triangulation, and it even includes a 50+ employee company (I7) where one informant is clearly not enough. This threatens internal validity: the five categories could capture the mental model of a founder rather than the actual distributed practice. Additionally, the semi-structured guide is not quoted in the manuscript, and the codebook is only in supplementary material, so the degree to which the interview questions pre-structured the five parts cannot be assessed from the paper. Even a representative sample would not fix this if the framework is an artifact of the instrument or of the informant's role.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory qualitative interview survey of ten indie game developers, one per company, about pre-release experimentation practices. Through open and axial coding, the authors synthesize an emerging continuous experimentation framework with five key parts: goal definition, design strategy, experiment object, sampling strategy, and execution strategy. They report that pre-release experimentation is centered on qualitative data, with playtesting and observation as dominant methods, and that resource constraints and limited access to participants shape sampling and execution. The paper positions the framework as a first step to be validated in future case studies.","tokens_in":11463,"tokens_out":5771,"duration_ms":52988,"significance":"If the framework holds, it fills a gap in continuous experimentation research by characterizing experimentation in indie game development, a context with limited user data and resources. The study's methodological reporting is a strength: pairwise coding with merged codebooks, peer review of coding, audit trail, saturation noted by the tenth interview, member checking of an early article version, and a supplementary codebook. The framework is explicitly exploratory and the authors include a candid limitations paragraph. The significance is nevertheless bounded by the sample (10 companies, mostly Swedish, mostly desktop), the single-informant design, and the lack of triangulation; the result is best read as a hypothesis-generating synthesis rather than a validated company-level account.","major_comments":[{"comment":"The paper asserts that one representative per indie company was deemed sufficient to understand practice at these 'very small companies,' but the sample includes I7 with 50+ employees, and all informants are founders, CEOs, or lead developers. Because the framework describes company-level experimentation practice, a single retrospective self-report from one role cannot validate distributed practices, and the five parts may capture founder mental models rather than what the team actually does. Section 4.2 further attributes I7's mature A/B testing descriptions to experience from larger mobile game companies, so that interview may not describe current practice at the sampled indie company. Please either narrow the claims to individual-level perceptions, add triangulation for larger companies, or justify the single-informant assumption for each company.","section":"Section 3, Table 1"},{"comment":"The semi-structured interview questionnaire is not quoted or included in the manuscript, and the final codebook appears only in the supplementary material. Since the framework's five parts are derived from coded interview responses, the absence of the instrument makes it impossible to assess whether categories such as 'design strategy' or 'execution strategy' were prompted by the question wording rather than emergent from the data. Please include the interview guide (or at least the core questions) in an appendix and indicate how each framework part maps to the questions that elicited the supporting quotes.","section":"Section 3 and supplementary material [15]"},{"comment":"The limitations paragraph appropriately cautions that the sample is small and mostly Swedish and that generalization should be anecdotal. However, several Discussion statements generalize beyond this evidence base, for example the claim that split testing is often performed in the initial idea stage and the assertion that the results 'may be of value also for larger game companies, and for software intensive organisations in other industries.' Given the explicitly emerging status of the framework, these inferences should be presented as hypotheses for future validation rather than as conclusions.","section":"Section 5"}],"minor_comments":[{"comment":"The number of employees is written with both a comma and a period ('2,5' and '1.5'), and 'Rougelite' should be 'Roguelite.'","section":"Table 1"},{"comment":"The sentence 'Representatives from ten indie company were identified' should read 'ten indie companies.'","section":"Section 3"},{"comment":"The word 'perofmred' should be 'performed' in the sentence about the cadence with which experiments are performed.","section":"Section 4.2"},{"comment":"The phrase 'present their game ideas to publishers early one' should be 'early on.'","section":"Section 4.3"},{"comment":"Please ensure that Figure 1 is legible in print and in grayscale, since it is central to understanding the proposed framework.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The study fits the journal's scope and the framework is plausible, but the single-informant design and the opaque instrument make the company-level framing vulnerable. I would not require new data collection for revision; the framework can stand as an emerging model if the claims are carefully scoped and the interview guide is made inspectable. The I7 case, with its 50+ employees and externally sourced A/B testing experience, should be handled explicitly in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a careful, modest exploratory study that gives indie game pre-release experimentation a workable vocabulary. The new bit is the five-part framework — goal definition, design strategy, experiment object, sampling strategy, execution strategy — built inductively from ten interviews, plus the observation that pre-release experimentation in this setting is essentially qualitative. The emphasis on understandability as an evaluation aspect alongside mechanics, aesthetics, and fun factor is a useful addition. The method is transparent: conference sampling, semi-structured interviews, pairwise coding with a merged codebook, member checking, and an explicit note about saturation. The authors are honest that this is an emerging framework, not a validated one.\n\nSoft spots are the usual ones for interview surveys at this scale. One representative per company, mostly founder-level, interviewed for about 30 minutes, gives a picture of the founder's mental model more than distributed team practice. The sample skews Swedish and was recruited at a single conference. I7, with 50+ employees and experience at larger mobile studios, sits awkwardly in an 'indie' sample. The interview guide is not in the paper, so the degree to which the five parts were pre-structured by the instrument is not fully checkable from the text; the codebook is supplementary, which helps. None of this is fatal for an exploratory qualitative study, and the authors already frame the result as something to refine with cases and larger companies.\n\nIn proportion: this is a useful, well-reported qualitative contribution to the continuous experimentation and game software engineering literature. It doesn't overclaim. It deserves a serious referee. I'd ask the authors to make the interview guide available in the supplementary material and to be more explicit about the single-informant limitation, but I wouldn't hold the paper hostage to those.\n\nRecommendation: send to peer review. It's the kind of paper that gets cited for its framework in future CE-in-games work.","headline":"Useful inductive framework for indie pre-release experimentation, honestly reported, but the single-informant sample and hidden interview guide mean the five parts are provisional.","tokens_in":11988,"tokens_out":1927,"would_cite":true,"duration_ms":17866,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"From ten indie-studio interviews, this paper proposes an emerging continuous-experimentation framework with five parts, and claims that pre-release experimentation is centered on qualitative data from observation and playtesting.","keywords":["continuous experimentation","indie game development","game user research","playtesting","qualitative data","pre-release development","interview survey","game development framework"],"falsifier":"A replication with a larger, geographically varied sample of indie studios, coded without prior knowledge of the five categories, would falsify the framework if the categories did not emerge or if most studios reported quantitative telemetry rather than qualitative playtesting as their primary pre-release data source.","tokens_in":11073,"feed_emoji":"🎮","tokens_out":6920,"duration_ms":58503,"temperature":0.7,"pith_summary":"This paper claims that pre-release experimentation in indie game development can be captured in a single emerging framework with five parts: goal definition, design strategy, experiment object, sampling strategy, and execution strategy. Based on interviews with ten indie developers, it argues that before a game is released, experimentation is centered on qualitative data, with observation of play and playtesting doing most of the work, rather than on large-scale quantitative telemetry. The paper also finds that time and resource limits constrain how many participants indie studios can reach, and that these studios manage bias and representativeness with small, carefully chosen samples. A sympathetic reader would care because most software experimentation frameworks assume abundant user data, whereas indie studios must validate game ideas with very little; the proposed framework describes how experimentation can still function under that constraint and may extend to larger game companies and other software-intensive organizations.","feed_headline":"Indie pre-release experiments run on playtests, not big data","feed_subtitle":"Interviews with ten studios yield a five-part framework: goal, design, object, sampling, execution.","key_machinery":"The load-bearing object is the five-part CE framework for game development itself. It functions as a classification and planning vocabulary: each part names a decision an indie studio must make when running an experiment, and the interview data supplies the options observed for each part, such as split/sequential/exploratory testing for design strategy and internal team/friends and family/publishers/external unknowns/communities/streamers for sampling. The framework's role is to turn scattered qualitative interview findings into a reusable structure that can be checked against future cases.","core_discovery":"The central discovery is an emerging continuous-experimentation framework for game development, derived from ten interviews. The framework says that any pre-release experiment in an indie studio can be planned and understood through five key parts: goal definition combines the experiment's purpose (scoping a game idea or a feature) with the game aspect being evaluated (aesthetics, mechanics, fun factor, or understandability); design strategy chooses among split, sequential, and exploratory testing; experiment object selects the medium (sketches, video, playable minimum viable game) and refinement level (conceptual, functional, content); sampling strategy picks participant demographics and numbers while weighing bias and representativeness; and execution strategy collects feedback mainly through observation, screen recordings, surveys, and occasional interviews. Across these parts, the study finds that pre-release experimentation runs primarily on qualitative data from a small number of participants, with split testing concentrated in early ideation and prototyping, and sequential and exploratory testing taking over as development matures.","pith_inferences":["The five parts amount to a domain-specific restatement of a general experimental-design checklist; an implicit next step would be to test whether the same five parts organize pre-release experimentation in non-game software startups with scarce user data.","The interviews suggest an implicit maturity path in sampling: studios move from internal teams and friends/family toward dedicated communities and external unknowns; a testable extension would be whether studios that reach community and streamer samples make better feature decisions.","Because observation and screen recordings carry most of the evidence, cheap remote-playtest tooling that records sessions and annotates events could lower the cost barrier and make sequential experimentation more systematic in small studios.","A hypothesis-driven comparison is left untested: the data describe exploratory open-minded playtesting as common, but do not establish whether explicitly writing hypotheses before playtests improves the quality of decisions; that would be a natural next experiment."],"forward_implications":["Pre-release indie experimentation is primarily a qualitative activity: observing people play and recording sessions gives useful feedback even with 8-10 participants, so studios need not wait for large user bases to validate ideas.","Using the five parts as a checklist can help a studio see which decision is weak; for example, a playtest with unclear goals or an unrepresentative sample can be diagnosed as a sampling or goal-definition problem.","Split (A/B) testing is most affordable in early ideation and prototyping, while sequential and exploratory testing fit later stages, which reverses the common assumption that A/B testing belongs mainly to mature products.","Since pre-release experimentation is centered on qualitative data, quantitative telemetry remains a later-stage or mobile-game practice whose tooling costs currently block small studios.","If the framework holds, larger game companies and other software-intensive organizations with limited early user data can adopt the same five-part structure for pre-release experimentation."],"supporting_citations":[{"why":"Supplies the definition of continuous experimentation and hypothesis-driven evaluation that frames the whole study.","marker":"[8]"},{"why":"Prior work showing qualitative methods and playtesting dominate early-stage game development, which this paper extends to indie pre-release practice.","marker":"[24]"},{"why":"Closest prior study of experimentation in early-stage video game startups, used as a starting point for the indie focus.","marker":"[6]"},{"why":"Links playtesting and A/B testing to development stages and minimum viable games, positioning the design-strategy findings.","marker":"[11]"},{"why":"Source for the game aspects of mechanics, dynamics, and aesthetics that the goal-definition part extends with fun factor and understandability.","marker":"[13]"},{"why":"Supports the contrast that quantitative game analytics dominate freemium and mobile contexts, unlike the qualitative indie practice observed here.","marker":"[14]"},{"why":"The empirical standard for qualitative interview surveys that defines the study's research design.","marker":"[17]"},{"why":"Provides the open and axial coding method used to synthesize interview themes into the framework.","marker":"[20]"}],"fun_headline_variants":["Indie games: playtests, not big data, guide pre-release experiments","Five-part framework for pre-release indie game experimentation","Indie devs lean on qualitative playtests for pre-release experiments","Pre-release indie experiments run on small-scale playtest feedback","How indie studios test games pre-release: a playtest-driven framework"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework rests on the assumption that one roughly 30-minute online interview per company, with developers recruited at a single indie-game conference and mostly based in Sweden, gives an accurate picture of how that studio experiments before release.","fun_headline_variants_meta":{"raw":{"variants":["Indie games: playtests, not big data, guide pre-release experiments","Five-part framework for pre-release indie game experimentation","Indie devs lean on qualitative playtests for pre-release experiments","Pre-release indie experiments run on small-scale playtest feedback","How indie studios test games pre-release: a playtest-driven framework"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1247,"prompt_tokens":956,"completion_tokens":291,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":204}},"tokens_in":572,"tokens_out":291,"duration_ms":3456,"temperature":1.0,"reasoning_tokens":204,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:23:34.111076+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with a larger, geographically varied sample of indie studios, coded without prior knowledge of the five categories, would falsify the framework if the categories did not emerge or if most studios reported quantitative telemetry rather than qualitative playtesting as their primary pre-release data source.","supporting_citations":[{"cited_title":"Journal of Systems and Software 123, 292–305 (2017)","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of continuous experimentation and hypothesis-driven evaluation that frames the whole study."},{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"Prior work showing qualitative methods and playtesting dominate early-stage game development, which this paper extends to indie pre-release practice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior study of experimentation in early-stage video game startups, used as a starting point for the indie focus."},{"cited_title":"In: 17th IFIP WG 6.11 Conf","cited_arxiv_id":null,"evidence_quote":"Links playtesting and A/B testing to development stages and minimum viable games, positioning the design-strategy findings."},{"cited_title":"In: Software Business: 6th Int","cited_arxiv_id":null,"evidence_quote":"Source for the game aspects of mechanics, dynamics, and aesthetics that the goal-definition part extends with fun factor and understandability."},{"cited_title":"on e-Business, e- Services, and e-Society","cited_arxiv_id":null,"evidence_quote":"Supports the contrast that quantitative game analytics dominate freemium and mobile contexts, unlike the qualitative indie practice observed here."},{"cited_title":"Sage (2021)","cited_arxiv_id":null,"evidence_quote":"Provides the open and axial coding method used to synthesize interview themes into the framework."}],"review_version":1}