{"id":"c45743ea-fe59-443c-845a-be5f3dd57ae9","arxiv_id":"1908.02586","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A regression model combined with simulations of random giving identifies ingroup favouritism and reciprocity in token-exchange games, and tracks how these behaviors strengthen over rounds.","lead":"This paper presents a statistical method for analyzing who gives tokens to whom in online social experiments, where standard models fail because people's choices are not independent. The method tests whether behaviors like favoring your own group or returning favors are real, and it tracks how these behaviors strengthen over time.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The time-dependent reciprocity claim is confounded by the use of cumulative outcome variables; early non-significance does not establish that reciprocity strengthens over time.","rationale":"The reader's weakest assumption (the uniform random null model) is a valid concern about the baseline for significance tests, but it does not capture the most load-bearing flaw I see. My identified concern is more specific: the paper's key temporal claim about reciprocity rests on cumulative coefficient trajectories that are confounded with power and cumulative aggregation. This does not overturn the reader's conditional verdict, but it sharpens the condition: the time-dependent reciprocity claim needs a per-round lagged analysis before it can be accepted. The ingroup favouritism claim is comparatively robust because it appears early and is supported by descriptive proportions, although even that should be checked at the earliest rounds with an appropriate model. I therefore agree partially with the reader: both are methodological concerns about the inference, but mine focuses on the temporal interpretation rather than the null model. The proposed concrete test would settle whether the reciprocity trend is real or an artefact.","tokens_in":13709,"tokens_out":11885,"duration_ms":139876,"concrete_test":"Reanalyse the raw data (or a simulated dataset with known constant per-round reciprocity) using a per-round logistic regression of whether i gives a token to j in round t, with predictors: an indicator that j gave to i in any earlier round, the group indicator G_ij, and an interaction between the lagged-receipt indicator and round number t. If the interaction coefficient is not significantly positive (or if the cumulative ρ_t trajectory from a simulation with constant per-round reciprocity reproduces the observed Fig 10 pattern), then the 'becomes more pronounced' claim is not supported. This test directly separates the cumulative-aggregation artefact from a genuinely strengthening behavioural tendency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim includes that reciprocity 'becomes more pronounced as the game progresses.' The evidence for this is the trajectory of the cumulative reciprocity coefficient ρ_t in Section 4.3 and Fig 10. But in Eq (3.2), Y_ijt is the cumulative number of tokens i received from j up to round t, and Y_jit is the cumulative number i gave to j up to round t. At small t, these cumulative counts are small and mostly zero, so any per-round reciprocity effect has very low statistical power; fitting the model at each t and observing that ρ_t is non-significant early and significant later is exactly what one would expect under a constant per-round reciprocity effect as information accumulates. Thus the increase in ρ_t over time does not, by itself, demonstrate that reciprocal behaviour strengthens; it may simply reflect the growing sample size within each cumulative sum. Moreover, because both Y_ijt and Y_jit are contemporaneous cumulative counts, ρ_t captures a symmetrical association (i and j giving more to each other over the whole game) rather than a temporal response where a prior gift increases the probability of a later gift. This makes the 'it takes time for a reciprocal relationship to develop' interpretation unsupported. The ingroup favouritism part of the claim is better supported by the early and persistent group effect, but the temporal differentiation between ingroup and reciprocity effects in the Discussion rests on this cumulative-coefficient comparison. The paper does not fit a per-round model with lagged predictors that would directly measure how reciprocity changes with time.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a regression-based modelling framework for temporal network data arising from the VIAPPL social-interaction platform. The response variable is the cumulative number of tokens player i has received from player j up to round t, regressed on the cumulative number of tokens i has given to j and on a same-group indicator, after removing self-giving. Inference is based on comparing observed regression coefficients with coefficients fitted to many simulations from a null model in which every player gives one token per round uniformly at random. The method is applied to four 14-player games. The authors report that ingroup favouritism is present from early rounds, that reciprocity becomes stronger over time, that one game is dominated by two unusually reciprocal players, and they introduce an influence metric and network visualisation for identifying players whose behaviour departs from the norm.","tokens_in":13950,"tokens_out":5165,"duration_ms":58080,"significance":"The proposed framework is a reasonable and useful step for analysing individual-level interaction data in controlled experiments, where standard independence assumptions fail. The use of a parameter-free null model with 10,000 simulations is a strength, as is the explicit treatment of self-giving and the model-based influence metric. If the substantive conclusions were fully supported, the paper would provide a valuable methodology for VIAPPL and similar platforms. However, the central temporal claim about reciprocity is not supported by the cumulative-response model as presented, and the significance statements are not adjusted for the large number of tests performed. The group-favouritism result is more robust and is consistent across three of the four games.","major_comments":[{"comment":"The claim that reciprocity 'becomes more pronounced as the game progresses' is not supported by the reported analysis. In Eq. (3.2), both the response Y_ijt and the predictor Y_jit are cumulative counts up to the same round t. At small t these counts contain very little information and are mostly zero, so the early non-significance of rho_t is exactly what would be expected under a constant per-round reciprocity effect as information accumulates. Moreover, because both variables are contemporaneous cumulative counts, rho_t captures a symmetric association between how much i and j give to each other over the whole game, not a temporal response in which a prior gift raises the probability of a later gift. The authors should test the strengthening claim directly, for example by fitting a model to per-round or differenced data, or by using a lagged predictor such as Y_{ji,t-1}.","section":"Section 4.3, Eq. (3.2), Fig. 10"},{"comment":"No correction is made for multiple testing. The authors report significance of coefficients for four games, four models, two effect types, and up to forty rounds. The 95% bands in Fig. 10 are pointwise bands from the null simulations, and the p-values in Table 3 are unadjusted. Consequently, the statement that the group effect is 'significant at almost all rounds' and the comparison of early versus late reciprocity are familywise claims that could be driven by the large number of tests. The authors should either provide multiple-testing-corrected bands or clearly state that all conclusions are pointwise and assess robustness to the number of comparisons.","section":"Sections 4.2 and 4.3, Table 3 and Fig. 10"},{"comment":"The null model assumes that, in the absence of reciprocity and ingroup favouritism, every player chooses every other player with equal probability in every round. This is a strong behavioural assumption. If participants have other baseline preferences, such as a tendency to give to players in certain screen positions or to players with fewer tokens, those preferences would be absorbed into the estimated rho and gamma coefficients and could make the reciprocity or group effects appear significant. Because the null model is not fitted to the data, the p-values are conditional on this specific null. The authors should test alternative null models, for example one that preserves each player's observed marginal giving rates or one that includes positional choice probabilities, to confirm that the substantive conclusions are robust.","section":"Section 3.1"}],"minor_comments":[{"comment":"The text says that when the two unusual players in game 3 were 'removed from the data' the results became more similar to the other games, but in Section 4.4 the players were not removed; they were replaced by simulated null players. The Discussion should use the more precise terminology of replacement rather than removal.","section":"Section 4.4 and Discussion"},{"comment":"The caption says the dotted line and shaded region are 'explained in the Game 3 section', but the figure appears before Section 4.4. Rephrase the caption or move the explanation so the figure is self-contained.","section":"Figure 10 caption"},{"comment":"The threshold of 2 for flagging influential players is arbitrary, and the influence metric itself is not given a null distribution. This is acceptable as an exploratory tool, but the authors should state that the cut-off is a descriptive choice and not a formal test.","section":"Section 4.5, Table 4"},{"comment":"No data- or code-availability statement is provided. Making the anonymised data and simulation code available would substantially strengthen the reproducibility of the results.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically honest and the null-model inference is a sensible choice for the stated problem. The main concern is that the temporal reciprocity conclusion rests on cumulative variables and therefore does not follow from the reported coefficients; this is fixable with a reanalysis using per-round or lagged formulations. The group-effect result appears more robust, so the paper is likely salvageable after substantial revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely useful: it adapts a null-model regression approach to individual-level token-exchange data and shows that ingroup favouritism is strong and immediate. The group effect is robust across games and survives scrutiny; that part of the analysis deserves credit. The influence metric, which replaces a player with a random null player and measures the distance between coefficient curves, is a clever and practical way to flag anomalous participants, and it correctly identifies the unusual pair in game 3.\n\nThe soft spot is the reciprocity-over-time claim. The model uses cumulative outcomes: Y_ijt is the total number of tokens i received from j up to round t, and the predictor is the cumulative total i gave to j. Fitting this at every t and watching the coefficient become significant later is exactly what you would expect when counts accumulate and power grows, even if the per-round reciprocity effect is constant. The statement that 'it takes time for a reciprocal relationship to develop' is not supported by this analysis; it is a cross-sectional correlation on cumulative totals, not a temporal response. To make that claim, the authors need lagged or per-round predictors, or at least an explicit model of increments.\n\nThere are smaller issues. The significance testing ignores multiple testing across 40 rounds and several models. The influence threshold of 2 is arbitrary. Only one null model is used, and while that is fine for a first pass, the conclusion would be stronger if they tested a null with ingroup baseline giving or other alternatives. No data or code is provided, which makes the reproducibility hard to assess. These are not fatal, but they are real.\n\nOverall, this is a solid methodological contribution that deserves a serious referee. The group-effect finding and the general framework are worth publishing, but the paper needs substantial revision to reframe or re-analyze the reciprocity trajectory. I would send it out.","headline":"The null-model regression is a useful addition to the social-interaction toolkit, and the ingroup-favouritism result is solid, but the 'reciprocity strengthens over time' claim is an artifact of cumulative outcomes.","tokens_in":14480,"tokens_out":1560,"would_cite":false,"duration_ms":16123,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J05","62P25","91D30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Pair-by-pair token exchanges show ingroup bias appears immediately, while reciprocity strengthens over rounds.","keywords":["agent-based model","null model","regression","simulation","Virtual Interaction Application","temporal network data","ingroup favouritism","reciprocity"],"falsifier":"Run the same four-game analysis with an alternative null model, for example one where each player chooses a recipient uniformly at random but with self-giving at the observed rate, or where giving is biased toward physically nearby nodes in the displayed network; if the observed reciprocity and group coefficients then fall inside the 95% simulation bands, the equal-probability null is what creates the significant effects.","tokens_in":13528,"feed_emoji":"🤝","tokens_out":6440,"duration_ms":67414,"temperature":0.7,"pith_summary":"This paper presents a way to model the stream of token exchanges in an online social-interaction experiment, where standard statistical assumptions of independent, normally distributed observations do not hold. The authors fit a regression explaining how many tokens one player has received from another by how many they gave back and whether they share a group, and they test significance against a simulated null model in which everyone gives at random. Applied to four experimental games, the model finds that favouring one's own group is strong and present from the first rounds, while reciprocity is weaker at first and grows more pronounced as the game proceeds. It also provides a systematic way to flag players whose behaviour strongly distorts the estimated effects.","feed_headline":"Ingroup bias appears instantly; reciprocity builds over time","feed_subtitle":"A regression-plus-simulation model reads pair-level giving in online experiments and flags outlier players.","key_machinery":"The central object is the pair-level cumulative-count regression $Y_{ijt} = \\alpha + \\rho Y_{jit} + \\gamma G_{ij} + \\varepsilon_{ijt}$ for $i \\neq j$, where $Y_{ijt}$ is the number of tokens player $i$ has received from player $j$ up to round $t$, $Y_{jit}$ measures reciprocation, $G_{ij}$ indicates whether $i$ and $j$ are in different groups, and the error term is left unspecified. Because the errors are non-normal and interdependent, significance is assessed by an agent-based null model: simulate games where every player chooses a recipient uniformly at random, refit the regression 10,000 times, and compare the observed coefficients to the simulated null distribution. A companion influence metric replaces one player's actions with a random giver, refits the model, and measures the $\\ell^1$ distance of the resulting coefficient trajectories to flag unusual participants.","core_discovery":"The paper's central claim is that individual-level token exchanges in the VIAPPL games reveal two distinct normative dynamics: ingroup favouritism is present and statistically significant from the very first rounds, while reciprocity starts weak and becomes stronger as players build relationships. Across three of the four games the estimated coefficients follow very similar trajectories, indicating a repeatable pattern; the fourth game is an outlier because two players from different groups exchanged tokens with each other in almost every round. The model-based influence metric, which replaces a player's actions with random giving and measures the resulting shift in coefficient paths, identifies exactly those two players as unusually influential, and removing their influence brings the game's results closer to the others.","pith_inferences":["The equal-probability null is a strong assumption; testing the same model against a spatially biased or position-based null would show whether part of the estimated group effect is actually an artefact of node layout.","The same influence metric could serve as a bot-detection tool on online social platforms, as the authors hint; a concrete check would be whether known bot accounts receive influence scores above the same threshold.","The cumulative response variable hides recency; a rolling-window version of the response could test whether reciprocity strengthens within a game or simply accumulates mechanically.","The null-model simulation could be inverted into a model-selection device: instead of testing only 'no effect', one could compare the observed data against several candidate behavioural rules and see which rule's simulated coefficient distribution best contains the observed coefficients."],"forward_implications":["Ingroup favouritism is measurable at the individual-interaction level and is present from the start of the game, not only in aggregate.","Reciprocity is a real but slower-forming norm; it becomes significant only after several rounds, suggesting reciprocal ties build on group context.","The dynamics are repeatable across independent groups: three of the four games show very similar coefficient trajectories.","Unusual players can be detected systematically by their influence on the coefficient estimates, rather than only by inspecting scatterplots.","Because the method avoids independence and normality assumptions, it transfers to other rule-based interaction settings such as iterated prisoner's dilemma experiments."],"supporting_citations":[{"why":"Prior aggregated analysis of VIAPPL data that found ingroup favouritism; the baseline result this paper's individual-level findings extend.","marker":"[9]"},{"why":"Guide to agent-based modelling for social psychologists; supplies the simulation approach used to generate null games.","marker":"[16]"},{"why":"Overview of agent-based modelling methods; underpins the null-model simulations.","marker":"[3]"},{"why":"Discusses random graphs as null models for networks; motivates significance testing against simulated behaviour.","marker":"[26]"},{"why":"Guide to null models for animal social networks; supports the general null-model methodology.","marker":"[10]"},{"why":"Source of dfbetas influence diagnostics; conceptual basis for the paper's influence metric.","marker":"[5]"},{"why":"Force-directed graph layout algorithm used to position nodes in the visualisation.","marker":"[12]"}],"fun_headline_variants":["Ingroup bias instant, reciprocity builds – model spots outliers","Bias immediate, reciprocity delayed: model identifies odd players","Gift exchanges reveal: ingroup bias immediate, reciprocity delayed","Reading social experiments: ingroup bias first, reciprocity later"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results depend on the null model that, without reciprocity or group bias, every player would give their token to every other player with equal probability in every round; if players have other default preferences, those defaults would be counted as ingroup bias or reciprocity.","fun_headline_variants_meta":{"raw":{"variants":["Ingroup bias instant, reciprocity builds – model spots outliers","Bias immediate, reciprocity delayed: model identifies odd players","Gift exchanges reveal: ingroup bias immediate, reciprocity delayed","Reading social experiments: ingroup bias first, reciprocity later"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000698,"raw_usage":{"total_tokens":3097,"prompt_tokens":830,"completion_tokens":2267,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":2199}},"tokens_in":446,"tokens_out":2267,"duration_ms":17712,"temperature":1.0,"reasoning_tokens":2199,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:40:04.299754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four-game analysis with an alternative null model, for example one where each player chooses a recipient uniformly at random but with self-giving at the observed rate, or where giving is biased toward physically nearby nodes in the displayed network; if the observed reciprocity and group coefficients then fall inside the 95% simulation bands, the equal-probability null is what creates the significant effects.","supporting_citations":[{"cited_title":"Tredoux, Kim Titlestad, an d Larry Tooke","cited_arxiv_id":null,"evidence_quote":"Prior aggregated analysis of VIAPPL data that found ingroup favouritism; the baseline result this paper's individual-level findings extend."},{"cited_title":"No rton, and Kurt Gray","cited_arxiv_id":null,"evidence_quote":"Guide to agent-based modelling for social psychologists; supplies the simulation approach used to generate null games."},{"cited_title":"Bonabeau","cited_arxiv_id":null,"evidence_quote":"Overview of agent-based modelling methods; underpins the null-model simulations."},{"cited_title":"On the use of random graphs as nu ll model of large connected networks","cited_arxiv_id":null,"evidence_quote":"Discusses random graphs as null models for networks; motivates significance testing against simulated behaviour."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Guide to null models for animal social networks; supports the general null-model methodology."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of dfbetas influence diagnostics; conceptual basis for the paper's influence metric."},{"cited_title":"Graph drawing b y force-directed placement","cited_arxiv_id":null,"evidence_quote":"Force-directed graph layout algorithm used to position nodes in the visualisation."}],"review_version":1}