{"id":"2efa4622-dbdd-4b78-8896-1a7167930ed3","arxiv_id":"2508.19848","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Elimination-style tournaments yield more circular win-loss patterns and less apparent hierarchy than round-robin play, and network centrality scores predict match winners as accurately as Elo or official rankings.","lead":"Tennis and fencing match records were turned into directed networks of who beat whom, and the networks' hierarchical shape and circular patterns were compared across tournament formats. The findings suggest that round-robin style play produces stronger dominance structure and fewer rock-paper-scissors cycles than single-elimination brackets, and that network-based scores predict future winners about as well as Elo ratings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cycle-enrichment conclusion depends on a single global (α, β) fit; phase-specific fits could dissolve the DE excess.","rationale":"The paper has real strengths: out-of-sample prediction, public data, multiple sports/disciplines, and consistent qualitative patterns across five ranking scores. The single most load-bearing point is the validity of Eq. (2) as an unbiased null for the cycle-enrichment calculation. The reader's weakest assumption is close to this, but the specific wording 'fitted to the same year's matches' does not match the paper's out-of-sample prediction protocol; the more precise concern is that one global (α, β) pair is applied to both pool and DE subnetworks. If the phases have different intrinsic upset rates, the DE enrichment in Figure 3d could be a pooling artifact. The proposed test—phase-specific fits and a phase-calibrated expectation—settles this. The secondary observation that tennis (also elimination) shows enrichment near 1 suggests the abstract overgeneralizes; this would remain an issue even if the sabre DE result survives. Overall, the conclusion is plausible but not yet established; the paper should be accepted only after the phase-calibration check and a more measured abstract. Hence CONDITIONAL.","tokens_in":34416,"tokens_out":15880,"duration_ms":188026,"concrete_test":"Fit Eq. (2) to pool-phase matches only (out-of-sample: fit on year t−1, apply to year t) and use these pool-calibrated parameters to generate the expected cycle count for the DE-phase network; compare to the observed DE cycle count. If the DE excess persists, it is not an artifact of pooling phases with different upset rates. As a direct check, report phase-specific fitted α and β values and their uncertainties: a significantly larger α for DE than for pool directly supports the claim. Additionally, run the same cycle-enrichment procedure on Grand Slam tennis matches only (pure single-elimination); if enrichment ≈ 1 there, the abstract's general 'elimination tournaments' claim should be narrowed to fencing DE.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that elimination tournaments increase cycle probability rests on the cycle-enrichment analysis in Figure 3, where the expected cycle count is obtained by redrawing each match direction independently using Eq. (2) with α and β fitted per year to all matches pooled together. The same fitted (α, β) are then applied to the pool-phase and DE-phase subnetworks. If the true upset propensity differs by phase, this pooled null is miscalibrated: DE matches will appear enriched whenever their true upset rate exceeds the pooled average, even if a phase-specific model would predict exactly the observed cycle frequency. The paper never reports phase-specific fitted α/β or a goodness-of-fit of Eq. (2) separately for pool and DE, so the 'increased probability of cycles' in DE is inferred from a null that is assumed, not shown, to be unbiased. The consistency across five ranking scores in the SI is reassuring but does not remove this risk, since all five are pushed through the same single-(α, β) pipeline. Additionally, the abstract generalizes to 'elimination tournaments' at large, yet men's tennis—also single-elimination—shows enrichment near 1 in Figure 3a, so the paper's own data support a fencing-DE effect at most, not a universal format effect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs time-evolving directed networks from elite tennis and fencing match data (nodes = players, directed edges = winner to loser) and studies hierarchy and ranking. It applies three global hierarchy measures (flow hierarchy, global reaching centrality, random walk hierarchy), cycle-abundance statistics, and five node-level ranking scores (official sports federation score, Elo, reversed PageRank, 2-reach centrality, random walk centrality). The two central claims are: (1) networks built from elimination-format matches show a smaller level of hierarchy and an increased probability of cyclic (upset) outcomes compared with round-robin-style matches, and (2) hierarchy-based network metrics predict future match outcomes with accuracy comparable to official rankings and Elo. The prediction analysis uses a temporal out-of-sample split: scores are computed from matches before the prediction year, the parameters of the model in Eq. (2) are fitted on a preceding fitting period, and accuracy is evaluated on the next year's matches.","tokens_in":34780,"tokens_out":5870,"duration_ms":63066,"significance":"If the hierarchy/cycle claim were established, the paper would make a useful contribution to the literature on tournament design and network-based ranking: it would show that the format of competition itself shifts the directed network away from transitivity, and that simple network centralities are competitive with established rating systems. The manuscript has clear strengths: it uses publicly available data across eight competition categories, it compares multiple hierarchy measures and multiple ranking scores, and the prediction protocol is a genuine out-of-sample evaluation. The central cycle-enrichment claim, however, currently rests on a null model whose calibration is not demonstrated, and the abstract generalizes the finding beyond what the data show (men's tennis, itself an elimination format, shows enrichment near 1, not above 1). These issues are fixable but require new evidence or a substantial claim restriction.","major_comments":[{"comment":"The cycle-enrichment analysis uses a single set of fitted parameters (α, β) from Eq. (2), apparently pooled across all matches in a given year, to generate the expected cycle count for both pool-phase and DE-phase subnetworks. If the true upset propensity differs by phase, the pooled null is miscalibrated: DE matches will appear enriched whenever their upset rate exceeds the pooled average, even if a phase-specific model would predict the observed cycle count exactly. The paper does not report phase-specific α/β fits or any goodness-of-fit of Eq. (2) separately for pool and DE. It also does not state clearly whether the α/β used in Figure 3 are fitted in-sample on the same year's matches (descriptive goodness-of-fit) or out-of-sample as in the prediction analysis. Please provide phase-specific fits, an explicit statement of the fitting window, and a calibration check; without this, the c","section":"Methods: 'Cycle abundance and cycle enrichment'; Figure 3"},{"comment":"The abstract states that 'elimination tournaments lead to networks with a smaller level of hierarchy and thus, importantly, to an increased probability of circular win-loss situations (cycles).' The data in Figure 3 do not support this as a general statement about elimination formats. Men's tennis, which the text itself describes as single-elimination, shows cycle enrichment near 1 (Figure 3a), not above 1, whereas the elevated enrichment appears only for the DE phase of fencing (Figure 3d). The claim should be restricted to the fencing DE phase, or the authors must supply an explanation for why tennis's elimination format behaves differently and why the universal wording is justified.","section":"Abstract and Results, Figure 3(a) vs Figure 3(d)"},{"comment":"The headline claim of 'smaller level of hierarchy' in elimination networks is in tension with the raw hierarchy measures: FH and RWH/RWC are larger for the DE-phase network than for the pool-phase network (Figure 2g vs Figure 2e). The authors attribute this to the tree-shaped skeleton of DE graphs, but the abstract and Discussion do not carry this caveat. The comparison to Erdős–Rényi ratios is suggestive, but the paper should either report a structure-controlled hierarchy comparison explicitly or temper the abstract's unqualified statement. The current text moves from raw measures to ratios to cycle enrichment without a single, clearly defined quantity that supports 'smaller level of hierarchy' across all measures.","section":"Results, Figure 2(e)–(h)"},{"comment":"The rows for Women's épée and Women's foil are identical: FM 0.329, SE 0.213, LE 0.417 for 2RC, RWC, SFS, ELO, and RPR. This is almost certainly a data-entry or copy-paste error. Since Table 4 is the main evidence for the prediction claim across all categories, these duplicated rows need to be corrected before the comparison can be assessed. If the values are genuinely identical, the manuscript should explain why.","section":"Table 4"},{"comment":"The text states that cycle enrichment values are 'significantly larger' for DE matches, but no confidence intervals, standard errors, or significance tests are reported for the enrichment ratios. The shaded regions are only described as standard deviations of an unspecified number of relinking trials; the Methods section describes relinking in the singular ('we relinked the network'), so the number of null realizations is unclear. Please state the number of relinking samples and provide error bars or confidence intervals for the enrichment values, and preferably a formal test comparing pool and DE enrichment distributions.","section":"Figure 3; Results paragraph on 'significantly larger'"}],"minor_comments":[{"comment":"The caption says 'the tennis tournament is in a round-robin format,' which contradicts the Results section where tennis is correctly described as single-elimination. Tennis Grand Slams and tour-level events are not round-robin; this is also relevant to the abstract's wording about 'round-robin data.' Please correct the caption and adjust the abstract to avoid conflating sport identity with tournament format.","section":"Figure 1 caption"},{"comment":"The text reads 'FH and RWC values are firmly larger for the direct elimination network' but the corresponding measures in the figure and Methods are FH, GRC, and RWH (random walk hierarchy), not RWC (random walk centrality). Please correct the acronym.","section":"Results, paragraph after Figure 2(e)"},{"comment":"The phrase 'revered pagerank' appears several times (e.g., Figure 4, Figure 6, SI captions); it should be 'reversed PageRank.' Also, Figure 7 caption and one SI caption contain 'rank-socre' instead of 'rank-score.'","section":"Throughout"},{"comment":"Several SI captions mislabel the year: e.g., Figure S41 says 'men's tennis, 2015' in the caption text while the panel header says 2023. Similar copy-paste issues occur in other SI captions. Please standardize.","section":"SI captions S40–S55"},{"comment":"The description of the fitting procedure is somewhat indirect ('trial probabilities for matches at t'). A precise statement of which matches are in ts, which in tf, and how they relate to the calendar-year scores would help reproducibility. Also, the Elo update formula and the PageRank equation contain notation (e.g., M(i), L(j)) that should be defined explicitly, since L(j) is also used for the number of links in the FH definition.","section":"Methods: 'Making predictions using score differences'"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely topic in sports network science and the out-of-sample prediction protocol is a real strength. However, the headline claim about elimination tournaments currently rests on a null model whose phase-specific calibration is not shown, and the paper's own tennis results undercut the universal wording. The duplicated rows in Table 4 also make the prediction comparison unreliable as presented. These are fixable with additional analyses and claim restrictions, so I do not recommend rejection, but the current version is not ready without substantive revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a serious empirical study, not a gimmick. What's new: it exploits the fact that fencing tournaments contain both a pool (round-robin) phase and a direct elimination (DE) phase, so you can compare formats while holding sport, players, and season fixed. That's a clean setup, and the dataset is large and public: tennis 1968-2024, fencing 2015-2024, eight categories. The prediction comparison is also well done: five ranking scores, three accuracy metrics, genuinely out-of-sample (fit on year t, predict t+1). The finding that official fencing scores predict better than Elo is a real result. So there is solid work here.\n\nThe soft spot is the headline. The abstract says elimination tournaments lead to smaller hierarchy and more cycles, but their own Figure 3 shows men's tennis—a single-elimination sport—has cycle enrichment near 1, not above it. The effect appears only in the fencing DE phase. So the claim is overgeneralized. Worse, the cycle-enrichment analysis uses a single (α, β) fitted to all matches of a year, then applies it to pool and DE subnetworks. If the true upset propensity differs by phase, the null is miscalibrated and the DE enrichment could be an artifact. They never report phase-specific fits or a goodness-of-fit for the model per phase. That's a real hole. The hierarchy measures are also in tension: raw FH and RWH are higher in DE networks, and the 'smaller hierarchy' conclusion comes only from the ratio against Erdős-Rényi baselines, which is a structural caveat.\n\nThere are smaller issues: the random walk parameter f for RWH/RWC is never given a value, which hurts reproducibility, and the shaded error regions in Figure 3 don't come with significance tests, so we don't know if the pool-DE difference is meaningful.\n\nNone of this makes the paper worthless. The prediction results are solid and the pool-vs-DE comparison is worth having. But the central claim needs to be re-analyzed, likely with phase-specific fits, and the abstract should be scaled back to what the data actually support.\n\nMy recommendation: send it to peer review, but with a referee who will demand the phase-specific fits and explicit significance testing. If those don't come out, the paper is still publishable as a prediction-comparison study, but not with the current abstract.\n\nWould I bring it to reading group? Maybe—it's a good case study in why null models need to be calibrated per subgroup. I'd cite it for the prediction results, but not for the cycle-format claim without qualification.","headline":"A solid empirical study of round-robin vs elimination formats, but the headline claim about elimination increasing cycles is overreached and rests on a pooled null that needs phase-specific checks.","tokens_in":35202,"tokens_out":6263,"would_cite":true,"duration_ms":68904,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Elimination tournaments make match networks less hierarchical and more cyclic than round-robin formats.","keywords":["sports networks","tournament design","hierarchy","cycles","ranking","match prediction","tennis","fencing"],"falsifier":"Take the same datasets and repeat the cycle-enrichment calculation with an alternative outcome model (for example, one that adds surface effects or a different rating formula); if the direct-elimination phase no longer shows cycle enrichment above 1 while pool phases stay below 1, the claimed format effect would not survive the change of model.","tokens_in":34392,"feed_emoji":"🏆","tokens_out":5017,"duration_ms":53687,"temperature":0.7,"pith_summary":"The paper builds directed networks from years of tennis and fencing matches, with each node a player and each directed edge pointing from the winner to the loser. Its central claim is that the tournament format itself, not just player ability, shapes the network's structure: single-elimination phases show weaker hierarchy and more three-player circular upsets than round-robin pools, even after accounting for strength differences. A second claim is that position in the network hierarchy, including a reversed PageRank score, predicts match winners about as well as Elo or official federation points. If both hold, the design of a competition changes how much dominance rankings can reveal, with practical consequences for seeding, ranking fairness, and forecasting outcomes.","feed_headline":"Elimination tournaments breed more upsets than round-robin ones","feed_subtitle":"Match networks from tennis and fencing show the knockout format itself weakens hierarchy and boosts circular upsets.","key_machinery":"Directed match networks built from a one-year sliding window over past results, combined with three global hierarchy measures (flow hierarchy FH, global reaching centrality GRC, and random walk hierarchy RWH), cycle abundance metrics (cycle density and cycle-FFL ratio), and a cycle-enrichment statistic that compares observed 3-cycles with those produced by resampling each match direction from a two-parameter logistic model of win probability. Reversed PageRank (RPR) serves as both a hierarchical centrality and a ranking score.","core_discovery":"On the paper's own terms, the key result is that elimination tournaments lead to networks with a smaller level of hierarchy and thus an increased probability of circular win-loss situations (cycles). When match directions are resampled according to the fitted two-parameter Bradley-Terry-type model of Eq. (2), the observed number of 3-cycles in the direct-elimination phase exceeds the expected number (cycle enrichment above 1), while round-robin pools show enrichment below 1. The paper interprets this as evidence that knockout formats generate more upsets relative to estimated player strengths, and that a substantial part of the perceived hierarchy in such sports comes from the tree-shaped or","pith_inferences":["The cycle-enrichment statistic could be used as a design tool: organizers wanting fewer upsets could favour round-robin phases, while those wanting higher upset rates would use knockout formats.","The same analysis could be applied to other head-to-head sports (badminton, boxing, MMA) to test whether the format effect generalizes beyond tennis and fencing.","Because the enrichment calculation relies on a specific parametric model of match outcomes, the claim should be re-tested with alternative outcome models (including surface-dependent or time-varying strengths) before being taken as a universal law of tournament design.","The finding that reversed PageRank has a two-segment power-law distribution suggests a natural, data-driven threshold for 'elite' status that does not require arbitrary point cutoffs."],"forward_implications":["Knockout-style tournaments can be expected to produce more non-transitive outcomes among the same field of players than round-robin pools.","Official rankings built primarily from elimination events may overstate the true dominance of top players, because the format itself creates a pyramid of wins.","Network-derived scores such as reversed PageRank can predict match outcomes with accuracy close to Elo and official points, offering a ranking method independent of federation points.","The power-law tail observed in reversed PageRank scores identifies a 'doorstep' separating the very best players from other elite athletes."],"supporting_citations":[{"why":"Supplies the global reaching centrality (GRC) hierarchy measure used in the comparisons.","marker":"[29]"},{"why":"Supplies the random walk hierarchy measure (RWH) and random walk centrality used for ranking and hierarchy analysis.","marker":"[31]"},{"why":"Defines the Elo rating system that serves as one of the prediction baselines.","marker":"[35]"},{"why":"Provides the Bradley-Terry paired-comparison model underlying Eq. (1) and the fitted Eq. (2).","marker":"[40]"},{"why":"Introduces the two-parameter extension (alpha, beta) of the win-probability formula used for fitting and cycle enrichment.","marker":"[41]"},{"why":"Supplies the men's tennis match dataset used throughout the study.","marker":"[42]"},{"why":"Supplies the women's tennis match dataset used throughout the study.","marker":"[43]"},{"why":"Supplies the fencing datasets for all six weapon/sex categories.","marker":"[44]"},{"why":"Defines flow hierarchy (FH), the fraction of links not participating in any cycle.","marker":"[45]"},{"why":"Defines the PageRank algorithm, reversed in this paper to score players who beat strong opponents.","marker":"[47]"}],"fun_headline_variants":["Elimination tournaments weaken hierarchy, fuel upsets","Knockout rounds boost circular upsets over round-robin","Tournament format skews hierarchy: knockout lifts upsets","Round-robin shows stronger hierarchy than knockout","Knockout style reduces hierarchy, increases upsets"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim that elimination phases contain more upsets than expected rests on the fitted two-parameter model being an unbiased generator of every match's direction; if that model is misspecified, the expected cycle count is wrong and the format effect conclusion does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Elimination tournaments weaken hierarchy, fuel upsets","Knockout rounds boost circular upsets over round-robin","Tournament format skews hierarchy: knockout lifts upsets","Round-robin shows stronger hierarchy than knockout","Knockout style reduces hierarchy, increases upsets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1364,"prompt_tokens":776,"completion_tokens":588,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":511}},"tokens_in":520,"tokens_out":588,"duration_ms":6705,"temperature":1.0,"reasoning_tokens":511,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:23:56.697007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same datasets and repeat the cycle-enrichment calculation with an alternative outcome model (for example, one that adds surface effects or a different rating formula); if the direct-elimination phase no longer shows cycle enrichment above 1 while pool phases stay below 1, the claimed format effect would not survive the change of model.","supporting_citations":[{"cited_title":"& Vicsek, T","cited_arxiv_id":null,"evidence_quote":"Supplies the global reaching centrality (GRC) hierarchy measure used in the comparisons."},{"cited_title":"& Palla, G","cited_arxiv_id":null,"evidence_quote":"Supplies the random walk hierarchy measure (RWH) and random walk centrality used for ranking and hierarchy analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Elo rating system that serves as one of the prediction baselines."},{"cited_title":"Atp tennis data","cited_arxiv_id":null,"evidence_quote":"Supplies the men's tennis match dataset used throughout the study."},{"cited_title":"Wta tennis data","cited_arxiv_id":null,"evidence_quote":"Supplies the women's tennis match dataset used throughout the study."},{"cited_title":"Fie fencing data","cited_arxiv_id":null,"evidence_quote":"Supplies the fencing datasets for all six weapon/sex categories."},{"cited_title":"& Magee, C","cited_arxiv_id":null,"evidence_quote":"Defines flow hierarchy (FH), the fraction of links not participating in any cycle."}],"review_version":1}