{"id":"d4c19e10-0cac-4b4b-9e67-f982a94314e9","arxiv_id":"2504.20774","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In stateless mean-field congestion games, the price of anarchy is infinite under exponential discounting but equals 2 under power-law discounting whenever a stationary equilibrium exists, independent of discounting intensity.","lead":"This paper analyzes how myopic discounting changes the efficiency of mean-field congestion games, where agents repeatedly choose actions that congest shared resources. It finds exponential discounting can drive the price of anarchy to infinity, while power-law discounting keeps it at 2 whenever a stationary equilibrium exists.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The α≤1 power-law regime is not a well-defined game: utilities in (1) are infinite, so (6) is vacuous and the PoA=2 theorem rests on an unstated overtaking criterion.","rationale":"The exponential-discounting half is well-supported: the Bellman reduction, Kakutani fixed-point argument, and PoA=M+1 bound are internally coherent. The α>1 power-law analysis has a separate, less fundamental gap: Proposition 8 only establishes best-response after some initial time, and the proof of Proposition 5 does not explicitly show that a stationary equilibrium must be globally optimal from time 0; this is likely fixable via a tail-switching argument, since a non-maximum reward-rate action can be abandoned at large T. The dominant concern is the α≤1 regime: the model in Section 2 defines utility as a possibly infinite sum, not as a limit of partial sums. The text admits divergence and uses asymptotic comparison (14), effectively changing the solution concept without updating Definition 1. This is not a matter of convention: under the stated objective, all strategies have infinite utility, so every distribution is an equilibrium and the PoA can be unbounded. A conditional acceptance requiring a formal overtaking criterion is therefore the right level; the authors need to amend Section 2/Definition 1 and re-derive the α≤1 characterization. The concrete test above settles whether the concern lands: if the literal definitions are kept, the claimed PoA=2 theorem is false; if overtaking is intended, the paper must state it and revise the equilibrium definition accordingly.","tokens_in":30975,"tokens_out":15302,"duration_ms":168820,"concrete_test":"Take α=1/2, A={1,2}, τ1=1, τ2=ε, r1=1, r2=1/ε, with constant sojourn times satisfying Assumption 1. Under the literal definitions (1)-(2), V(σ,τ)=∞ for every strategy, so every µ∈U_m(A) is a stationary equilibrium. The social optimum is all mass on action 2, with SW=m·r2/τ2=m/ε^2, while the equilibrium all mass on action 1 has SW=m, giving PoA=1/ε^2→∞. Since Proposition 5 states PoA=2, this instance disproves the theorem as written. If the authors intend to replace (2) by an overtaking or limsup criterion, the test is to state that criterion explicitly in Definition 1 and re-run the proof of Theorem 2; the current text never defines a preference order over infinite-utility strategies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that power-law discounting gives PoA=2 whenever a stationary equilibrium exists is not established for 0≤α≤1 under the model as stated. Section 2 defines agent payoff as the infinite sum V(σ,τ) in (1) and optimality via the supremum in (2). For α≤1, the paper itself notes in §4.1 that 'the series diverges' and switches to comparing truncated sums VT(i,τ) in (14). But the equilibrium condition (6) is defined using V, not VT or an overtaking limit. When V=∞ for every strategy that accrues infinitely many rewards, the supremum in (2) is attained by all strategies. Consequently, under Definition 1 every compatible mass distribution is a stationary equilibrium, and the characterization (15) in Theorem 2—that only maximum reward-rate actions lie in the support—does not follow from optimality; it is an extra asymptotic-dominance assumption. The PoA=2 upper bound in Proposition 5 uses exactly (15), so it is not a consequence of the stated game. A concrete instance (α=1/2, τ=(1,ε), r=(1,1/ε)) has all distributions as equilibria under the literal definition and gives PoA=1/ε^2→∞, so the broad claim is false without an overtaking criterion. No such criterion is stated in Section 2 or Definition 1.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies a stateless mean-field congestion game in which a continuum of agents repeatedly choose actions with congestion-dependent sojourn times and maximize the discounted present value of an infinite reward stream. The paper defines stationary equilibria, proves existence for exponential discounting and for power-law discounting with α in [0,1], and computes the price of anarchy: infinity for exponential discounting and 2 for power-law discounting whenever a stationary equilibrium exists. It also analyzes stability under projection dynamics, showing that exponential discounting can create unstable equilibria while average-reward equilibria are always locally asymptotically stable.","tokens_in":31156,"tokens_out":9675,"duration_ms":102702,"significance":"If the results hold, the paper makes a clean and surprising contribution: the efficiency cost of selfish dynamic routing depends only on the qualitative form of time preference, not on its intensity. The exponential-discounting analysis is coherent and self-contained: existence via Kakutani's fixed point theorem, an equivalence to stable population games, and a sharp PoA expression in terms of a characteristic index. The stability comparison between discounted and undiscounted dynamics is also valuable, and the fixed-point and inequality arguments are reproducible from the text. However, the power-law α≤1 branch is currently not a theorem of the stated model, so the headline ''PoA = 2 whenever stationary equilibria exist'' is only conditionally established.","major_comments":[{"comment":"The characterization of stationary equilibria for α≤1 is not derived from the model as stated. For α≤1 and any strategy that accrues infinitely many rewards, the sum in (1) diverges, so the supremum in (2) is attained by every such strategy and condition (6) in Definition 1 imposes no restriction on the stationary strategy σ†. The paper notices in §4.1 that 'the series diverges' and switches to the truncated comparison in (14), but no overtaking or catching-up optimality criterion is added to Definition 1. Consequently, Theorem 2's condition (15) does not follow from optimality, and the upper bound in Proposition 5, which uses (15), is not a consequence of the stated game. Concretely, take α=1/2, m=1, τ=(1,ε), r=(1,1/ε). Under the literal definition every stationary distribution is an equilibrium; the distribution μ=(1,0) is compatible and optimal, with SW=1, while the social optimum at μ=(0,1) has SW=1/ε², so the ratio is 1/ε²→∞, contradicting the PoA=2 claim. The model must either state an overtaking criterion explicitly or restrict the power-law results to α>1.","section":"§2, Eq. (1), Definition 1, §4.1, Theorem 2"},{"comment":"The upper-bound proof of Proposition 5 for α>1 relies on the assertion that at a stationary equilibrium 'it is optimal for the agents to choose the action with the largest possible reward rate' by Proposition 8. Proposition 8 establishes that a particular deterministic strategy is the unique best response after some initial time T, but it does not by itself show that this strategy maximizes the full infinite-horizon utility in (1); the proof in Appendix I compares rewards only from time T onward. Since Proposition 5 is the central efficiency claim for α>1, the argument needs to spell out why, if a stationary equilibrium exists, its optimal stationary strategy must coincide with the action identified by Proposition 8 on the entire horizon and why no nonstationary deviation during the initial period can improve on it.","section":"§4.3, Proposition 5, Appendix J"}],"minor_comments":[{"comment":"The power-law discounting parameter is introduced as α>0 in Section 2 but then stated as α≥0 in the first paragraph of Section 4; these should be made consistent.","section":"§2 and §4"},{"comment":"The displayed formula in (14) is typeset incorrectly: the summation limits involving T/τ_i appear garbled and the floor function is not rendered; please fix the notation.","section":"§4.1, Eq. (14)"},{"comment":"There are several typos that should be corrected: 'equilibirum' in Section 1, 'discouting' in Section 3.3, 'indepenedent' in the caption of Fig. 2, and 'The parameter are given as follows' in Section 5.3.","section":"Throughout"},{"comment":"The proof of Theorem 1 asserts continuity of V∗(μ) in μ without proof; this is plausible under the standing assumptions but should be justified or cited.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The exponential-discounting results and the stability analysis appear sound and publishable. The power-law α≤1 branch requires either an explicit overtaking solution concept or a restriction to α>1; the authors should also re-derive the α≤1 equilibrium characterization under the corrected model. If the overtaking criterion is added, the paper is likely acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the exponential-discounting half of this paper is in good shape, and the α>1 power-law results are coherent. The soft spot is load-bearing: for 0≤α≤1, the equilibrium definition as written is vacuous, so the PoA=2 theorem is not a consequence of the stated model. The paper deserves a serious referee, but only after that gap is fixed.\n\nWhat is genuinely new: a stateless mean-field game with congestion-dependent sojourn times, analyzed under two discounting types. For exponential discounting, Theorem 1 gives existence and the explicit characterization (10); Proposition 2 gives the sharp PoA bound M+1, yielding PoA=∞. The equivalence to a stable population game and the stability analysis are nice additions. For α>1, the utility is finite, the nonexistence example (Proposition 4) is convincing, and the PoA=2 result in that regime is a real advance. The proofs are analytic, no fitted parameters, and the self-citations are to results they re-derive in the appendix.\n\nThe problem is in the low-discounting power-law regime. Section 2 defines payoff by the infinite sum (1), optimality by the supremum (2), and equilibrium by condition (6). For α≤1 that infinite sum diverges, as the paper itself notes in Section 4.1. The authors then switch to truncated utilities VT in (14) and characterize equilibria by maximum reward rate ri/τi in Theorem 2. But Definition 1 still uses V, and with V=∞ for effectively every strategy, the supremum in (2) is attained by all strategies. Literally, every compatible distribution is a stationary equilibrium under the stated definition, and (15) is an extra asymptotic-dominance assumption rather than a derived equilibrium condition. The stress-test example (α=1/2, τ=(1,ε), r=(1,1/ε)) shows the collapse: under the literal definition the PoA is unbounded, so Proposition 5's universal PoA=2 claim is false without an overtaking criterion. This is fixable: state in Section 2 that for α≤1 payoffs are compared via the overtaking or long-run-average criterion that VT captures, and adjust Definition 1 accordingly. But without that revision, the paper's central dichotomy—PoA=∞ vs PoA=2 depending only on discounting type—is not established for the low-discounting power-law regime.\n\nThe stability example is suggestive but limited to a two-action one-resource model; that is a minor issue. The citation pattern looks fair; no code or data, but none is needed for this kind of analytic work.\n\nWho should read it: researchers in mean-field games and algorithmic game theory interested in time preferences and congestion efficiency. I would cite the exponential-discounting characterization, but not the power-law PoA=2 claim in its current form.\n\nRecommendation: send to peer review, but tell the authors the α≤1 definition must be repaired before the PoA=2 theorem can be evaluated. As it stands, the manuscript is conditionally acceptable.","headline":"Exponential-discounting results are solid, but the power-law α≤1 regime is formally undefined as stated; the PoA=2 claim there needs an explicit overtaking criterion.","tokens_in":31771,"tokens_out":4174,"would_cite":false,"duration_ms":45010,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A16","91A10","91A26","90B22"],"pacs":[],"model":"deepseek-v4-flash","headline":"The shape of time discounting, not its intensity, sets the price of anarchy in congested systems.","keywords":["mean-field games","price of anarchy","exponential discounting","power-law discounting","stationary equilibrium","congestion","stable population games","learning dynamics"],"falsifier":"Construct the paper's two-action constant-execution-time family with fixed $\\beta>0$, take $t_1\\to 0$, $t_2\\to\\infty$, and choose $r_1, r_2$ so that $r_1/(e^{\\beta(t_1+w_1)}-1)=r_2/(e^{\\beta t_2}-1)$ with $r_2/t_2\\to\\infty$. Corollary 1 says the ratio of socially optimal to equilibrium welfare diverges in this family, so a parameter sequence with bounded ratio would refute the paper's central claim; alternatively, any undiscounted two-action one-resource instance whose equilibrium is unstable under projection dynamics would refute Proposition 6.","tokens_in":30664,"feed_emoji":"⏳","tokens_out":10835,"duration_ms":105523,"temperature":0.7,"pith_summary":"This paper claims that in a system of many selfish agents sharing congested resources, the functional form of time discounting—exponential versus power-law—determines efficiency, while the strength of the discounting parameter does not. The model is a stateless mean-field game: a continuum of agents repeatedly choose actions, each action has a reward and a completion time, and completion times grow with how many agents use the same actions. The authors define stationary equilibria and measure the price of anarchy as the worst ratio between the socially optimal long-run reward rate and the equilibrium reward rate. They find that exponential discounting makes this ratio infinite for every discount parameter $\\beta>0$, while power-law discounting gives ratio $2$ whenever a stationary equilibrium exists—matching the no-discounting benchmark. They also find that exponential discounting can make equilibria unstable under learning dynamics, whereas undiscounted play leaves them stable.","feed_headline":"Discounting shape, not discount rate, decides congestion efficiency","feed_subtitle":"Exponential discounting drives the price of anarchy to infinity; power-law discounting keeps it at 2 whenever an equilibrium exists.","key_machinery":"The carrying object is the stationary best-response index. For exponential discounting, the total discounted reward of repeating action $i$ forever is $r_i/(e^{\\beta\\tau_i(\\mu)}-1)$, so a stationary equilibrium is a mass distribution $\\mu$ in which every used action attains the maximum of these indices; the rate-monotonicity assumption—if an action's usage rate is higher, its sojourn time is no lower—makes the induced population game stable. For power-law discounting, the relevant index is the reward rate $r_i/\\tau_i(\\mu)$ whenever a stationary best response exists, because slow discounting makes the asymptotic reward rate dominate. The PoA proofs split actions into those whose usage rates grow or shrink relative to the social optimum; on shrinking actions, monotonicity lets the lost welfare be charged to equilibrium welfare, producing the factor $M+1$ in the exponential case (with $M=\\chi(G)$) and a factor $2$ in the power-law case. The characteristic number $\\chi(G)$ is the mechanism through which the exponential upper bound is attained: it measures how much exponential discounting can inflate the value of a short fast action relative to its average-reward rate.","core_discovery":"The paper's discovery is a dichotomy: the functional form of time discounting, not its intensity, sets the worst-case efficiency of a congested mean-field system. Under exponential discounting, a stationary equilibrium always exists and is characterized by ties among per-action indices $r_i/(e^{\\beta\\tau_i(\\mu)}-1)$; the price of anarchy is $M+1$ for instances whose characteristic number $\\chi(G)=\\max_i \\sup_{\\mu}(e^{\\beta\\tau_i(\\mu)}-1)/(\\beta\\tau_i(\\mu))$ equals $M$, and because $M$ can be made arbitrarily large, the PoA is $\\infty$ for every $\\beta>0$. Under power-law discounting with $\\alpha\\le 1$, stationary equilibria always exist and reduce to the no-discounting condition $\\max_i r_i/\\tau_i(\\mu)$, giving PoA 2; for $\\alpha>1$ a stationary equilibrium may fail to exist, but if one exists the same PoA 2 bound applies. The contrast is traced to time consistency: exponential discounting keeps preferences stable, so a long, high-reward action is permanently undervalued by a factor that can grow without bound, whereas power-law discounting's slow decay makes the far future decisive and reproduces average-reward behavior. The paper further shows that exponential discounting can create additional unstable equilibria under projection dynamics in a shared-resource model, while the undiscounted game's equilibria are always locally asymptotically stable.","pith_inferences":["If the same model were run with agents whose discount functions are estimated experimentally, the predicted efficiency distribution would be bimodal: power-law discounters should cluster near PoA 2 while exponential discounters should show unbounded worst-case loss, independent of estimated discount parameters. This is a testable prediction the paper does not itself make.","Because the social optimum is defined with an undiscounted long-run average, the paper measures efficiency from a patient planner's perspective; replacing that criterion with a discounted planner would likely compress the exponential-case loss and could break the power-law guarantee.","For $\\alpha\\le 1$ the formal game's utility sum diverges, so the PoA=2 result tacitly adopts an overtaking or asymptotic-growth comparison; making that criterion explicit—or choosing a different tie-break—could change which stationary distributions count as equilibria.","The stability contrast suggests that in real repeated-choice systems, exponential discounters may converge to low-efficiency equilibria, while power-law discounters should converge to the unique stable outcome; interventions that change the shape of agents' time preferences might therefore be more effective than adjusting incentives by a constant factor."],"forward_implications":["Under exponential discounting, changing $\\beta$ cannot fix efficiency; only reducing the sojourn time of the bottleneck action lowers the worst-case loss, and even the best case leaves PoA at least 2.","Under power-law discounting, worst-case efficiency is unaffected by the degree of myopia: whenever a stationary equilibrium exists, the equilibrium reward rate is at least half the socially optimal rate.","For $\\alpha>1$ power-law discounting, designers cannot assume stationary equilibria always exist; nonexistence is a real possibility even though efficiency would be good if one did exist.","In a two-action shared-resource model, exponential discounting can create multiple equilibria, some unstable, while the undiscounted game has only locally asymptotically stable equilibria; learning dynamics will accordingly behave differently."],"supporting_citations":[{"why":"Supplies the undiscounted baseline—stationary equilibria with congestion-dependent sojourn times and PoA=2—that the power-law result is compared against.","marker":"[6]"},{"why":"Defines stable population games and their convergence properties, used to prove equilibrium stability and equivalence.","marker":"[11]"},{"why":"Introduces the stateless mean-field game framework that the average-reward reduction builds on.","marker":"[26]"},{"why":"Provides the existence-of-equilibrium approach for dynamic games with many players underlying the fixed-point arguments.","marker":"[1]"},{"why":"Defines mean-field equilibria, the concept adapted to stationary strategies in this paper.","marker":"[17]"},{"why":"Supplies the projection dynamics whose stability is analyzed in the two-action one-resource model.","marker":"[18]"}],"fun_headline_variants":["Discounting shape sets congestion game inefficiency: infinity or 2","Exponential discounting inflates price of anarchy to infinity","Power-law discounting keeps price of anarchy at 2, unlike exponential","Discounting form dictates efficiency: exponential gives infinity, power-law gives 2"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that congestion is monotone—an action used at a higher rate never completes faster—and, for weak power-law discounting, that divergent reward streams are compared by their far-future growth, a criterion the formal model never states.","fun_headline_variants_meta":{"raw":{"variants":["Discounting shape sets congestion game inefficiency: infinity or 2","Exponential discounting inflates price of anarchy to infinity","Power-law discounting keeps price of anarchy at 2, unlike exponential","Discounting form dictates efficiency: exponential gives infinity, power-law gives 2"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00134,"raw_usage":{"total_tokens":5533,"prompt_tokens":1117,"completion_tokens":4416,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":4350}},"tokens_in":733,"tokens_out":4416,"duration_ms":32371,"temperature":1.0,"reasoning_tokens":4350,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:20:10.037089+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct the paper's two-action constant-execution-time family with fixed $\\beta>0$, take $t_1\\to 0$, $t_2\\to\\infty$, and choose $r_1, r_2$ so that $r_1/(e^{\\beta(t_1+w_1)}-1)=r_2/(e^{\\beta t_2}-1)$ with $r_2/t_2\\to\\infty$. Corollary 1 says the ratio of socially optimal to equilibrium welfare diverges in this family, so a parameter sequence with bounded ratio would refute the paper's central claim; alternatively, any undiscounted two-action one-resource instance whose equilibrium is unstable under projection dynamics would refute Proposition 6.","supporting_citations":[{"cited_title":"In: AAMAS ’23: Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (2023)","cited_arxiv_id":null,"evidence_quote":"Supplies the undiscounted baseline—stationary equilibria with congestion-dependent sojourn times and PoA=2—that the power-law result is compared against."},{"cited_title":"Journal of Economic Theory 144(4), 1665–1693.e4 (2009)","cited_arxiv_id":null,"evidence_quote":"Defines stable population games and their convergence properties, used to prove equilibrium stability and equivalence."},{"cited_title":"∞X k=1 rake−β Pk j=1τaj (µ) # = nX i=1 σ1(i) rie−βτi(µ) + Eσ","cited_arxiv_id":null,"evidence_quote":"Introduces the stateless mean-field game framework that the average-reward reduction builds on."},{"cited_title":"Journal of Economic Theory156, 269–316 (2015)","cited_arxiv_id":null,"evidence_quote":"Provides the existence-of-equilibrium approach for dynamic games with many players underlying the fixed-point arguments."},{"cited_title":"Springer New York, NY (1996)","cited_arxiv_id":null,"evidence_quote":"Supplies the projection dynamics whose stability is analyzed in the two-action one-resource model."}],"review_version":1}