{"id":"36574a57-d44f-4854-9ea2-91b083549e66","arxiv_id":"2509.02493","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In feedback Stackelberg incentive design, a principal's cost is concave in her belief, so free information always weakly helps her, and both agent persuasion and costly experiments reduce to convex-hull problems.","lead":"This paper studies how a principal whose incentive policy must react to an agent's actions should handle the fact that the agent knows more about the world. It shows the principal always benefits from free information, and compares two ways to get it: letting the agent design a signal, or paying for a signal herself.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"QG G3/G4 results rest on unproved affine/Gaussian restrictions; the matrix-game concavity core is sound.","rationale":"The reader's CONDITIONAL verdict is appropriate. I agree with the QG-scope concern: the paper's own Section 4 disclaims optimality of its restricted policies, so the QG 'solutions' are conditional. I do not see a fatal flaw in the matrix-game core: Lemma 1's concavity is valid because J*_{P,2} is a minimum of affine functions over a prior-independent feasible set, so the P-never-hurt claim (17) is robust. The tie-breaking issue raised by the reader is real for the exact value of G3 when the receiver is indifferent, but it does not threaten the central P-cost bound, since any tie-breaking among P-optimal policies leaves J*_{P,2}(µ_s) unchanged. The biggest unqualified claim that needs correction is the QG section's scope. Hence no verdict change from CONDITIONAL is needed.","tokens_in":13764,"tokens_out":18813,"duration_ms":193161,"concrete_test":"For a grid of β>0, solve the QG G2 problem within the full class of affine policies γ(v)=L1 v+L2: minimize P's expected cost over (L1,L2) subject to A's unique best response, and compare the optimum with the mean-feedback cost (30). If any β yields strictly lower cost, then the §4 baseline J*_{P,2} is not the affine-optimal value, and the G3/G4 comparisons in that section must be recomputed with the optimal affine policy before claims about QG information acquisition can stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core matrix-game claim survives: J*_{P,2} in Lemma 1 is the pointwise minimum over a µ0-independent feasible set of objective functions affine in µ0, so it is concave; (17) then correctly implies the principal's expected cost cannot increase with any signal, regardless of the agent's objective. The load-bearing weakness is the paper's scope claim. Section 4 explicitly restricts P to affine mean-feedback policies and G3 signaling to additive Gaussian noise channels, and admits 'affine mean feedback policies may not constitute an optimal affine incentive policy'. Consequently J*_{P,2}, J*_{P,3}, and J*_{P,4} in the QG section are not equilibrium values of G2-G4; they are costs of specific heuristic strategies. All G3/G4 conclusions in §4, including the 'either full or no revelation' dichotomy and the Figure 4 optimization, are conditional on these unproved restrictions. The abstract's 'providing solutions to all these cases' therefore overstates what is established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies feedback-Stackelberg incentive design (a principal commits to a policy that reacts to an agent's action) under information asymmetry. It defines four games: G1 with full information, G2 with the agent observing the state and the principal only holding a prior, G3 in which the agent chooses a persuasion signal before the principal commits, and G4 in which the principal purchases a costly signal. For two-state finite matrix games, Section 3 formulates G1 and G2 as linear programs, proves that the principal's G2 value J*_{P,2} is piecewise affine and concave, derives from concavity that any signal weakly reduces the principal's expected cost (Eq. 17), and expresses the G3 and G4 solutions via biconjugates, with two worked numerical examples. Section 4 restricts attention to scalar quadratic-Gaussian costs, affine mean-feedback incentive policies, and additive Gaussian noise channels, gives closed-form G1/G2 costs, shows the agent's optimal G3 disclosure is either full revelation or no revelation depending on the sign of f(beta), and numerically analyzes the principal's G4 monitoring cost.","tokens_in":13872,"tokens_out":11387,"duration_ms":105225,"significance":"If the concavity result holds, Eq. (17) is a clean and useful bridge between Bayesian persuasion and incentive design: a principal who can tailor her policy after observing any signal is never hurt in expectation by information, even when the agent designs the signaling scheme. The linear-programming treatment of the matrix games in Section 3 is transparent, and the beta-dependent full/no-revelation dichotomy in the quadratic-Gaussian example is a sharp falsifiable prediction. The paper also gives concrete numerical examples that illustrate the trade-off between agent-designed persuasion and costly principal-designed channels. However, two central lemmas are stated without proof, tie-breaking is left unresolved, and the quadratic-Gaussian section explicitly works under unproved restrictions while using equilibrium-style notation. The core matrix-game claim is defensible, but the presentation currently overstates what is established.","major_comments":[{"comment":"The final step of the proof of Lemma 1 is not valid as written. The proof says J*_{P,2} takes the minimum over the piecewise affine and concave functions V_ij(mu0), but the minimum of concave functions is not generally concave. The correct argument is that J*_{P,2}(mu0) = min_{gamma in union_{i,j} F_{ij}} [mu0(theta1) gamma(v_i)^T C_P(theta1) e_i + mu0(theta2) gamma(v_j)^T C_P(theta2) e_j], and the feasible union is independent of mu0, so J*_{P,2} is an infimum of a family of affine functions in mu0 and hence concave. Since Eq. (17) is the main matrix-game payoff, this proof gap should be repaired.","section":"Section 3.2, Lemma 1"},{"comment":"Lemma 2 is described as having its proof omitted, and Lemma 3 is also stated without proof. These are central results: Lemma 2 characterizes the agent-optimal persuasion value J*_{A,3} as the biconjugate of J*_{A,2}, and Lemma 3 does the same for the costly-information problem. Citing Kamenica and Gentzkow (2011) is not by itself sufficient because the continuation game here is the principal's incentive-design problem, not a fixed receiver action. The equivalence between the Bayesian-persuasion formulation in (18) and the recommendation-based linear program in (19) also needs a proof. Please supply these proofs or a precise derivation showing how the cited results apply.","section":"Section 3.2, Lemmas 2 and 3"},{"comment":"The linear program in (19) enforces that the recommended vertex is at least as good for the principal as any alternative, but it does not specify what happens when the principal is indifferent. If an indifferent principal breaks ties against the agent, the claimed value J*_{A,3}(mu0) may not be achievable. Footnote 1 addresses non-unique responses by the agent but not non-unique optimal policies for the principal. The paper should state a tie-breaking rule (for example, that an indifferent principal selects the policy that minimizes the agent's cost) and verify that the reported values are consistent with it.","section":"Section 3.2, Eq. (19)"},{"comment":"The quadratic-Gaussian results are explicitly restricted to affine mean-feedback policies and additive Gaussian noise channels, and the text admits that 'affine mean feedback policies may not constitute an optimal affine incentive policy from P's standpoint.' Nevertheless, the paper uses the notation J*_{P,2}, J*_{P,3}, and J*_{P,4} as if these were equilibrium values, and the abstract claims 'providing solutions to all these cases.' These symbols actually denote costs under a restricted class of heuristic policies, not equilibrium values of G2-G4. The abstract, Section 4 wording, and notation should be revised so that the conditional nature of the QG comparisons is explicit.","section":"Section 4, Eqs. (28)-(31) and abstract"}],"minor_comments":[{"comment":"The minimum in Eq. (11) is written over [m]^2, but the indices i and j denote the agent's actions in V, so the minimum should be over [n]^2, as correctly used in Eqs. (12) and (19).","section":"Section 3.1, Eq. (11)"},{"comment":"The function eH(mu_s) is used in Eq. (23) and Lemma 3 but never explicitly defined; please write out the formula for the mapping from mu_s to H(mu'_s) under the stated posterior transformation.","section":"Section 3.2, Eq. (23) and Lemma 3"},{"comment":"The sentence about Bergemann, Heumann, and Morris (BHM22) appears twice in the introduction, and one occurrence contains 'can can'; please delete the duplicate and fix the typo.","section":"Introduction"},{"comment":"The phrase 'Providing solutions to all these cases' overstates the quadratic-Gaussian contribution, which the authors themselves describe as a specific example with restricted policies; please soften the wording to reflect the conditional nature of the QG results.","section":"Abstract"},{"comment":"Figure 4 is a numerical plot; the caption could state the fixed parameter values (z0=1, sigma0=2) and clarify that no closed-form optimality is claimed for the plotted G4 cost.","section":"Section 4, Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern about the QG section is legitimate: the QG results are conditional on unproved restrictions and should not be presented as equilibrium solutions. The matrix-game core is sound once the Lemma 1 proof is repaired, and the omitted proofs and tie-breaking issue are fixable within the manuscript's scope. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe genuinely new thing in this paper is Lemma 1: in a feedback Stackelberg incentive game where the principal does not observe the state, her equilibrium cost as a function of the prior is piecewise affine and concave. The proof is a one-liner—min of affine functions—but the consequence is useful. From concavity and Bayes plausibility, (17) follows: any signal that arrives before the principal commits, even one designed by the agent, weakly lowers her expected cost. That is a compact principle for information acquisition in contract and monitoring design.\n\nWhat else is solid: the matrix-game treatment is careful. G3 and G4 are honest reductions to the Kamenica-Gentzkow convexification of the right continuation value, and the examples are worked out. The example with costs (21) where the agent reveals at some priors and both parties improve is a nice concrete illustration.\n\nThe soft spots are real but manageable. Lemma 2 states the biconjugate result and simply omits the proof; it is a direct corollary of KG11, so no one should panic, but an omitted proof in a main result is bad form. The LP in (19) uses weak obedience constraints. If the principal is indifferent among policies, the agent's recommended policy may not actually be played; the paper does not say how ties are broken. The quadratic Gaussian section is the bigger gap. The authors are transparent that they use affine mean-feedback policies and additive Gaussian noise channels, and they admit these may not be optimal. That is fine for a specific example, but the abstract's 'providing solutions to all these cases' oversells what is established. The QG conclusions are conditional on those restrictions. Several numerical claims, including the noisy-monitoring optimum in Figure 4, are asserted from plots rather than derived. Minor: a sentence is duplicated in the introduction.\n\nNet assessment: the matrix-game concavity core holds up and is worth knowing. The G3/G4 contribution is the identification of the right continuation value to convexify, not a new theorem in persuasion. The paper deserves a serious referee. I would send it out with a request to either prove or properly cite Lemma 2, address the tie-breaking assumption, and scale back the abstract and QG claims to match the restricted policy classes.","headline":"A clean concavity principle for information in incentive design, wrapped in a paper whose QG section is more conditional than the abstract admits.","tokens_in":14466,"tokens_out":4813,"would_cite":true,"duration_ms":45057,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A65","91B44","90C05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Free information never hurts a principal who can commit after the signal.","keywords":["incentive design","Bayesian persuasion","feedback Stackelberg games","information asymmetry","information acquisition","concavity","quadratic Gaussian games","biconjugate"],"falsifier":"In the quadratic Gaussian game, optimize the principal's G2 policy over all affine rules $\\gamma(v)=L_1 v+L_2$ instead of the mean-feedback form; if the resulting cost is strictly below (30) for some $\\beta$, then the reported G2/G3/G4 comparisons are artifacts of the policy restriction rather than equilibrium values. In the matrix setting, a direct test of Lemma 1 is to evaluate $J^\\star_{P,2}$ at the prior and at the two posteriors of any binary signal: any instance with $\\mathbb{E}[J^\\star_{P,2}(\\mu_s)]>J^\\star_{P,2}(\\mu_0)$ would falsify the central inequality (17).","tokens_in":13503,"feed_emoji":"🎯","tokens_out":11814,"duration_ms":113620,"temperature":0.7,"pith_summary":"The paper asks whether a principal who commits to an incentive policy—a rule for reacting to the agent's action—can use information to cut the cost of dealing with an agent who privately knows the state. In two-state matrix games it proves that the principal's Bayesian value function is piecewise affine and concave, and concavity makes any free signal, even one designed by the agent, weakly reduce the principal's expected cost. It gives the agent's optimal persuasion as the convex hull of the agent's cost, and the principal's optimal paid channel as a convex hull of the principal's cost net of an entropy penalty. The same questions are worked out in quadratic Gaussian games under affine mean-feedback policies and Gaussian noise channels. If the results hold, a principal should postpone commitment until after signals arrive and should actively compare agent disclosure against buying her own noisy channel.","feed_headline":"Signals before commitment can only lower the principal's cost","feed_subtitle":"A principal should not fear agent-designed disclosures if she can commit to a policy after seeing them.","key_machinery":"The load-bearing object is the principal's Bayesian value function $J^\\star_{P,2}:\\Delta(\\Theta)\\to\\mathbb{R}$, the least expected cost she can guarantee in the Bayesian feedback Stackelberg game when her belief is $\\mu$. Lemma 1 shows this function is piecewise affine and concave: it is the minimum of finitely many affine functions, one per vertex of the state-independent polytopes that define feasible incentive policies. Concavity combines with Bayes plausibility, $\\mathbb{E}[\\mu_s]=\\mu_0$, to deliver inequality (17), making Jensen's inequality the engine that converts any free information into a weak win for the principal. The companion machinery for G3 and G4 is the Fenchel bi-conjugate, the convex hull of a function's epigraph: the agent's persuasion value is the biconjugate of $J^\\star_{A,2}$, and the principal's acquisition value is the biconjugate of $J^\\star_{P,2}-\\kappa\\tilde H$ plus $\\kappa$.","core_discovery":"The central discovery is that information asymmetry in feedback Stackelberg incentive games can be harnessed through concavity. The principal's cost value function $J^\\star_{P,2}$ is piecewise affine and concave on the belief simplex, and Bayes plausibility fixes the expected posterior at the prior, so $\\mathbb{E}[J^\\star_{P,2}(\\mu_s)] \\le J^\\star_{P,2}(\\mu_0)$: free information before commitment weakly helps the principal regardless of who designs the signal. Both parties' optimal information games are then convexification problems: the agent's persuasion value equals the biconjugate of $J^\\star_{A,2}$, and the principal's channel value equals the biconjugate of $J^\\star_{P,2}-\\kappa\\tilde H$ plus $\\kappa$. In the quadratic Gaussian example the agent's optimal signal is extreme—full revelation or none—depending on the sign of $f(\\beta)$, while the principal's optimal paid channel can be noisy; the paper also exhibits matrix games where the agent's persuasion strictly helps the principal attain the full-information cost.","pith_inferences":["The proof of Lemma 1 never uses the two-state restriction, so the 'free information weakly helps' inequality (17) should extend to any finite state space, provided the same state-independent polytope structure holds.","If an indifferent principal can choose which optimal incentive policy to implement, the agent's persuasion LP (19) assumes a favorable tie-break; with a tie-break against the sender, the reported G3 value is an upper bound on what the agent can guarantee.","In the quadratic Gaussian analysis, restricting the channel to additive Gaussian noise collapses persuasion to a one-dimensional variance choice; non-Gaussian channels may yield richer posteriors and could overturn the bang-bang conclusion.","The concavity-plus-Jensen mechanism is a general template: any incentive-design game whose principal's indirect value function is concave in beliefs inherits the property that free information before commitment weakly helps, so identifying other concave-value classes is a natural next question."],"forward_implications":["In every two-state matrix game, a principal who can write her incentive policy after observing a free signal is weakly better off, even when the agent chooses the signal; her cost is at most $J^\\star_{P,2}(\\mu_0)$.","The agent profits from persuasion exactly when her G2 cost function is nonconvex; if $J^\\star_{A,2}$ is convex, signaling has no value and the principal can ignore the agent's disclosure.","In the quadratic Gaussian game with affine mean-feedback policies, the agent's optimal signal is bang-bang (full revelation or no revelation), so the principal's G3 cost equals either the full-information value or the no-information value.","When the principal buys information, the optimal channel can be noisy, and if the price $\\kappa$ is high enough she will buy none; an agent-designed free signal can dominate a paid channel, as in the $\\kappa=2$ matrix example and the $\\beta=1$ Gaussian example."],"supporting_citations":[{"why":"Supplies the Bayesian-persuasion machinery: the Markov property of posteriors, sufficiency of small signal spaces, and the biconjugate characterization used for Lemma 2 and G3.","marker":"[KG11]"},{"why":"Supplies the affine incentive scheme and indirect team-problem method used to solve G1 and to motivate the affine mean-feedback policies in the quadratic Gaussian analysis.","marker":"[Bas ¸84]"},{"why":"Supplies the entropy-reduction cost model for information channels and the costly-persuasion biconjugate formulation that Lemma 3 adapts for the principal's channel purchase.","marker":"[GK14]"},{"why":"Supplies the Bayes-rule transformation between posteriors under two priors, used to rewrite the entropy cost as a function of the posterior under the principal's prior in G4.","marker":"[AC16]"},{"why":"Supplies the quadratic-preference persuasion result whose sign condition the paper's full-versus-no revelation threshold $f(\\beta)$ mirrors in the Gaussian setting.","marker":"[Tam18]"},{"why":"Supplies the soft-policy incentive design setup that frames the feedback Stackelberg game with commitment to a reaction rule.","marker":"[Bas ¸24]"}],"fun_headline_variants":["Free info before commitment always weakly helps the principal","Concavity shows agent-designed signals can lower principal costs","Principal gains from listening before committing to incentives","Information asymmetry reversed: principal benefits from agent's signal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes the linear-program value $J^\\star_{P,2}$ correctly describes the continuation equilibrium, which requires the agent's best response to be unique and presumes a definite tie-breaking rule for the principal; the quadratic-Gaussian comparisons additionally assume affine mean-feedback policies are the relevant class.","fun_headline_variants_meta":{"raw":{"variants":["Free info before commitment always weakly helps the principal","Concavity shows agent-designed signals can lower principal costs","Principal gains from listening before committing to incentives","Information asymmetry reversed: principal benefits from agent's signal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1276,"prompt_tokens":896,"completion_tokens":380,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":319}},"tokens_in":512,"tokens_out":380,"duration_ms":4422,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:39:10.840318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the quadratic Gaussian game, optimize the principal's G2 policy over all affine rules $\\gamma(v)=L_1 v+L_2$ instead of the mean-feedback form; if the resulting cost is strictly below (30) for some $\\beta$, then the reported G2/G3/G4 comparisons are artifacts of the policy restriction rather than equilibrium values. In the matrix setting, a direct test of Lemma 1 is to evaluate $J^\\star_{P,2}$ at the prior and at the two posteriors of any binary signal: any instance with $\\mathbb{E}[J^\\star_{P,2}(\\mu_s)]>J^\\star_{P,2}(\\mu_0)$ would falsify the central inequality (17).","supporting_citations":[],"review_version":1}