REVIEW 2 major objections 6 minor 30 references
Markov Information Processes
T0 review · 2 major / 6 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Far-sighted agents who move a Markov state obey recommendations only if the designer prices the future; better models of the dynamics earn rent that grows only logarithmically under excitation.
desk verdict Solid controlled-Markov extension of BCE with recursive design and clean LQG forms; the log-rent theorem is conditional on the paper’s own Assumption 2, which it already flags as Conjecture 11.4. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The dynamic obedience condition (Definition 6.1): the static stage-payoff comparison plus a discounted continuation-value differential Δ that prices how a deviation today changes the law of tomorrow’s state; this is the object that couples constraints across time, reduces to ordinary BCE when the differential vanishes, and becomes a covariance condition with modified interaction matrix in the LQG case.
What would settle it
Simulate or solve a small LQG instance in which excitation is purely endogenous, agents learn from their own possibly off-path deviations, and each tracks the others’ estimators; if the cumulative rent grows faster than any log-squared envelope, the rate claim fails outside the paper’s maintained assumptions.
Extended reading notes
Core claim
Markov Bayes correlated equilibrium is the controlled-Markov generalisation of Bayes correlated equilibrium for far-sighted agents. It is characterised by a dynamic obedience condition that augments the static Bergemann–Morris condition with an explicit continuation-value differential; recommending actions is without loss, the designer’s recursive problem in promised utilities attains an optimum between no-disclosure and first-best, and an agent’s rent from a superior model of the dynamics is non-negative, zero at the known-dynamics benchmark, and ˜O(log T) cumulative under persistent excitation.
Load-bearing premise
The logarithmic cumulative-rent bound assumes agents estimate only from on-path obedient data, the designer injects exogenous persistent excitation, and agents treat the common environment as fixed without tracking one another’s evolving models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops information design for far-sighted, strategically interacting agents whose actions control a persistent Markov state. It introduces the Markov Bayes correlated equilibrium (Markov BCE), characterised by a dynamic obedience condition (Definition 6.1, eq. 5) that augments Bergemann–Morris static obedience with a continuation-value differential Δ^i_h and reduces to the static condition when H=1, B_h=0, or δ=0 (Proposition 6.2). A dynamic revelation principle (Proposition 7.1) justifies action recommendations under unmonitored within-stage deviations. The designer’s problem is cast recursively in promised continuation utilities and solved by an APS-style set-valued backward induction (Algorithm 1); existence and value bounds between no-disclosure and first-best are proved (Propositions 7.2–7.4). In the LQG class, dynamic obedience becomes a covariance condition with a modified interaction matrix Φ̃_h (Theorem 8.2), and the stationary case is a discounted algebraic Riccati fixed point (Proposition 8.3). Part II defines an agent’s rent from a superior model of the transition (non-negative, zero at the known-dynamics benchmark), gives a closed-form LQG expression, and proves a ˜O(log T) cumulative-rent bound under Assumption 2 (Theorem 11.3), leaving the unconditional case as Conjecture 11.4. Two scalar LQG examples (evacuation, power coordination) and a numerical study illustrate the theory.
Significance. If the results hold, the paper supplies a clean controlled-Markov specialisation of dynamic Bayes correlated equilibrium that is missing from both the static BCE literature and the myopic Markov-persuasion stream. The dynamic obedience condition, the recursive promised-utility formulation, and the LQG covariance/Riccati characterisation are genuine contributions that make multi-agent, far-sighted information design tractable in a standard control setting. The rent analysis in Part II is a novel object—an agent’s informational advantage over a committed designer’s model of the dynamics—and the conditional logarithmic bound is carefully derived from self-normalised least squares plus Riccati Lipschitz continuity. Strengths include explicit reduction to the static case, existence under standard compactness/Feller assumptions, a complexity statement for the LQG recursion (Proposition 8.4), and transparent isolation of the open unconditional learning problem as a conjecture. The framework is directly usable for congestion and resource-coordination applications of the type sketched in Section 12.
major comments (2)
- [Abstract; §3 Contributions; Theorem 11.3 / Assumption 2] Theorem 11.3 and the corresponding contributions bullet / abstract sentence state a ˜O(log T) cumulative-rent bound. The proof relies on Assumption 2 (L2)–(L4): on-path (counterfactual) estimation, exogenous persistent excitation λ_min(Λ_t) ≥ λ_0 t, and non-entangled estimators of a common fixed environment. The paper correctly isolates the unconditional case as Conjecture 11.4, but the abstract and the fifth contributions bullet present the logarithmic rate without an explicit qualifier that it holds only under those three modelling restrictions. Because the design tension noted after the theorem (excitation both sustains obedience and accelerates learning) is precisely what makes (L3) non-innocuous, the abstract and contributions list should state the conditioning assumptions in the same sentence as the rate claim so that the result is not over-read.
- [Proposition 7.1 / Remark 1; §12] Proposition 7.1 (dynamic revelation principle) and Remark 1 correctly restrict the principle to the class in which the designer cannot monitor or contract on within-stage actions; deviations propagate only through the realised next state. The two worked examples (evacuation staggering, power reserve coordination) are drawn from settings in which a planner often can observe egress rates or reserve provision. The manuscript does not discuss how much of the implementable set would change if the designer could condition continuation policy on observed within-stage actions, nor whether the Markov BCE recommendations remain approximately optimal under partial monitoring. A short paragraph in §7 or §12 clarifying the scope for the motivating applications would prevent misapplication of the recommendation-policy formulation.
minor comments (6)
- [Abstract; throughout] Abstract and several body paragraphs contain irregular word-internal spaces (e.g., “charact erised”, “de signer’s”, “s et-valued”, “covaria nce”). These appear to be line-break artefacts and should be cleaned for the camera-ready version.
- [§12.1 / Figure 1] Figure 1 (right panel) is described as tracking the c log² t envelope of Theorem 11.3, but the plotted object is cumulative squared estimation error ∑∥θ̂_k−θ∥² rather than cumulative rent. A one-sentence clarification that the estimation error is the driver of the rent bound (Rent_t ≤ C∥G_t∥²) would make the panel self-contained.
- [§8, eq. (10)] In the LQG payoff (10), the term d_i(a_{-i},γ) is said not to affect best responses; this is correct for pure-strategy best responses but should be noted as potentially affecting the designer’s objective when the designer cares about opponents’ payoffs or when public randomisation is used.
- [Proposition 8.3 / 8.4] Proposition 8.3’s contraction condition δ∥A∥² + δL < 1 is used both for existence and for the geometric rate in Proposition 8.4(ii). A brief remark on how L (the Lipschitz constant of the policy-dependent map Π ↦ R_i) can be bounded a priori from the spectrahedron of second moments would help implementers.
- [§12.1 / Table 1] The numerical study imposes E[a_i²] ≤ 4 to restore compactness (consistent with Proposition 7.4). Reporting sensitivity of Table 1 to this bound (or to the asymmetric γ-variances) would strengthen the illustration.
- [§2.5; References] References [24] and [25] are the author’s own continuous-time Stackelberg papers; the positioning in §2.5 is appropriate, but the arXiv identifiers and dates should be double-checked for consistency with the July 2026 manuscript date.
Circularity Check
No circularity: dynamic obedience, LQG covariance/Riccati, and rent bounds are derived from stated primitives and FOCs, not fitted or self-defined into their conclusions.
full rationale
The paper’s load-bearing claims are definitional extensions or direct derivations, not circular. Dynamic obedience (Def. 6.1 / eq. 5) is the static Bergemann–Morris condition plus an explicit continuation differential Δ^i_h; Prop. 6.2 shows reduction to the static case by direct cancellation when H=1, B_h=0 or δ=0. The dynamic revelation principle (Prop. 7.1) is a standard one-shot-deviation + pooling argument under the model’s monitoring structure. The designer’s problem is the APS recursion (Alg. 1) in promised utilities; existence and sandwich bounds (Props. 7.2–7.4) follow from compactness/Feller continuity and feasibility of Bayes–Nash recommendations, not from self-reference. In LQG, Thm. 8.2 obtains the covariance condition by differentiating the quadratic stage-plus-continuation payoff and taking second moments under joint normality; the modified matrix Φ̃_h and the Riccati recursion (Lem. 8.1, Prop. 8.3) are the ordinary controlled-LQ objects, collapsing to the static covariance condition when B=0 or δ=0. Part II’s rent (21) is defined as perceived one-shot deviation gain; Props. 10.1–10.2 and 11.1 show non-negativity and vanishing exactly when the agent’s model matches the designer’s commitment—by the definition of true-model obedience, not by fitting. The ˜O(log T) cumulative-rent theorem (Thm. 11.3) is a standard self-normalised LS + Riccati-Lipschitz argument under the paper’s own Assumption 2; the unconditional case is correctly left as Conjecture 11.4. Self-citations [24,25] supply only the strategic structure of the worked examples and are explicitly not used for any theorem. No fitted parameter is renamed a prediction, no uniqueness theorem is imported from the author, and no ansatz is smuggled via citation. The derivation chain is therefore self-contained against its stated primitives.
Assumptions & free parameters
free parameters (2)
- Numerical LQG primitives (h0, h1, A, B, C, δ, σ²_ε, σ²_γ, action bound E[a_i²]≤4) =
h0=1, h1=0.4, A=0.4, B=0.25, C=0.3, δ=0.9, σ²_ε=0.05, σ²_γ1=1.5, σ²_γ2=0.5, E[a_i²]≤4
- Ridge regularizers λ^{(i)} and prior means for heterogeneous estimators =
λ=1 in the numerical learning panel
assumptions (6)
- standard math One-shot-deviation principle for finite-horizon discounted payoffs with unmonitored within-stage deviations that propagate only through the next state
- domain assumption Designer commits ex ante to a recommendation policy and cannot condition continuation on unobserved within-stage actions (Remark 1)
- standard math Abreu–Pearce–Stacchetti promised-utility recursion and set-valued backward induction
- domain assumption Φ+Φ^⊤ ≻ 0 (strict concavity of the LQG stage game) and joint Gaussianity of (γ,s,a) under linear-Gaussian recommendations
- ad hoc to paper Assumption 2 (L1)–(L4): stationary contractive primitives; on-path estimation; exogenous persistent excitation λ_min(Λ_t)≥λ_0 t; common fixed environment with heterogeneous ridge estimators and sub-Gaussian noise
- standard math Self-normalized least-squares concentration (Abbasi-Yadkori et al.) and local Lipschitz continuity of the Riccati map θ↦Π(θ)
invented entities (3)
-
Markov information process (MKIP)
-
Markov Bayes correlated equilibrium (Markov BCE)
-
Rent from superior knowledge of the dynamics (Renth,i)
Cite this review
Pith. "Pith review of Markov Information Processes." pith.science (2026). https://pith.science/paper/PBF7UZAW
@misc{pith2026260704308,
author = {Pith},
title = {Pith review of: Markov Information Processes},
year = {2026},
howpublished = {\url{https://pith.science/paper/PBF7UZAW}},
note = {Machine review of arXiv:2607.04308}
}
read the original abstract
We study information design when a designer with commitment shapes the information of strategically interacting, far-sighted agents whose actions drive a persistent, controlled Markov state. We introduce the Markov Bayes correlated equilibrium (Markov BCE), the controlled-Markov generalisation of the BCE of Bergemann and Morris (2016), characterised by a dynamic obedience condition that adds a continuation-value term to the static one and reduces to it when actions cannot move the state. Recommending actions is without loss; the designer's problem is recursive in the agents' promised continuation utilities and is solved by a set-valued backward-induction algorithm whose optimum exists and lies between the no-disclosure and first-best values. For linear-quadratic-Gaussian payoffs the obedience condition becomes a covariance condition with a modified interaction matrix, and the stationary case reduces to an algebraic Riccati equation. When agents instead learn the transition, we identify the rent an agent earns from a model of the dynamics sharper than the designer anticipates: it is non-negative, zero at the known-dynamics benchmark, and deterred only by building slack into obedience. Under persistent excitation the cumulative rent grows logarithmically as heterogeneous agents' estimates converge. Two worked examples, in congestion and resource coordination, together with a numerical study illustrate the theory.
Figures
Reference graph
Works this paper leans on
-
[1]
Bayesian persuasio n
Emir Kamenica and Matthew Gentzkow. Bayesian persuasio n. American Economic Review , 101(6):2590–2615, October
-
[2]
URL https://www.aeaweb.org/articles?id=10.1257/aer.101.6.2590
doi: 10.1257/aer.101.6.2590. URL https://www.aeaweb.org/articles?id=10.1257/aer.101.6.2590. 16 Markov Information Processes
-
[3]
Information design: A unified perspective
Dirk Bergemann and Stephen Morris. Information design: A unified perspective. Journal of Economic Literature, 57(1):44– 95, March 2019. doi: 10.1257/jel.20181489. URL https://www.aeaweb.org/articles?id=10.1257/jel.20181489
-
[4]
Bayes correlated equ ilibrium and the comparison of informa- tion structures in games
Dirk Bergemann and Stephen Morris. Bayes correlated equ ilibrium and the comparison of informa- tion structures in games. Theoretical Economics , 11(2):487–522, 2016. doi: 10.3982/TE1808. URL https://onlinelibrary.wiley.com/doi/abs/10.3982/TE1808
doi:10.3982/te1808 2016
-
[5]
Robust predictions i n games with incomplete information
Dirk Bergemann and Stephen Morris. Robust predictions i n games with incomplete information. Econometrica, 81(4):1251– 1308, 2013. doi: 10.3982/ECTA11105. URL https://onlinelibrary.wiley.com/doi/abs/10.3982/ECTA11105
-
[6]
Jibang Wu, Zixuan Zhang, Zhe Feng, Zhaoran Wang, Zhuoran Yang, Michael I. Jordan, and Haifeng Xu. Se- quential information design: Markov persuasion process an d its efficient reinforcement learning, 2022. URL https://arxiv.org/abs/2202.10678
arXiv 2022
-
[7]
Markov persuasion processes: Learning to persuade f rom scratch
Francesco Bacchiocchi, Francesco Emanuele Stradi, Mat teo Castiglioni, Alberto Marchesi, and Nicola Gatti. Markov persuasion processes: Learning to persuade f rom scratch. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P . Koniusz, M. Ghassemi, and N. Chen, edit ors, Advances in Neural Infor- mation Processing Systems , volume 38, pages 70123–70160. Curran...
2025
-
[8]
Markov per suasion processes with endogenous agent be- liefs
Krishnamurthy Iyer, Haifeng Xu Xu, and Y ou Zu. Markov per suasion processes with endogenous agent be- liefs. Proceedings of The 19th Conference On W eb And InterNet Econo mics (WINE 2023) , 2023. URL https://par.nsf.gov/servlets/purl/10532237
arXiv 2023
Show all 30 references
-
[9]
Toward a theory of discounted repeated games with imperfect monito ring
David Pearce, Dilip Abreu, and Ennio Stacchetti. Toward a theory of discounted repeated games with imperfect monito ring. Econometrica, 58(5):1041–1063, 1990. URL https://www.jstor.org/stable/2938299
1990
-
[10]
A continuous- time version of the princi pal: Agent problem
Y uliy Sannikov. A continuous- time version of the princi pal: Agent problem. The Review of Economic Studies , 75(3): 957–984, 2008. ISSN 00346527, 1467937X. URL http://www.jstor.org/stable/20185061
2008
-
[11]
Information desig n in multistage games
Miltiadis Makris and Ludovic Renou. Information desig n in multistage games. Theoretical Economics, 18(4):1475–1509,
-
[12]
URL https://onlinelibrary.wiley.com/doi/abs/10.3982/TE4769
doi: https://doi.org/10.3982/TE4769. URL https://onlinelibrary.wiley.com/doi/abs/10.3982/TE4769
-
[13]
Robert J. Aumann. Correlated equilibrium as an express ion of bayesian rationality. Econometrica, 55(1):1–18, 1987. URL https://www.jstor.org/stable/1911154
1987
-
[14]
On inf ormation design in games
Laurent Mathevet, Jacopo Perego, and Ina Taneva. On inf ormation design in games. Journal of Political Economy , 128(4): 1370–1404, 2020. doi: 10.1086/705332. URL https://www.journals.uchicago.edu/doi/abs/10.1086/705332
2020 doi
-
[15]
Jeffrey C. Ely. Beeps. American Economic Review , 107(1):31–53, 2017. doi: 10.1257/aer.20150218. URL https://www.aeaweb.org/articles?id=10.1257/aer.20150218
2017 doi
-
[16]
Op timal dynamic information provision
J´ erˆ ome Renault, Eilon Solan, and Nicolas Vieille. Op timal dynamic information provision. Games and Eco- nomic Behavior , 104:329–349, 2017. ISSN 0899-8256. doi: https://doi.org /10.1016/j.geb.2017.04.010. URL https://www.sciencedirect.com/science/article/pii/S089982561730074X
2017 doi
-
[17]
Ely and Martin Szydlowski
Jeffrey C. Ely and Martin Szydlowski. Moving the goalpo sts. Journal of Political Economy , 128(2):468–506, 2020. doi: 10.1086/704387. URL https://www.journals.uchicago.edu/doi/abs/10.1086/704387
2020 doi
-
[18]
Laura Doval and Jeffrey C. Ely. Sequential information design. Econometrica, 88(6):2575–2608, 2020. doi: https://doi.org/ 10.3982/ECTA17260. URL https://onlinelibrary.wiley.com/doi/abs/10.3982/ECTA17260
2020 doi
-
[19]
Spear and Sanjay Srivastava
Stephen E. Spear and Sanjay Srivastava. On repeated mor al hazard with discounting. The Review of Economic Studies , 54 (4):599–617, 10 1987. ISSN 0034-6527. doi: 10.2307/229748 4. URL https://doi.org/10.2307/2297484
1987 doi
-
[20]
Dynami c mechanism design: A myersonian ap- proach
Alessandro Pavan, Ilya Segal, and Juuso Toikka. Dynami c mechanism design: A myersonian ap- proach. Econometrica, 82(2):601–653, 2014. doi: https://doi.org/10.3982/ECT A10269. URL https://onlinelibrary.wiley.com/doi/abs/10.3982/ECTA10269
2014 doi
-
[21]
Th e value of the information in the moral hazard setting
Ishak Hajjej, Caroline Hillairet, and Mohamed Mnif. Th e value of the information in the moral hazard setting. Stochastics, 0 (0):1–35, 2025. doi: 10.1080/17442508.2025.2587756. URL https://doi.org/10.1080/17442508.2025.2587756
2025 doi
-
[22]
Improved algorithms for linear stochas- tic bandits
Yasin Abbasi-yadkori, D´ avid P´ al, and Csaba Szepesv´ ari. Improved algorithms for linear stochas- tic bandits. In J. Shawe-Taylor, R. Zemel, P . Bartlett, F. Pe reira, and K. Weinberger, editors, Ad- vances in Neural Information Processing Systems , volume 24. Curran Associ...
2011
-
[23]
Jor dan, and Benjamin Recht
Max Simchowitz, Horia Mania, Stephen Tu, Michael I. Jor dan, and Benjamin Recht. Learning without mixing: Towards a sharp analysis of linear system identification. In S´ ebasti en Bubeck, Vianney Perchet, and Philippe Rigollet, editors , Proceed- ings of the 31st Conference On ...
2018
-
[24]
On the sample complexity of the linear quadr atic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht , and Stephen Tu. On the sample complexity of the linear quadr atic regulator. F oundations of Computational Mathematics , 20(4):633–679, 2020. doi: 10.1007/s10208-019-09426-y. URL https://link.springer.com/article/10.1007...
2020 doi
-
[25]
Fudenberg and D.K
D. Fudenberg and D.K. Levine. The Theory of Learning in Games . Economics Learn- ing and Social Evolution Series. MIT Press, 1998. ISBN 97802 62061940. URL https://mitpress.mit.edu/9780262529242/the-theory-of-learning-in-games/ . 17 Markov Information Processes
1998
-
[26]
Distributionally robust joint informat ion and mechanism design for multi-area power system coordi nation,
Furkan Sezer. Distributionally robust joint informat ion and mechanism design for multi-area power system coordi nation,
-
[27]
URL https://arxiv.org/abs/2606.24015
-
[28]
Continuous-time information design for hurricane evacuation: Disclosure, congestion, and optima l phasing under model uncertainty, 2026
Furkan Sezer. Continuous-time information design for hurricane evacuation: Disclosure, congestion, and optima l phasing under model uncertainty, 2026. URL https://arxiv.org/abs/2606.30320
2026 arXiv
-
[29]
Continuous-time persuasion by filter- ing
Ren´ e A¨ ıd, Ofelia Bonesini, Giorgia Callegaro, and Lu ciano Campi. Continuous-time persuasion by filter- ing. Journal of Economic Dynamics and Control , 176:105100, 2025. doi: 10.1016/j.jedc.2025.105100. URL https://www.sciencedirect.com/science/article/pii/S0165188925000661
2025 doi
-
[30]
Judd, Sevin Yeltekin, and James Conklin
Kenneth L. Judd, Sevin Yeltekin, and James Conklin. Com puting supergame equilibria. Econo- metrica, 71(4):1239–1254, 2003. doi: https://doi.org/10.1111/1 468-0262.t01-1-00445. URL https://onlinelibrary.wiley.com/doi/abs/10.1111/1468-0262.t01-1-00445 . 18
2003 doi
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.