Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Endogenous Vindication: Reputation and Effort in Expert Advice

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that an expert's advice follows a reputation-dependent cutoff: the higher the public reputation, the higher the private-signal threshold for recommending a risky action, because higher reputation makes implementers work ha

desk verdict Genuinely new reputation-effort feedback mechanism and a useful framework, but the central conservatism theorem is not proven: Appendix B has a false lemma and Condition 6 is assumed close to the conclusion. read the letter →

arxiv 2508.19676 v3 pith:FGOZY6TQ submitted 2025-08-27 econ.TH

classification econ.TH
keywords dynamicdelegationexpertadvicereputationfeedbackmoralhazardexperimentationreputationalconservatismcareerconcernsbelief-basedequilibrium
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

An expert's advice is evaluated only when a client acts, and the quality of that evaluation depends on how hard the client works. The paper models a long-lived expert who repeatedly advises short-lived clients, with effort responding to the expert's reputation: trusted experts are heeded and their recommendations are executed more diligently. The paper's central claim is that, under a diagnosticity condition, the competent expert's optimal policy is a cutoff in her private signal that rises with reputation—reputational conservatism—and that reputation itself becomes a submartingale for competent types and a supermartingale for less competent types, hitting trust or distrust regions almost surely. If true, this explains why highly reputed surgeons or analysts may recommend risky actions less often, why early successes can suppress later experimentation, and why more informative tests need not produce more learning overall.

What carries the argument

The engine is the effort response e*(1,π): because effort cost is convex, the implementer chooses effort equal to the posterior probability that the state is good, which rises with the expert's reputation. Higher effort raises the success probability and makes a failure much more informative about the expert's type, while leaving the informativeness of success unchanged. With a convex continuation value V(π), this asymmetry makes the risky-minus-safe payoff difference ΔH(s;π) have decreasing differences in (s,π), so the crossing point s*(π) increases with π. The formal load-bearing condition is Condition 6, log L+(π) ≤ -log L-(π), comparing the log-likelihood jumps of success and failure at

What would settle it

Estimate the log-likelihood jumps after success and failure from panel data on recommendations and outcomes at each reputation level; if at the equilibrium cutoff log L+ > -log L-, Condition 6 fails and the predicted rising cutoff can reverse. Equivalently, estimate the risky-signal threshold as a function of reputation: a range in which the threshold falls with reputation would contradict Theorem 7.

Watch

Extended reading notes

Core claim

At any public reputation π, the High-type expert recommends the risky action exactly when her private signal s is at least a cutoff s*(π) (Theorem 5). Under Condition 6—failures are at least as diagnostic as successes at the cutoff—this cutoff is weakly increasing in π (Theorem 7): more highly reputed experts are more conservative. Reputation dynamics follow a martingale law: posterior beliefs drift up under a competent expert, down under a less competent one, and hit boundary regions with probability one (Theorem 12). Comparative statics are transparent: better private information or a higher good-state prior lowers the cutoff and raises experimentation, whereas more patience raises conserv

Load-bearing premise

Reputational conservatism stands on Condition 6—that, at the equilibrium cutoff, failure is at least as diagnostic of the expert's type as success—which the paper assumes rather than derives from primitives, and the proof also relies on an unproved composition claim about convex value functions and logistic posteriors.

Editorial extensions

If this is right

  • Higher reputation predicts fewer risky recommendations and, conditional on a risky recommendation, higher success rates, because implementation effort is higher.
  • Failures are more damaging to reputation at high reputation than at low reputation; transitory good news reduces experimentation frequency while making subsequent failures more revealing.
  • Improvements in the expert's signal precision or in the prior probability that the risky state is good lower the cutoff; greater patience raises it.
  • Competent experts' reputations drift up and can end in full vindication, while less competent experts drift down; with positive probability even a competent expert ends permanently distrusted, and boundary absorption is almost sure.
  • Individually more informative tests can coexist with less learning overall, because reputational incentives reduce the number of tests the expert is willing to run.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If Condition 6 fails—successes more diagnostic than failures—the model's logic implies the reputation-advice slope could reverse, so estimating the sign of that slope from data is a direct test of the diagnosticity asymmetry.
  • The paper's monitoring result implies a concrete cross-setting prediction: environments with visible implementation effort (checklists, adherence logs) should show a flatter reputation-conservatism relationship than opaque environments, because visible effort reduces the reputational downside of failure.
  • The continuous-time limit suggests a reduced-form estimation strategy: recover the log-odds drift and jump sizes from observed recommendation frequencies and outcome jumps, without solving for the full value function.
  • Boundary absorption means long panels should show reputation distributions that are bimodal conditional on true ability—near full trust or near permanent distrust—rather than mean-reverting; this is a sharp, testable fingerprint of the model.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies a dynamic expert-advice model in which a long-lived expert of unknown ability advises a sequence of short-lived implementers. Implementer effort responds to the expert's current reputation, making outcomes more informative about ability when reputation is high. The expert anticipates this and chooses recommendations accordingly. The authors characterize a recursive belief-based equilibrium in which the high type uses a signal cutoff, show that the cutoff is increasing in reputation ('reputational conservatism'), derive comparative statics in precision, state prior, and patience, characterize reputation dynamics as a submartingale/supermartingale, and provide a Gaussian–quadratic benchmark, a surgery application, committee and monitoring extensions, and a success-bonus policy design. The headline result is Theorem 7, which asserts that the high type's risky-signal cutoff is weakly increasing in the public reputation under A1–A3 and an increasing convex value function.

Significance. The mechanism studied—reputation increases implementation effort, which in turn makes failures more diagnostic—is interesting and potentially important for advice markets, medical decision-making, finance, and organizational design. If Theorem 7 and the dynamic results were established, the paper would provide a clean, testable theory with explicit comparative statics, a tractable Gaussian benchmark, and a calibration recipe. The paper also contains many useful extensions (committees, monitoring, endogenous exit, continuous-time approximation) and a detailed empirical measurement appendix. These are real strengths. However, the central comparative static is supported by proofs that, as written, contain a false composition lemma and rely on an assumption (Condition 6) that is close to the conclusion. The martingale/boundary theorem is also misstated. The contribution is therefore currently not established to the standard required for publication.

major comments (4)
  1. [Appendix B / Section 3.5] There is a fundamental inconsistency in the proof of Theorem 7. In the model, the signal s is private: the implementer's effort e*(1,π) in (3) depends only on the public recommendation and reputation, and the public post-outcome reputations π+(π), π-(π) in (6) do not condition on the expert's private signal. Appendix B, however, repeatedly writes π±(π,s), J±(π,s), and e*(1,π;s), and Lemma 31 states that J- is decreasing in e*(1,π;s) with e* increasing in s. This treats s as publicly observed. Once the s-dependence is removed, the 'effort amplification' channel in Lemma 32(i) and the LLR-asymmetry channel used to obtain decreasing differences of Γ1 collapse. The proof of Theorem 7 as written therefore does not apply to the model that is actually specified.
  2. [Appendix B, Lemma 30] Lemma 30 is false as stated. It claims that Γ0(s,π)=V(σ(logit π+J0(s))) has decreasing differences because σ is increasing and concave and V increasing preserves the property. The logistic σ is not globally concave, and an increasing convex V does not preserve decreasing differences under a non-concave transformation. For the logistic link, the cross derivative is V''(σ)(σ')^2+V'(σ)σ''. Taking V(t)=e^{10t} and σ=0.4 (so x+z=logit(0.4)) makes this quantity strictly positive, so Γ0 has increasing differences at that point. Thus the claimed preservation property is not valid without additional restrictions. Since Lemma 30 is a load-bearing step, Theorem 7 is not established by the argument given.
  3. [Condition 6 / Theorem 7] Condition 6—failures at least as diagnostic as successes at the equilibrium cutoff—is assumed rather than derived from A1–A3 or the Gaussian primitives. It is essentially the mechanism producing reputational conservatism. Moreover, Theorem 7 states its hypotheses as A1–A3 plus increasing convex V, but the convexity of V is itself obtained in Theorem 4 only under Condition 6 (see OA.1, Lemma 64). As stated, Theorem 7 either silently assumes a condition very close to its conclusion or has insufficient hypotheses. The binary-signal reduction of Condition 6 to q_H(1-q_H)≤q_L(1-q_L) illustrates its restrictiveness, but no analogous characterization is supplied for the continuous-signal case. This needs to be made explicit and, ideally, Condition 6 replaced by a primitive condition.
  4. [Theorem 12] Theorem 12 contains a misstatement and an unsupported claim. Part (1) says that under θ=H the process is a submartingale and writes E[π_{t+1}|H_t]=π_t. The equality is the martingale property under the public prior with θ random; conditional on θ=H the correct inequality is >, as Lemma 33 itself proves. Part (3) asserts that the process hits a high-trust or low-trust region with probability one. The proof sketch says that if informative periods cease, the process is 'eventually constant and trivially hits a boundary region,' but an interior constant path does not hit a boundary region. The boundary-hitting claim therefore needs additional assumptions (e.g., experimentation occurs infinitely often on the relevant event) or a different formulation.
minor comments (5)
  1. [Title] The manuscript is submitted under the title 'Endogenous Vindication: Reputation and Effort in Expert Advice,' but the full text uses 'Dynamic Delegation with Reputation Feedback.' This should be reconciled.
  2. [Section 4.2] Theorem 4 refers to 'Condition (6)' before Condition 6 is defined in Section 4.3. The numbering and order need adjustment.
  3. [Sections 4.5 and 4.7] Propositions 9–11 and 14–16 are duplicate statements with identical proofs. This appears to be a manuscript assembly error.
  4. [Figure 1 / OA.3] The caption of Figure 1 reports baseline parameters (0,1,1,1.7,0.5,0.9), while OA.3.3 reports (0,1,0.8,1.6,0.5,0.95). The numerical values should be consistent.
  5. [Notation throughout] The notation e*(1,π) and e*(1,π;s) is used interchangeably; given the model's information structure, e* should not depend on the private signal s. This ambiguity should be resolved in the main text and appendices.

Circularity Check

0 steps flagged · score 0.0 of 10

No construction-level circularity: the main comparative static is an explicit if-then result under a stated diagnosticity condition, and the proof gaps are correctness concerns rather than input-output circularity.

full rationale

I walked the derivation chain from the Bellman equation through cutoff existence (Theorem 5), value convexity (Theorem 4/OA.1), reputational conservatism (Theorem 7), and the comparative statics. The central result is explicitly conditional: Condition 6 requires failures to be at least as diagnostic as successes at the equilibrium cutoff, and Theorem 7 is proved from that condition together with increasing convex V. This is an assumption, not a renamed conclusion: Condition 6 is not defined in terms of the cutoff's monotonicity, and the theorem is a substantive if-then statement rather than a tautology. The proof of Lemma 30 in Appendix B contains an invalid claim (composition with increasing convex V does not generally preserve decreasing differences through the logistic map), and Lemma 32's 'Jensen's inequality' step is asserted rather than proved; these are mathematical-validity gaps, not circularity. The self-citation to Lukyanov et al. (2025) is confined to a pointer about a one-shot companion setup and does not carry the argument. No fitted parameter is relabeled as a prediction, no known result is merely renamed, and no load-bearing premise is justified only by the authors' own prior work. The paper is self-contained relative to its stated assumptions; the honest circularity finding is therefore zero.

Assumptions & free parameters 1 free parameters · 7 assumptions · 0 invented entities

The model introduces no new physical entities, states, or forces. Its central claim rests on standard Bayesian updating, the MLRP/Blackwell assumptions, a convex effort cost, and the ad hoc diagnosticity Condition 6. The Gaussian comparative statics also assume a specific Low-type mixing probability. No fitted numerical parameters are required for the qualitative theorems.

free parameters (1)
  • Low-type mixing probability p(pi) in Gaussian benchmark = not specified; signal-independent in (0,1)
    Ad hoc assumption in Appendix D.4.1 used to keep on-path beliefs Bayes-consistent and make the algebra transparent; the closed-form Gaussian comparative statics (Props. 40, 43, 45) are derived under this choice.
assumptions (7)
  • domain assumption A1: Signals satisfy MLRP and High type Blackwell-dominates Low type.
    Imposed in Section 3.8; underpins cutoff structure and posterior monotonicity.
  • domain assumption A2: Effort cost c is strictly convex, increasing, c(0)=0, c'(0)=0.
    Imposed in Section 3.8; gives unique agent best response and monotone effort.
  • domain assumption A3: Flow utility u is weakly increasing; discount factor delta in (0,1).
    Imposed in Section 3.8; used for value monotonicity and contraction.
  • domain assumption A4: Initial reputation and good-state prior are common knowledge.
    Imposed in Section 3.8; defines the public belief state.
  • ad hoc to paper Condition 6: failures are at least as diagnostic as successes at the equilibrium cutoff.
    Stated in Section 4.3; not derived from primitives in the continuous-signal model. Used to prove convexity of V (Theorem 4) and reputational conservatism (Theorem 7).
  • domain assumption V is increasing and convex on (0,1) in Theorem 7.
    Theorem 7 states V is increasing and convex; convexity is partly derived under Condition 6 but is also directly assumed in the theorem statement.
  • domain assumption Positive signal densities at the cutoff and limited liability for the bonus theorem.
    Theorem 23 and Proposition 24 assume strictly positive densities at the cutoff to get continuous and monotone bonus response.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Endogenous Vindication: Reputation and Effort in Expert Advice." pith.science (2026). https://pith.science/paper/FGOZY6TQ

@misc{pith2026250819676,
  author       = {Pith},
  title        = {Pith review of: Endogenous Vindication: Reputation and Effort in Expert Advice},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FGOZY6TQ}},
  note         = {Machine review of arXiv:2508.19676}
}
read the original abstract

An expert's advice is often evaluated only when a client acts, and the quality of that evaluation depends on implementation effort. We study repeated advice by an expert whose fixed ability is unknown to both the expert and the public. A favorable recommendation may induce a short-lived client to undertake a costly project and choose effort. Higher reputation elicits greater effort, making outcomes more informative about ability. The expert therefore controls whether a reputation-dependent performance test occurs, while the client controls its precision. With diminishing career returns to reputation, the expert may withhold favorable advice even when implementation creates positive surplus. Monotone exposure follows from broad smooth primitives when implementation requires sufficiently high reputation, and from an explicit open family of quadratic environments. Advice then follows a reputation cutoff; stronger career concerns weakly raise it and can stop testing earlier. A low-ability expert eventually enters an absorbing distrust region, while a high-ability expert is permanently distrusted with positive probability and otherwise fully vindicated. Individually more informative tests can therefore coexist with less learning overall because reputational incentives reduce the number of tests conducted.

Figures

Figures reproduced from arXiv: 2508.19676 by the authors.

Figure 1
Figure 1. Simulated reputation dynamics in the Gaussian benchmark. Baseline parameters [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Designing Silence: Peer Feedback under Reputational Concerns

    econ.TH 2025-09 reject novelty 5.0 of 10

    The abstract's concealment-ray theorem and 71.08% optimal revelation threshold are not present in the submitted full text, which is an unrelated paper.

  2. Contrarian Incentives and Costly Social Learning

    econ.TH 2025-08 conditional novelty 5.0 of 10

    In a Gaussian social-learning model, contrarian preferences expand the set of public beliefs where agents invest in private information, as long as the no-signal action is the observed majority.

  3. Paying for Failure in Expert Advice

    econ.TH 2025-08 reject novelty 4.0 of 10

    Higher reputation can make experts recommend risky actions less often, but the paper's key condition is assumed rather than derived, and its single-cutoff characterization is not proven for different ability types.

Reference graph

Works this paper leans on

29 extracted references · 28 canonical work pages · cited by 3 Pith papers

  1. [1]

    I Don’t Know,

    Backus, M. and A. T. Little (2020): “I Don’t Know,”American Political Science Review, 114, 724–743. Birkhäuer, J., J. Gaab, J. Kossowsky, S. Hasler, P. Krummenacher, C. Werner, and H. Gerger (2017): “Trust in the Health Care Professional and Health Outcome: A Meta-Analysis,” PLOS ONE, 12, e0170988

  2. [2]

    Surgical Skill and Complication Rates after Bariatric Surgery,

    Birkmeyer, J. D., J. F. Finks, A. O’Reilly, M. Oerline, A. M. Carlin, A. R. Nunn, J. Dimick, M. Banerjee, N. J. O. Birkmeyer, and M. B. S. Collaborative (2013): “Surgical Skill and Complication Rates after Bariatric Surgery,”The New England Journal of Medicine, 369, 1434–1442. 64

  3. [3]

    Surgeon Volume and Operative Mortality in the United States,

    Birkmeyer, J. D., A. E. Siewers, E. V. A. Finlayson, T. A. Stukel, F. L. Lucas, I. Batista, H. G. Welch, and J. E. Wennberg (2003): “Surgeon Volume and Operative Mortality in the United States,”The New England Journal of Medicine, 349, 2117–2127

  4. [4]

    The Credit Ratings Game,

    Bolton, P., X. Freixas, and J. Shapiro (2012): “The Credit Ratings Game,”Journal of Finance, 67, 85–111

  5. [5]

    Strategic Experimentation,

    Bolton, P. and C. Harris (1999): “Strategic Experimentation,”Econometrica, 67, 349–374

  6. [6]

    Imperfect Monitoring and Impermanent Reputation,

    Cripps, M., G. Mailath, and L. Samuelson (2004): “Imperfect Monitoring and Impermanent Reputation,” Econometrica, 72, 407–432

  7. [7]

    (Bad) Reputation in Relational Contracting,

    Deb, R., M. Mitchell, and M. M. Pai (2022): “(Bad) Reputation in Relational Contracting,” Theoretical Economics, 17, 763–800

  8. [8]

    Ethier, S. N. and T. G. Kurtz (1986): Markov Processes: Characterization and Convergence, Wiley Series in Probability and Mathematical Statistics, New York: John Wiley & Sons

Show all 29 references
  1. [9]

    Optimal Contracts for Experimentation,

    Halac, M., N. Kartik, and Q. Liu (2016): “Optimal Contracts for Experimentation,”The Review of Economic Studies, 83, 1040–1091

  2. [10]

    Physician Communication and Patient Adherence to Treatment: A Meta-Analysis,

    Hall, P. and C. C. Heyde (1980): Martingale Limit Theory and Its Application, New York: Academic Press. Haskard Zolnierek, K. B. and M. R. DiMatteo (2009): “Physician Communication and Patient Adherence to Treatment: A Meta-Analysis,”Medical Care, 47, 826–834. Holmström, B. (1...

  3. [11]

    Financial Advice,

    Inderst, R. and M. Ottaviani (2012): “Financial Advice,”Journal of Economic Literature, 50, 494–512

  4. [12]

    Jacod, J. and A. N. Shiryaev (2003): Limit Theorems for Stochastic Processes, Grundlehren der Mathematischen Wissenschaften, Berlin: Springer, 2 ed

  5. [13]

    Bayesian Persuasion,

    Kamenica, E. and M. Gentzkow (2011): “Bayesian Persuasion,”American Economic Review, 101, 2590–2615

  6. [14]

    Strategic Experimentation with Exponential Bandits,

    Keller, G., S. Rady, and M. Cripps (2005): “Strategic Experimentation with Exponential Bandits,” Econometrica, 73, 39–68

  7. [15]

    When Are Analyst Recommendation Changes Influential?

    Loh, R. K. and R. M. Stulz (2011): “When Are Analyst Recommendation Changes Influential?” Review of Financial Studies, 24, 593–627

  8. [16]

    Risky Advice and Reputational Bias,

    Lukyanov, G., A. Vlasova, and M. Ziskelevich (2025): “Risky Advice and Reputational Bias,” arXiv preprint arXiv:2508.19707, submitted

  9. [17]

    Motivating Innovation,

    Manso, G. (2011): “Motivating Innovation,”Journal of Finance, 66, 1823–1860

  10. [18]

    Political Correctness,

    Morris, S. (2001): “Political Correctness,”Journal of Political Economy, 109, 231–265

  11. [19]

    Will Truth Out?—An Advisor’s Quest to Appear Competent,

    Mylovanov, T. and N. Klein (2017): “Will Truth Out?—An Advisor’s Quest to Appear Competent,” Journal of Mathematical Economics, 72, 112–121. 65

  12. [20]

    Reputational Cheap Talk,

    Ottaviani, M. and P. N. Sørensen (2006): “Reputational Cheap Talk,”RAND Journal of Economics, 37, 155–175

  13. [21]

    A Fraudulent Expert and Short-Lived Customers,

    Ozyurt, S. (2016): “A Fraudulent Expert and Short-Lived Customers,” Tech. rep., Sabanci University Economics Working Paper, available at SSRN: 10.2139/ssrn.2671388

  14. [22]

    Good lies,

    Pavesi, F. and M. Scotti (2022): “Good lies,”European Economic Review, 141, 103965

  15. [23]

    The Wrong Kind of Transparency,

    Prat, A. (2005): “The Wrong Kind of Transparency,”American Economic Review, 95, 862–877

  16. [24]

    A Theory of

    Prendergast, C. (1993): “A Theory of "Yes Men": Honesty and Opinion Herding in Organizations,” American Economic Review, 83, 757–770

  17. [25]

    Impetuous Youngsters and Jaded Old-Timers: Acquiring a Reputation for Learning,

    Prendergast, C. and L. Stole (1996): “Impetuous Youngsters and Jaded Old-Timers: Acquiring a Reputation for Learning,”Journal of Political Economy, 104, 1105–1134

  18. [26]

    Central Limit Theorems for Local Martingales,

    Rebolledo, R. (1980): “Central Limit Theorems for Local Martingales,”Zeitschrift für Wahrschein- lichkeitstheorie und Verwandte Gebiete, 51, 269–286

  19. [27]

    Herd Behavior and Investment,

    Scharfstein, D. and J. Stein (1990): “Herd Behavior and Investment,”American Economic Review, 80, 465–479. Schottmüller, C. (2019): “Too Good to Be Truthful: Why Competent Advisers are Fired,” Journal of Economic Theory, 181, 333–360

  20. [28]

    The Market for Reputations as an Incentive Mechanism,

    Tadelis, S. (2002): “The Market for Reputations as an Incentive Mechanism,”Journal of Political Economy, 110, 854–882

  21. [29]

    Experimentation with Reputation Concerns: Dynamic Signalling with Changing Types,

    Thomas, C. (2019): “Experimentation with Reputation Concerns: Dynamic Signalling with Changing Types,”Journal of Economic Theory, 179, 366–415. 66

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.