{"id":"42d3064f-0524-4ad0-90bb-dc4b07ad27bd","arxiv_id":"2508.19707","paper_version":5,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Higher reputation can make experts recommend risky actions less often, but the paper's key condition is assumed rather than derived, and its single-cutoff characterization is not proven for different ability types.","lead":"This economics paper models how a reputation-conscious expert picks between a risky and a safe recommendation, and asks whether high reputation makes experts more cautious. It claims to show that conservatism can emerge at the top, but the key condition is essentially assumed, and the single-cutoff characterization is not proven for different ability types.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1's common-cutoff claim is unproven and generically false: it proves single-crossing only for type H, and with H more informative than L, the two types' indifference conditions have different roots, so the equilibrium characterization collapses.","rationale":"The reader's weakest assumption is exactly the load-bearing issue. The paper's abstract and intro promise a sharp behavioral characterization, but the proof only establishes each type's best response is a cutoff (single-crossing). The step from 'type-specific cutoff' to 'common cutoff' is a non sequitur. Because p_H and p_L differ, the cutoffs differ. This is not merely a technical gap: the one-to-one success-bonus mapping in Propositions 2, 4, and 6 is defined through the H-type's risky frequency; if L's cutoff differs, the mapping is no longer a pure function of a single threshold, and the design conclusion 'least-cost instrument is protection after failure' (abstract) is unsupported. The relative-diagnosticity condition in Theorem 2 is also essentially the conclusion, but the common-cutoff failure is even more fundamental. I agree with the reader's REJECT verdict.","tokens_in":14289,"tokens_out":3960,"duration_ms":41829,"concrete_test":"Use the Gaussian parameterization in Table 2 (µ0=0, µ1=1, σH=1, σL=1.7, α=π=0.5, V=π², ϕ=0). Drop the common-cutoff restriction and solve the two-type fixed point: find (c_H,c_L) such that ∆_H(c_H;c_H,c_L)=0 and ∆_L(c_L;c_H,c_L)=0, where π_1,1, π_1,0, π_0,0 are computed from Bayes' rule using both cutoffs. If the solution has c_H≠c_L (as expected), Lemma 1 is false. A simpler diagnostic: at any common c, evaluate ∆_H(c) and ∆_L(c) with posteriors generated by that common c; generically both cannot vanish.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central characterization rests on Lemma 1's assertion that both ability types use the same cutoff s*(π). Appendix A, Lemma 2 proves only that ∆_H(s;π) is strictly increasing; it never analyzes ∆_L. For a cutoff c, type θ's indifference condition is p_θ(c)[V(π_1,1)-V(π_0,0)] + (1-p_θ(c))[V(π_1,0)-V(π_0,0)] + ϕ = 0. If both types were to use the same c, we would need p_H(c) = p_L(c) = -(V(π_1,0)-V(π_0,0)+ϕ)/(V(π_1,1)-V(π_1,0)). But under (A2), H is strictly more informative, so p_H(s)≠p_L(s) on a set of positive measure; the equality fails generically. The posterior values themselves depend on both cutoffs, but even allowing for that, the required equality of posterior success probabilities at a common threshold is a measure-zero coincidence. Hence the common-cutoff equilibrium does not exist, and Theorem 1, Theorem 2, and the design results that track only the H-type's experimentation rate lose their foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a one-shot expert-advice model in which a privately informed expert (ability type H or L) recommends a risky or safe action and is evaluated through a belief-based reputational payoff that also allows outcome-contingent transfers. The main claims are: (i) equilibrium advice is a cutoff rule with a single cutoff common to both ability types (Lemma 1, Theorem 1); (ii) under a 'relative-diagnosticity' condition the cutoff is weakly increasing in prior reputation, generating reputational conservatism (Theorem 2); and (iii) a success-only bonus implements any target experimentation rate via a one-to-one mapping (Propositions 2, 4, 6). Extensions cover implementation/gatekeeping, outcome noise, and Gaussian closed forms.","tokens_in":14626,"tokens_out":11484,"duration_ms":121723,"significance":"If correct, the framework would offer a clean, distribution-free characterization of reputational bias in expert advice and a simple design lever through success bonuses. The Gaussian benchmark and the computational details in the appendices are valuable and make the model easy to calibrate. However, the central common-cutoff lemma is not proven and is generically false as stated, and the reputational-conservatism theorem rests on a condition defined through equilibrium objects rather than primitives. These issues undermine the foundation of the characterization and the design results.","major_comments":[{"comment":"Lemma 1 asserts a single cutoff s*(π) for both types, but Appendix A, Lemma 2 proves strict single-crossing only for ∆_H. Under (A2), p_H(s) and p_L(s) are different functions, so the two types' indifference conditions generically have different roots. For a common cutoff c one would need p_H(c)=p_L(c), which is a measure-zero coincidence unless a fixed-point argument is supplied. No such argument is given. Theorem 1 and all subsequent results that track a single cutoff therefore lose their foundation.","section":"Lemma 1 / Appendix A, Lemma 2"},{"comment":"Theorem 2's 'relative-diagnosticity' condition is a sign restriction on the derivative of the equilibrium reputational return R(π)=p_c(V(π_1,1)-V(π_0,0))+(1-p_c)(V(π_1,0)-V(π_0,0)). Along the equilibrium path, ∆_H(s*(π);π)=0 implies R(π)=-ϕ identically, so the total derivative dR/dπ is zero. The proof of Theorem 4 instead uses the partial derivative ∂_π∆_H, and the paper does not define which derivative enters RD. Thus RD is an endogenous object, not a primitive condition, and the theorem is close to a tautology. A primitive condition on signal distributions or V is needed to establish reputational conservatism.","section":"Definition 1 / Theorem 2 / Appendix B"},{"comment":"The experimentation rate ρ(π;β1) is defined as the H-type's risky frequency only, while the equilibrium requires both H and L cutoffs and the posteriors depend on both. If the common-cutoff claim fails, L's behavior changes the posterior mapping, so the claimed one-to-one bonus-to-experimentation relationship is not identified. Even if the common cutoff were restored, the design results must account for both types' behavior to be an equilibrium statement.","section":"Section 5.2 / Propositions 4 and 6"}],"minor_comments":[{"comment":"The full-text title is 'Risky Advice and Reputational Bias', while the arXiv metadata and the provided abstract correspond to 'Paying for Failure in Expert Advice'. The abstract also emphasizes failure protection, which is not developed in the main text. Please align title, abstract, and content.","section":"Title and Abstract"},{"comment":"The proof says 'V is increasing ((A2))', but (A2) is the informativeness ordering; V is increasing and convex under (A4). Please correct the reference.","section":"Appendix A, Lemma 2 proof"},{"comment":"Table 2 reports negative β1 for ρ* = 0.65 and 0.80. Under the stated limited-liability assumption β1≥0, such targets are not implementable. The text notes β1 may become negative but does not reconcile this with the limited-liability design result.","section":"Table 2 / Section 5.3"},{"comment":"The notation p_c(π) ≡ p_θ(s*(π)) is said to be 'well-defined by Lemma 1'. Since the common-cutoff lemma is unproven, the definition is ambiguous until this point is resolved.","section":"Notation after Lemma 1"}],"recommendation":"reject","confidential_remarks":"The central technical claim (common cutoff for both types) is unproven and appears false in generic environments. The reputational-conservatism theorem is stated through an endogenous derivative condition. These are load-bearing, not local issues. The Gaussian calibration and computational appendices are nice, but the model would need substantial reworking, possibly to a single-type setting or a proper equilibrium fixed-point analysis, before it could be published. There is also a mismatch between the title/abstract circulated on arXiv and the full text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for the report. My own reading lines up with yours: the model is clean and the design mapping is neat, but the main characterization has a hole that everything else rests on. The paper’s genuine contribution is a static career-concerns model with a continuous signal and a reduced-form reputational payoff, and it shows how a success-only bonus can implement any experimentation rate given a cutoff. That part is straightforward and correct. The extensions on gatekeeping and noisy outcomes are sensible, and the Gaussian calibration is reproducible.\n\nThe problem is Lemma 1. The text claims a single cutoff s*(π) for both types, but the appendix proves only that ∆_H(s;π) is increasing. Nothing is shown for type L. With H more informative than L, the two types face different success probabilities at any signal, so their indifference conditions have different roots generically. A common cutoff is a measure-zero coincidence, not a theorem. The stress test hits this exactly. Since Theorem 1 is used everywhere—conservatism, comparative statics, the bonus mapping—the paper’s foundation is shaky.\n\nTheorem 2 is also more circular than the authors let on. The RD condition is defined as the derivative of the expected reputational payoff being negative, and the result is just the implicit-function sign. That is a valid conditional statement, but calling it a “mild” condition without deriving it from primitives makes the economic content thin. The Gaussian microfoundation in OA2 is asserted, not proved.\n\nThe comparative statics in Proposition 1 are hand-wavy: the signs for informativeness and success prior are plausible but ignore that posterior beliefs depend on both cutoffs through the equilibrium strategies. The proof in Appendix C is a paragraph.\n\nThere’s also a title/abstract mismatch with the full text, which is sloppy.\n\nOverall: the paper is not a waste of time. The idea that failure protection is a cheaper instrument than success bonuses is interesting, and the design toolkit could be useful if the cutoff were rigorously justified. But as it stands, the main theorem is unproven and likely false. I would send it to a referee only if the authors can fix the common-cutoff issue or recast the equilibrium with type-specific cutoffs. Until then, I wouldn’t cite it.","headline":"Neat framework, but the common-cutoff lemma is unproven and generically false, and Theorem 2 is close to restating its assumption.","tokens_in":15049,"tokens_out":3369,"would_cite":false,"duration_ms":36420,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reputational pressure makes expert advice more conservative at the top, and a success bonus can tune the risk-taking rate.","keywords":["reputation","expert advice","career concerns","delegated experimentation","signal cutoff","success bonus","reputational conservatism","MLRP"],"falsifier":"In the Gaussian benchmark with σL > σH, solve Δ_H(c;π)=0 and Δ_L(c;π)=0 separately over a grid of π. If the two roots differ for any reputation, Lemma 1's common-cutoff premise is false. In real data, estimate the switching signal threshold separately for high- and low-ability experts and test whether they coincide.","tokens_in":14181,"feed_emoji":"🎯","tokens_out":9483,"duration_ms":104517,"temperature":0.7,"pith_summary":"The paper tries to establish that reputational incentives alone can explain and discipline expert risk-taking. It shows that equilibrium advice has a cutoff structure in the expert's private signal and that a mild diagnosticity asymmetry makes the cutoff rise with reputation, so top experts recommend risky actions less often but more accurately. It then proves that a simple success bonus provides a one-to-one lever on the experimentation rate, and that stricter gatekeeping or stronger career concerns push in predictable directions. The stated bigger claim is cost-side: because risky advice comes from the upper tail of confidence, paying for success also rewards recommendations that would happen anyway, whereas protecting against failure concentrates compensation on the marginal advice that actually needs incentivizing. If correct, this gives organizations a practical way to tune boldness with a single transfer.","feed_headline":"Reputation turns experts cautious — one cutoff explains when","feed_subtitle":"A success bonus maps one-to-one onto risk-taking, letting organizations pick any target experimentation rate.","key_machinery":"The risky–safe advantage Δθ(s;π) = ϕ + E[V(π′)|a=1,θ,s] − E[V(π′)|a=0,θ,s] is the central object. Under MLRP it is strictly increasing in the private signal s, so each type's optimal advice is a threshold s*(π). The relative-diagnosticity condition—failures are weakly more revealing than successes at high standing—turns the threshold into a rising function of reputation. A success bonus b adds αb to the risky branch, shifting the threshold downward and generating a one-to-one bonus-to-experimentation mapping.","core_discovery":"The paper's central claim is that a one-shot advisory relationship with belief-based reputation is completely described by a single cutoff in the expert's private signal, and that this cutoff moves in a disciplined way. Under MLRP, each type's risky-safe advantage is strictly increasing, so an expert recommends the risky action if and only if her signal clears a threshold s*(π). When a relative-diagnosticity condition holds—failures are weakly more revealing than successes at high standing—the threshold rises with reputation: high-standing experts are conservative, recommending risk less often but with a higher success probability at the margin. The same cutoff logic yields a one-to-one map","pith_inferences":["The abstract's 'least-cost' claim is an interpretation of margin-selection: the body's theorems prove implementability and monotonicity, not a formal cost comparison. Turning it into a theorem requires a principal objective and a budget constraint.","The single-cutoff assumption is load-bearing: if high- and low-ability types optimally use different thresholds, the conservatism theorem and the one-to-one bonus mapping must be re-derived. A Gaussian numerical check would reveal whether the difference is zero or small.","The relative-diagnosticity condition leaves an observable signature: at high reputation, a failed risky recommendation should move the market's posterior more than a success does. That asymmetry is estimable from data on recommendations and outcomes.","The committee extension yields a testable organizational prediction: raising the approval threshold should make the same experts recommend risky actions less often, visible in recommendation rates before and after a rule change."],"forward_implications":["At high reputation, risky recommendations become scarcer but more accurate; markets should read a risky recommendation from a top expert as a stronger signal of private confidence.","A principal can implement any target experimentation rate in the implementable range by choosing a unique success bonus; no dynamic contract is required.","If failure penalties are feasible, all interior experimentation rates are implementable, and the implementing contracts form an affine line in the bonus and penalty.","Lower implementation probability (stricter gatekeeping) raises the cutoff and lowers experimentation; if gatekeeping loosens as reputation rises, conservatism is amplified.","More informative signals or a higher prior success probability lower the cutoff, while stronger career concerns raise it, so the same transfer scheme works across environments once these primitives are known."],"supporting_citations":[{"why":"Establishes the career-concerns mechanism that makes the expert care about posterior reputation; the paper's V(π') payoff is built on this.","marker":"Holmström, 1999"},{"why":"The closest models of reputational expert advice that the cutoff characterization extends to continuous signals and event-contingent posterior updating.","marker":"Ottaviani and Sørensen, 2006a,b"},{"why":"Provides the delegated-experimentation benchmark where action-generated information and incentives interact, which this static model abstracts from.","marker":"Bolton and Harris, 1999"},{"why":"Defines the cheap-talk benchmark that contrasts with the present setting, where recommendations are implementable actions with observable successes and failures.","marker":"Crawford and Sobel, 1982"},{"why":"Serves as the persuasion-theory comparison: here the informativeness of evidence is endogenous to the chosen action.","marker":"Kamenica and Gentzkow, 2011"},{"why":"Motivates the financial-advice application and the committee/gatekeeping extensions.","marker":"Inderst and Ottaviani, 2012"}],"fun_headline_variants":["One cutoff decides when experts play it safe","Reputation makes experts cautious: a single threshold","Experts' caution tied to reputation: one cutoff rules","Why high-status experts hedge: a simple cutoff","The cutoff that makes experts conservative"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Both ability types are assumed to switch from safe to risky at exactly the same signal value, even though the high-ability type's signal is more informative; if their optimal thresholds differ, the single-cutoff characterization and the design results collapse.","fun_headline_variants_meta":{"raw":{"variants":["One cutoff decides when experts play it safe","Reputation makes experts cautious: a single threshold","Experts' caution tied to reputation: one cutoff rules","Why high-status experts hedge: a simple cutoff","The cutoff that makes experts conservative"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000141,"raw_usage":{"total_tokens":926,"prompt_tokens":595,"completion_tokens":331,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":339,"completion_tokens_details":{"reasoning_tokens":277}},"tokens_in":339,"tokens_out":331,"duration_ms":3545,"temperature":1.0,"reasoning_tokens":277,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:33:06.067951+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the Gaussian benchmark with σL > σH, solve Δ_H(c;π)=0 and Δ_L(c;π)=0 separately over a grid of π. If the two roots differ for any reputation, Lemma 1's common-cutoff premise is false. In real data, estimate the switching signal threshold separately for high- and low-ability experts and test whether they coincide.","supporting_citations":[{"cited_title":"Strategic Experimentation,","cited_arxiv_id":null,"evidence_quote":"Provides the delegated-experimentation benchmark where action-generated information and incentives interact, which this static model abstracts from."},{"cited_title":"Strategic Information Transmission,","cited_arxiv_id":null,"evidence_quote":"Defines the cheap-talk benchmark that contrasts with the present setting, where recommendations are implementable actions with observable successes and failures."},{"cited_title":"Bayesian Persuasion,","cited_arxiv_id":null,"evidence_quote":"Serves as the persuasion-theory comparison: here the informativeness of evidence is endogenous to the chosen action."},{"cited_title":"Financial Advice,","cited_arxiv_id":null,"evidence_quote":"Motivates the financial-advice application and the committee/gatekeeping extensions."}],"review_version":1}