{"id":"4c8ba897-99d9-4d73-b7ba-d919e820a665","arxiv_id":"2508.21446","paper_version":7,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a Gaussian social-learning model, contrarian preferences expand the set of public beliefs where agents invest in private information, as long as the no-signal action is the observed majority.","lead":"This paper models sequential social learning where agents pay a fixed cost for private information and receive a bonus for choosing the less popular action. It derives threshold decision rules and shows that, when the no-information choice follows the observed majority, stronger contrarian preferences expand the range of beliefs at which agents buy information.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prop. 4.4's monotonicity of Φ_t in k is unproven: the appendix's B_t(ρ)≥0 inequality does not imply increasing differences once the threshold c_t(k) shifts with k.","rationale":"The reader and I converge on the same soft spot. I checked the appendix claim: A.2 shows B_t(ρ)≥0, but this only establishes monotonicity of the standalone bonus term. The correctness term and the k-dependence of M_t through τ_t(k) are ignored. The total derivative above is the missing calculation. If it is nonnegative, Prop 4.4 is salvageable by replacing the sketch with this lemma, and the rest of the paper (Cor 5.1–5.2, Prop 5.3, and Prop 6.3 conditional on 4.4) is straightforward. If it can be negative, the central claim is false and the verdict should move toward reject. I do not see an independent fatal flaw: Lemma 4.1 is correct, Prop 4.3 is standard, and Corollaries 5.1–5.2 are direct derivatives. The abstract overclaims dynamic results (restart regions, stopped random walks, incomplete learning) that are absent from the body, but that is a framing issue, not the load-bearing technical concern. Therefore the current CONDITIONAL verdict is appropriate; I would not change it.","tokens_in":13792,"tokens_out":12149,"duration_ms":124572,"concrete_test":"Compute the total derivative at fixed ρ: dG_t/dk = M_t(ρ) + [(p_t−1/2)/(c_t(1−c_t))] ρ^{−1/2} [ (1−μ_t)φ(z_0)(1−k(1−2p_t)) − μ_t φ(z_1)(1+k(1−2p_t)) ], with z_0=(τ_t(k)−L_t+ρ/2)/√ρ, z_1=(τ_t(k)−L_t−ρ/2)/√ρ, and check its sign for all ρ≥0 under p_t>1/2, μ_t≥c_t (plus the symmetric case). If this expression is ever negative, Φ_t is not monotone in k and Prop 4.4 fails; if it is nonnegative, insert this lemma into A.2 and the proof is complete. Cross-check with a numerical grid over (μ,p,k,ρ) with c=1, F=0.1 that the set {k:Φ_t>F} is indeed [k_0,∞).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The first clause of Prop 4.4 is correct and nearly trivial: when the no-signal action is the majority action, Pr(a_t=1)∈[0,1] forces B_t(ρ)≥0. The load-bearing step is the next sentence: 'thus k ↦ Φ_t is increasing'. This does not follow. G_t(ρ,k) depends on k both through the k M_t term and through the action threshold c_t(k)=1/2+k(p_t−1/2) (and τ_t(k)), which enters the Gaussian probabilities in the correctness term and in M_t. The appendix (A.2) never computes ∂G_t/∂k at fixed ρ or ∂²G_t/∂ρ∂k; the displayed inequality only shows M_t(ρ)≥M_t(0) for the same k. Pointwise monotonicity of G_t in k—or a genuine increasing-differences condition on (ρ,k)—is what would make sup_ρ G_t monotone in k. Without it, the investment region need not be [k_0,∞), and Prop 6.3's informativeness jump collapses. This is the paper's central advertised result, so the proof gap is decisive even though the rest of the model is coherent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a sequential model of social learning with endogenous information acquisition and nonconformist preferences. Each agent observes the history of actions, chooses a Gaussian private signal precision ρ after paying a fixed cost F and convex cost (c/2)ρ², and then chooses a binary action. Utility is correctness plus k times the unpopularity of the chosen action among predecessors. The paper characterizes the optimal action as a posterior/LLR threshold tilting away from the majority, derives the FOC for precision, and proves existence/regularity of the precision choice. Its central result (Prop 4.4) claims that when the no-signal action coincides with the majority, the value of information Φ_t is weakly increasing in k, so the investment region expands with k. It also derives comparative statics for thresholds and action probabilities and a local welfare/informativeness analysis.","tokens_in":14076,"tokens_out":7721,"duration_ms":71751,"significance":"The model is transparent and fully parametric, yielding closed-form thresholds and a clean characterization of the information-acquisition margin. The central monotonicity insight—contrarian tastes expand experimentation when the default action is the majority—is interesting and economically plausible, and the paper's comparative statics (Corollaries 5.1–5.2) are correct. However, the main proof is currently incomplete, and the abstract advertises dynamic results (stopped random walk, restart region, long-run frequency) that do not appear in the body. Once the proof of Proposition 4.4 is completed and the abstract aligned with the actual content, the paper would make a meaningful contribution to the literature on social learning and information choice.","major_comments":[{"comment":"The proof of Proposition 4.4 is incomplete. The decomposition Φ_t = sup_{ρ≥0} {A_t(ρ)+k B_t(ρ)} treats A_t and B_t as independent of k, but the action probabilities P_1(ρ) depend on k through the threshold c_t(k)=1/2+k(p_t−1/2) and τ_t(k). The displayed inequality (1−2p_t)(P_1(ρ)−1)≥0 only shows that B_t(ρ)≥0 for each fixed k; it does not establish that G_t(ρ,k) has increasing differences in (ρ,k), nor that the supremum is increasing in k. The interval property of the investment region and Proposition 6.3 depend on this monotonicity. A correct argument must compute the total derivative dG_t/dk at fixed ρ (including the threshold-shift term) or use an envelope-theorem argument for Φ_t. Without this, the central claim is not proven.","section":"§4.4 / Appendix A.2"},{"comment":"The abstract claims results that are not in the manuscript: a 'binary-signal benchmark' with a 'restart region' and 'public log odds at information dates form a stopped random walk,' as well as long-run frequency statements about the frequency of correct actions. The main text is entirely Gaussian and presents single-history, local results (Propositions 4.4, 6.1, 6.3). No definition of a restart region or analysis of the dynamic belief process appears. The abstract must be revised to match the actual contributions, or the missing dynamic analysis must be supplied.","section":"Abstract / Introduction"}],"minor_comments":[{"comment":"In the expression for E[1−p_t(a_t)] − E[1−p_t(a_t)]|_{ρ=0}, the dependence of P_1(ρ) on k is not made explicit; this ambiguity obscures the proof gap.","section":"Appendix A.2"},{"comment":"There is a stray '1' in the sentence 'Our approach is different: 1 we add a taste-based bonus...' which should be removed.","section":"Section 2 (Introduction)"},{"comment":"The reference Lu, Meyer, and Rosenbaum (2019) contains an internal note 'Working paper; update authors, venue, and link if you have them' and is not a complete citation.","section":"References"},{"comment":"The claim that the qualitative results extend to any MLRP signal family is asserted without proof. Since the proof of Prop 4.4 relies on the specific form of the threshold c_t(k), the paper should state explicitly which results are proven only for Gaussian signals and which carry over to general MLRP families.","section":"Section 3.1 (footnote 2)"},{"comment":"The statement that 'upper hemicontinuity of the argmax yields right-continuity of ρ⋆_t(k) away from the entry point' is not immediate from UHC alone; a selection argument or further regularity is needed.","section":"Proposition 6.3"}],"recommendation":"major_revision","confidential_remarks":"The reader's report correctly identifies the proof gap in Proposition 4.4. The gap is load-bearing, but an envelope-theorem proof appears plausible, so a major revision rather than rejection is appropriate. The abstract also overstates the paper's content, and the authors should either add the missing dynamic analysis or revise the abstract and framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is worth taking seriously: add a taste for being in the minority, let agents choose whether to buy a signal and at what precision, and see how that reshapes herding and information acquisition. The cleanest parts of the paper are genuinely good. Lemma 4.1 and Proposition 4.2 give a transparent posterior cutoff that tilts against the majority, and Corollaries 5.1–5.2 are simple and correct. Proposition 4.3 (existence and FOC for precision) is standard but fine. The model avoids fixed-point issues by making the contrarian bonus depend only on observed histories, which is a sensible design.\n\nThe soft spot is Proposition 4.4, which underpins the paper's main message. The proof sketch says that because B_t(ρ) ≥ 0, the objective has increasing differences in (ρ, k), so the sup is increasing in k. That does not follow. B_t(ρ) is defined using the threshold c_t(k) = 1/2 + k(p_t − 1/2), so both A_t and B_t depend on k. Pointwise monotonicity of the objective in k for each ρ is what would be needed, but the appendix never shows it, and the cross-partial is never computed. Without this, the investment-region expansion and the informativeness jump in Proposition 6.3 are not established. I suspect the result may be true—there's an envelope-theorem route that might work—but the current proof doesn't get there, and a referee would need to see the actual argument.\n\nThere are two other significant issues. First, the abstract advertises restart regions, stopped random walks, and incomplete learning for general experiment menus, but none of that appears in the body. That's a mismatch a reader will hit immediately. Second, the conclusion claims that contrarian preferences \"lower chosen precision conditional on investing,\" which contradicts Remark 5.4's explicit statement that ρ*_t need not be monotone in k. The text needs to be reconciled.\n\nThe Gaussian–quadratic specification is a limitation, but they acknowledge MLRP extension claims; that's fine for a theory paper. The citation pattern is reasonable and not self-indulgent.\n\nNet: this is a paper that deserves referee time, not a desk reject. The model is coherent and several results are correct, but the central theorem needs a real proof and the abstract needs to match the actual content. I'd send it out, and I'd expect major revision.","headline":"The model and threshold results are solid, but the paper's central comparative static (Prop. 4.4) is not proven as written, and the abstract promises results that the body doesn't deliver.","tokens_in":14544,"tokens_out":10692,"would_cite":true,"duration_ms":91434,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When the default choice follows the crowd, contrarian tastes make agents more willing to pay for private information; extreme contrarianism then erodes accuracy.","keywords":["social learning","information cascades","contrarian preferences","nonconformity","endogenous information acquisition","popularity bonus","Bayesian thresholds","Gaussian signals"],"falsifier":"A numerical check settles it: compute the cross-partial of the agent's expected payoff with respect to (ρ,k) under Gaussian signals, and search over (µ_t, p_t, k, F) with the no-signal action equal to the majority for any case where Φ_t falls as k rises; a single counterexample refutes Proposition 4.4 and the investment-region expansion. A lab or field experiment raising the contrarian bonus while popularity stays observable should show information-purchase rates rising in discrete jumps as the bonus crosses the entry threshold; absence of jumps counts against the interval prediction.","tokens_in":13649,"feed_emoji":"🧠","tokens_out":24421,"duration_ms":203050,"temperature":0.7,"pith_summary":"The paper studies sequential social learning in which each agent decides whether to buy a private signal—and how precise it should be—and also earns a bonus for choosing the less popular action. It tries to establish that this nonconformist taste tilts decision thresholds against the observed majority and, whenever the default no-signal choice already follows that majority, weakly raises the value of private information; the resulting investment region is an interval [k0, ∞) in the contrarian-intensity dimension, so stronger contrarian motives expand the histories in which agents buy signals. If this claim is right, moderate contrarianism counteracts premature herding: it sustains experimentation and improves terminal beliefs and action accuracy in discrete steps, while extreme contrarianism eventually drives the long-run frequency of correct actions down toward one half. The paper also argues that for general experiment menus any positive fixed fee uniformly bounds expected information purchases and, with full-support signals, learning is incomplete. A sympathetic reader would care because this yields closed-form comparative statics for when minority-seeking preferences help or hurt collective learning, with applications to scientific priority races, cultural diffusion, and online platforms.","feed_headline":"Paying for private signals grows with contrarian tastes","feed_subtitle":"When acting on no information already follows the crowd, stronger nonconformity expands who pays for private signals.","key_machinery":"The contrarian bonus b(p_t(a)) = k(1 − p_t(a)) is tied to observed predecessor popularity only, so the equilibrium avoids any fixed point in anticipated popularity and keeps Bayesian updating standard. The action rule is the posterior cutoff c_t = 1/2 + k(p_t − 1/2), a threshold that tilts against the observed majority. The central object is the net value of information Φ_t(µ_t, p_t; k) = sup_ρ {A_t(ρ) + k B_t(ρ) − (c/2)ρ²}, where A_t is the correctness gain and B_t the bonus change from buying a signal. The crux is the sign of B_t: under the majority-aligned baseline, B_t(ρ) ≥ 0 for all ρ, giving increasing differences in (ρ,k) and monotone Φ_t. Gaussian LLR ℓ(s;ρ) = ρ(s − 1/2) puts thresho","core_discovery":"Optimal actions follow the posterior cutoff c_t = 1/2 + k(p_t − 1/2), tilting against the observed majority. The central claim (Prop. 4.4): whenever the no-signal action coincides with the observed majority, the net value of information Φ_t is weakly increasing in contrarian intensity k, so the investment region {k : Φ_t > F} is an interval [k0, ∞) — stronger contrarian motives expand the histories in which agents buy signals, and at the entry threshold actions jump from uninformative to informative (Prop. 6.3). The binary-signal benchmark (abstract) adds that contrarian incentives expand the restart region and improve terminal beliefs and action accuracy in discrete steps until an informati","pith_inferences":["Testable extension: a platform or lab experiment that raises the salience of popularity (or the social reward for distinctiveness) should show information-seeking behavior rising in discrete jumps as the contrarian incentive crosses a history-dependent boundary, not as a smooth gradient.","The two thresholds—the information-cost frontier and the accuracy-decline point—define a design window: visibility of popularity sustains experimentation inside the window and backfires outside it; an open question the paper leaves implicit is where real communities sit relative to these thresholds.","The threshold-tilt comparative statics (Corollaries 5.1–5.2) rest only on MLRP, so the directional prediction that majority choices decline with contrarian intensity is portable beyond Gaussian signals; the value-of-information and precision results are the Gaussian-dependent parts.","An immediate corollary the paper does not draw: with heterogeneous contrarian tastes in a population, aggregate acquisition rates should be increasing in the average taste but kinked wherever mass crosses k0, a micro-data fingerprint of this mechanism."],"forward_implications":["When the no-signal action follows the observed majority, stronger contrarian preferences weakly raise the net value of information, so the investment region is an interval [k0, ∞): more agents pay for signals as k grows.","Higher k shifts the posterior and signal thresholds against the majority, reducing the probability of choosing the majority action in both states (Corollaries 5.1–5.2).","Public action informativeness is zero below k0 and jumps to strictly positive at k0, then weakly increases with k (Lemma 6.2, Proposition 6.3).","In the binary-signal benchmark, contrarian incentives improve terminal beliefs and action accuracy in discrete steps up to an information-cost frontier; beyond a second threshold, long-run accuracy declines toward one half.","For any general experiment menu with a positive fixed fee, expected information purchases are uniformly bounded; with full-support signals, learning is incomplete."],"supporting_citations":[{"why":"supplies the baseline sequential herd model whose herding failure the paper revisits.","marker":"Banerjee (1992)"},{"why":"defines the informational-cascade equilibrium concept that the contrarian preference is shown to disrupt.","marker":"Bikhchandani et al. (1992)"},{"why":"characterizes when observational learning keeps actions informative, grounding the paper's informativeness measure.","marker":"Smith and Sørensen (2000, 2011)"},{"why":"motivates costly endogenous information acquisition in sequential environments.","marker":"Chamley and Gale (1994)"},{"why":"provides the Gaussian precision-choice toolkit used for closed-form thresholds and comparative statics.","marker":"Vives (2008)"},{"why":"supplies the fixed-plus-convex information technology adopted for the acquisition problem.","marker":"Veldkamp (2011)"},{"why":"is the conformity-preference benchmark the paper inverts into nonconformity.","marker":"Bernheim (1994)"},{"why":"models anti-conformist ('hipster') dynamics that the contrarian taste relates to.","marker":"Touboul (2019)"}],"fun_headline_variants":["Contrarian tastes expand who pays for private signals","Stronger nonconformity widens the signal-buying zone","Contrarian incentives grow the market for private info","Nonconformists buy data more when they buck the majority"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central message rests on an unproven step (Appendix A.2): the proof of Proposition 4.4 asserts, without computing the cross-partial and while treating the cutoff c_t as fixed though it shifts with k, that the objective has increasing differences in (ρ,k) when the no-signal action follows the majority; the Gaussian-quadratic specification is assumed throughout, with only a claim that MLRP families behave the same, and the abstract's binary-signal results are not derived in","fun_headline_variants_meta":{"raw":{"variants":["Contrarian tastes expand who pays for private signals","Stronger nonconformity widens the signal-buying zone","Contrarian incentives grow the market for private info","Nonconformists buy data more when they buck the majority"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000229,"raw_usage":{"total_tokens":1280,"prompt_tokens":670,"completion_tokens":610,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":541}},"tokens_in":414,"tokens_out":610,"duration_ms":6050,"temperature":1.0,"reasoning_tokens":541,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:20:38.978691+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A numerical check settles it: compute the cross-partial of the agent's expected payoff with respect to (ρ,k) under Gaussian signals, and search over (µ_t, p_t, k, F) with the no-signal action equal to the majority for any case where Φ_t falls as k rises; a single counterexample refutes Proposition 4.4 and the investment-region expansion. A lab or field experiment raising the contrarian bonus while popularity stays observable should show information-purchase rates rising in discrete jumps as the bonus crosses the entry threshold; absence of jumps counts against the interval prediction.","supporting_citations":[{"cited_title":"and Gale, D","cited_arxiv_id":null,"evidence_quote":"motivates costly endogenous information acquisition in sequential environments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the Gaussian precision-choice toolkit used for closed-form thresholds and comparative statics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the fixed-plus-convex information technology adopted for the acquisition problem."}],"review_version":1}