{"id":"e15be078-59ee-4c4a-a1b2-80458af8407c","arxiv_id":"2607.18300","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new Pareto-optimal behavior model shows exactly when sublinear regret is achievable in incentivized exploration with private external information and multiple priors.","lead":"This theory paper extends incentivized exploration—how a platform can nudge selfish users to try new options—to settings where users have private outside information the platform cannot see. It introduces a Pareto-optimal notion of user behavior and shows when sublinear regret is still possible, and when it is not, under Bayesian and non-Bayesian beliefs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's positive claim rests on an unproved posterior analysis for uninformed agents under the retry policy; 'in expectation over states' is not a valid substitute.","rationale":"The reader's weakest assumption correctly identifies that the posterior analysis for uninformed agents under the retry policy is missing and is load-bearing for Theorem 3. I do not fully agree with the specific mechanism stated: an agent cannot directly condition on 'previous agents ignored the message', because the agent only observes the current message and time, and the same message 2 is also sent in the state where R2=1 has already been observed and is optimal. Those confounded paths can make action 2 attractive, so the reader's assertion that conditioning on rejection necessarily makes action 1 better is not established. However, the proof in Appendix E.3 still fails to provide the required posterior calculation: 'in expectation over states' is not the same as conditioning on the agent's full information set, and the state distribution is affected by the retry dynamics. Therefore the central positive claim is not rigorously supported as written. The existing CONDITIONAL verdict is appropriate: the paper should be accepted only after the missing posterior lemma is supplied. My read does not move the verdict.","tokens_in":29052,"tokens_out":31707,"duration_ms":307429,"concrete_test":"Derive the exact posterior of an uninformed agent at time t who receives message 2 under the retry policy of Appendix E.3, including all paths where R2=1 has already been observed (optimal-message paths) and all paths where R2 is unknown (retry after one or more informed rejections). Compute E[R2 - R1 | σt=2, no external info, t] for t=2,3,4,5. If for any t this is ≤0, action 2 is not the unique undominated action and the proposed policy does not guarantee sublinear PO regret; if it is strictly positive for all t, the missing lemma can be added and Theorem 3(3) holds as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central positive claim of Theorem 3 requires that, in the retry phase, every uninformed agent receiving message 2 has action 2 as the unique undominated action. Appendix E.3 asserts that an uninformed agent has a distribution over states {S2,S3,S4} determined by whether previous agents were informed, and that the recommendation is SIC 'in expectation over states'. This is not a posterior analysis. The agent conditions on (t, σt=2). Under the stated retry policy, the event 'the principal is still sending message 2 at time t' is generated by multiple paths: (a) one or more informed agents rejected action 2, so R2 is unknown and R1>0.4; (b) R2=1 was already observed, so message 2 is the optimal recommendation. The mixture over these paths is not independent of R1 and R2, and it is not the same as the unconditional probabilities of S2/S3/S4. The proof never computes this posterior, so it does not establish that action 1 is dominated for the uninformed agent. If action 1 remains undominated, a worst-case PO behavior policy can choose action 1, exploration of action 2 fails, and regret is linear. This gap is load-bearing because it is the only argument for Theorem 3(3). This is a proof gap rather than a demonstrated counterexample, but the manuscript as written does not fill it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends Bayesian incentive-compatible exploration (Kremer et al. 2014) to settings with private external information and to non-Bayesian collections of priors. It separates the principal's signaling policy from the agents' behavior policy and defines three behavior notions: incentive-compatible (IC), strong incentive-compatible (SIC), and Pareto-optimal (PO). The main results characterize when sublinear regret is achievable under each notion. Theorem 3 is the headline positive result: in a Bayesian instance with private external information, no advice policy has sublinear SIC or IC regret, but some general (non-advice) signaling policy has sublinear PO regret. Theorems 2, 4, 5, and 6 give further separation results for full-information, non-Bayesian, and external-information settings, and Theorem 8 analyzes approximate PO behavior.","tokens_in":29403,"tokens_out":14923,"duration_ms":153596,"significance":"The conceptual contribution is valuable: introducing a principal/behavior-policy separation and a Pareto-optimality notion for incentivized exploration is a natural way to model agents who receive information the principal cannot observe. Several constructions (Theorems 2, 4, 5, 6) are explicit and appear correct, and the paper's comparative framework gives a useful map of when different behavioral assumptions permit sublinear regret. However, the central positive claim for external information, Theorem 3(3), rests on an incomplete posterior analysis. The paper does not currently establish that the proposed retry policy makes the recommended action undominated for uninformed agents. Since this theorem drives the claimed IC/PO separation in the external-information regime and is used in Figure 2 and Observation 3, the proof gap is load-bearing.","major_comments":[{"comment":"The proof of the positive PO-regret claim omits the required posterior analysis. Under the retry policy, an uninformed agent receiving message 2 at time t conditions on the event that the principal is still sending message 2, which can occur only if previous informed agents rejected action 2 when R1>0.4. This event is informative about R1 and R2; it shifts posterior mass toward R1>0.4. Appendix E.3 asserts only that the recommended action is SIC 'in expectation over states', but SIC/undominatedness is a pointwise statement given the agent's actual information, not an average over states. The distribution over states {S2,S3,S4} is not computed, and it is not independent of the rewards. Without a calculation of E[R1-R2 | sigma_t=2, t, retry event], the claim that every PO behavior policy must choose action 2 is unsupported. If action 1 remains undominated, a worst-case PO behavior policy c","section":"Section 5.3 / Appendix E.3 (Theorem 3(3))"}],"minor_comments":[{"comment":"Typo: 'uniformed agent' should be 'uninformed agent'.","section":"Appendix E.3"},{"comment":"The threshold 0.799 and the numerical posterior calculations are presented without derivation. Please include the algebra or an explicit reference to Kremer et al. so the reader can verify the claimed SIC constraints.","section":"Appendix B"},{"comment":"The claim that action 2 is never IC when R1>0.4 relies on the definition of IC requiring undominatedness for every possible private signal, not only for uninformed agents. This point should be stated explicitly, as it is central to the argument.","section":"Theorem 3(2) proof"},{"comment":"The proof is compressed. In particular, the induction that the all-action-2 sequence of behavior policies is epsilon-PO should be written out with explicit bounds: conditional on the all-2 history, any message can reveal at most R2, and E[R1-R2 | sigma_t] <= epsilon/2 < epsilon, so no action epsilon-dominates action 2.","section":"Theorem 8 proof"},{"comment":"The arrows marked with '≤' in Figure 2 are not defined in the caption; please state that they denote infima over the corresponding policy classes.","section":"Figure 2 and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for a theory-oriented venue and the framework is promising. The main obstacle is the incomplete proof of Theorem 3(3); the authors should be asked to supply a genuine posterior analysis or to redesign the policy. If the gap is repaired, the paper would make a solid contribution. I do not see a reason to doubt the other main constructions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful framework paper with a clean taxonomy, but the headline positive result for external information (Theorem 3) is not proven as written. The gap is in the posterior analysis for uninformed agents under the retry policy, and it is load-bearing.\n\nWhat's new and good: The paper separates advice policies from general signaling policies and defines Pareto-optimal (PO) behavior as 'choose any undominated action'. That is a sensible way to model agents with private external information and ties. The setting-by-setting feasibility table (Table 1) and the comparisons across SIC/IC/PO are genuinely informative. The examples are crisp. Theorems 2, 4, 5, and 6 are supported by explicit constructions; Theorem 6's induction showing the constant 'play action 3' policy is PO is clear and correct. The proof of Theorem 1 (full information: IC regret ≤ PO regret) is also reasonable, with the auxiliary existence proposition that is a bit terse but fixable.\n\nThe soft spot is Theorem 3(3). The appendix argues that an uninformed agent receiving a repeated recommendation of action 2 has a distribution over states S2/S3/S4, and that 'in expectation over states' the recommendation is SIC. That is not a posterior analysis. The very fact that the principal is still sending the retry message at time t is informative: it means informed agents have rejected the recommendation, which suggests R1 > 0.4, which makes action 1 better in expectation than action 2 for an uninformed agent who conditions on the history of rejections. The proof never computes the conditional distribution, so it does not establish that action 2 is uniquely undominated. If action 1 remains undominated, a worst-case PO behavior policy can choose it, exploration fails, and regret is linear. This is a proof gap rather than a demonstrated counterexample, but it is central: Theorem 3(3) is the only positive result showing that general policies beat advice policies in the external information setting.\n\nThe paper also leans a bit hard on the claim that the PO definition is well-defined; the existence of PO sequences in general settings is assumed more than proven, though Proposition 1 is a start.\n\nBottom line: I'd send this to a serious referee, but the revision should focus on either fixing the Theorem 3 posterior argument or replacing the construction. The framework and the other separations are worth preserving. For a reader in incentivized exploration, this is useful reading; I wouldn't cite the Theorem 3 claim until it's repaired.","headline":"A useful framework and taxonomy, but the central Theorem 3 positive result is missing its posterior analysis and should not be accepted as-is.","tokens_in":29858,"tokens_out":2307,"would_cite":false,"duration_ms":22986,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A principal who sends general signals rather than action recommendations can still achieve sublinear regret when agents hold private information, provided agents merely avoid dominated actions.","keywords":["incentivized exploration","Pareto-optimal behavior","external information","Bayesian incentive compatibility","multi-armed bandits","regret","non-Bayesian agents","tie-breaking"],"falsifier":"Compute, for the Example 3 instance, the posterior distribution of an uninformed agent who has seen the principal recommend action 2 for k consecutive rounds without any agent following it. If that posterior implies E[R1 | history] > E[R2 | history] for any k > 0, then the claim that action 2 remains undominated 'in expectation over states' is false, and the sublinear PO regret proof fails at that step.","tokens_in":28941,"feed_emoji":"📡","tokens_out":5719,"duration_ms":64017,"temperature":0.7,"pith_summary":"This paper argues that the classical Bayesian incentive-compatible exploration model, in which the principal knows everything agents know, breaks down once agents receive private external information. It tries to establish that exploration can still succeed with sublinear regret, but only if the principal stops relying on recommendations agents are expected to follow and instead sends general messages designed so that every action a reasonable agent might take is informative. The key move is to replace obedience-based incentives with Pareto-optimal behavior: agents may pick any action that is not strictly dominated given their knowledge. The paper proves this by exhibiting a two-action Bayesian instance where every advice policy has linear incentive-compatible regret, yet a general policy that repeats a recommendation until an uninformed agent accepts achieves sublinear Pareto-optimal regret. If correct, this means efficient learning does not require controlling agents' choices; it only requires shaping the set of undominated actions.","feed_headline":"Signals, not advice, keep exploration alive under private info","feed_subtitle":"Agents who ignore recommendations still drive sublinear regret when the principal designs messages around undominated actions.","key_machinery":"Pareto-optimal (PO) behavior policy: an agent may select any action that is not strictly dominated under all priors and all histories consistent with earlier agents behaving reasonably. This carries the argument because it lets the principal treat agents as autonomous choosers rather than obedient followers. It is paired with general principal policies, which are allowed to send arbitrary messages rather than only action recommendations. The concrete engine is the 'repeat until accepted' policy: the principal sends the same exploratory signal until some agent who finds the recommended action undominated takes it, exploiting the fact that informed and uninformed types respond differently to t","core_discovery":"The paper's central claim is a separation result: in the presence of private external information, advice policies that must be incentive-compatible (or strongly incentive-compatible) cannot achieve sublinear regret, but general principal policies that send arbitrary messages can, if agents are only assumed to choose undominated actions. Concretely, in a Bayesian instance where action 1 has a known uniform-distributed reward and action 2 is Bernoulli with mean 0.4, and each agent privately observes the history with probability 1/2, no advice policy can explore action 2 with sublinear regret, because an informed agent would reject the recommendation whenever action 1 is known to be better. Ye","pith_inferences":["The paper leaves implicit a natural characterization question: which external-information structures allow the 'repeat until some type accepts' mechanism to work? A plausible answer is that exploration remains possible whenever at least one agent type finds each unexplored action undominated with constant probability.","The proof of Theorem 3 relies on uninformed agents continuing to treat a repeated recommendation as undominated; a repair would be to use a policy whose messages are independent of previous rejections, so that uninformed agents' beliefs match the 'in expectation' calculation. Whether such a policy also achieves sublinear PO regret is a testable extension.","The IC/PO incomparability implies a practical design choice: a platform that cannot tell whether users are obedient rather than autonomous must choose which regret guarantee to optimize for. Measuring real user behavior under private signals could settle which behavioral model is more accurate."],"forward_implications":["If no advice policy can achieve sublinear regret in an external-information instance, then tuning recommendation strategies cannot help; the policy class itself is the bottleneck.","A general signaling policy can achieve sublinear Pareto-optimal regret in the same instance, so the principal's ability to shape information can substitute for incentive compatibility.","Under full information, sublinear PO regret implies sublinear IC regret, but under external information the two notions are incomparable, so designers must know which behavioral model their users satisfy.","Approximate rationality can destroy the exploration guarantee: if agents are only epsilon-approximately Pareto-optimal, there are instances where every general policy has linear regret.","The framework extends to non-Bayesian settings where agents do not share a single common prior, and there the same advice-versus-general-policy distinction determines whether sublinear regret is possible."],"fun_headline_variants":["Private info kills advice policies, but messages still curb regret","After Bayesian advice fails, arbitrary messages still explore","Undominated actions let principals ignore agent advice","No common prior? Messages still drive sublinear regret","Beyond Bayesian: advice useless, messages explore"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Theorem 3's proof assumes that an uninformed agent receiving a repeated recommendation of action 2 continues to treat action 2 as undominated, even after conditioning on the event that previous agents ignored that recommendation, although conditioning on that event is informative and could make action 1 strictly better; the paper does not supply the posterior calculation needed to justify this assumption.","fun_headline_variants_meta":{"raw":{"variants":["Private info kills advice policies, but messages still curb regret","After Bayesian advice fails, arbitrary messages still explore","Undominated actions let principals ignore agent advice","No common prior? Messages still drive sublinear regret","Beyond Bayesian: advice useless, messages explore"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2596,"prompt_tokens":607,"completion_tokens":1989,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":351,"completion_tokens_details":{"reasoning_tokens":1931}},"tokens_in":351,"tokens_out":1989,"duration_ms":14811,"temperature":1.0,"reasoning_tokens":1931,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:24:06.564742+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, for the Example 3 instance, the posterior distribution of an uninformed agent who has seen the principal recommend action 2 for k consecutive rounds without any agent following it. If that posterior implies E[R1 | history] > E[R2 | history] for any k > 0, then the claim that action 2 remains undominated 'in expectation over states' is false, and the sublinear PO regret proof fails at that step.","supporting_citations":[],"review_version":1}