{"id":"0e5902f4-eb57-4773-bfa0-53debaff16f2","arxiv_id":"2508.21162","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a field dataset from a mobile app store, the revenue-optimal exploration prior for new advertisers is 0.1 vs 0.002 for efficiency, a difference that changes revenue by about 32% and reveals a price-thickening channel unique to auction-bandit settings.","lead":"This paper studies how much a mobile app store should explore new advertisers in its sponsored-search auctions, using real data from an Asian app store that runs a Thompson Sampling second-price auction. It finds that the revenue-maximizing exploration level is far higher than the efficiency-maximizing level, because over-exploring entrants raises the price the winner pays, a mechanism absent in pure bandit problems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Revenue-maximizing prior is an untested boundary: footnote 5 admits revenue would rise above 0.1, and truthful bidding is only validated at the observed prior, so the 32% gain is not a robust optimum.","rationale":"The reader's weakest assumption flags the same core issue: truthful bidding is assumed across all counterfactual priors but tested only under the observed prior. My read sharpens this: the manuscript itself admits that revenue would continue to rise above 0.1, and that the range is capped specifically because the truthful-bidding assumption is expected to fail there. This makes the 'revenue-maximizing prior of 0.1' a constrained endpoint rather than an identified optimum, and the 32% gain is therefore not a robust estimate of the value of auction-aware exploration. A simple grid extension under the paper's own assumptions would reveal whether the model even supports 0.1 as a maximum. If it does not, the headline finding is an artifact of the truncation point. If it does, the behavioral concern still remains because the reason for truncation is strategic bidding, and no evidence is provided at the counterfactual prior. This is a genuine weakness, but it is acknowledged by the authors and is addressable with additional estimation and simulation. The paper's broader qualitative conclusion--that the auction side materially changes the exploration problem--is plausible and supported by the decomposition and thickness analyses. I therefore see no reason to change the reader's conditional verdict; the concern strengthens the conditions rather than overturning the paper.","tokens_in":22866,"tokens_out":5832,"duration_ms":65382,"concrete_test":"Re-run Algorithm 1 on the same 1,800 keywords for prior means 0.15, 0.2, 0.5, and 1.0 (with alpha=1 and beta adjusted), holding everything else fixed, and plot revenue versus prior mean. If revenue keeps increasing beyond 0.1, then Figure 8b does not identify a maximizer and the 32% comparison is conditional on an arbitrary truncation point; if revenue peaks at 0.1, the boundary concern would be resolved under the paper's own assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6.1.1 / Figure 8b reports that the revenue-maximizing prior mean is 0.1, at the endpoint of the tested grid, and uses this to claim a 32% revenue gain over the efficiency-maximizing prior of 0.002. Footnote 5 explicitly states that revenue is expected to increase further above 0.1 and that the range is truncated because sufficiently high priors would give incumbents an incentive to shade bids, letting entrants win and then have their scores revised downward. Thus, 0.1 is not identified as an optimum within the model; it is the largest prior for which the authors are willing to maintain Assumption 1 (per-period utility maximization / truthful bidding). The only empirical support for Assumption 1 (Section 5.2.1, Table 2) tests bid shading under the status quo prior mean of 0.1, not under the counterfactual regimes the paper recommends. Algorithm 1 then feeds truthful bids directly into allocation and payment rules, so if incumbents shade at or near the recommended prior, the simulated revenue--and the 32% headline--is overstated. The paper's own conclusion (Section 9) acknowledges this unexplored behavior. The central quantitative claim therefore rests on an unvalidated behavioral invariance at the boundary of the admissible policy space.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies sponsored-search auctions at a large Asian app store, where the platform allocates a single sponsored slot through a second-price auction with quality scores learned by Thompson Sampling with a Beta(1,9) prior. Using impression-level data with full bidder information, the authors estimate advertiser valuations per conversion (average bids) and conversion rates, then simulate counterfactual TS-SP mechanisms that vary the prior mean for entrant quality scores. They report an efficiency-maximizing prior mean of 0.002 and a revenue-maximizing prior mean of 0.1, with about 32% higher revenue under the latter, and attribute the gap to entrants finishing second and raising the winner's payment. They also analyze UCB-style exploration, decompose revenue by entrant rank and market thickness, propose keyword-specific prior customization, and provide a budget-cap robustness check.","tokens_in":23097,"tokens_out":7478,"duration_ms":84533,"significance":"If robust, this would be valuable first empirical evidence on the combined auction--bandit problem, showing that revenue-maximizing exploration can substantially exceed efficiency-maximizing exploration and identifying a concrete mechanism (entrant-second price pressure) that is absent from pure bandit problems. The paper's strengths include an unusually rich dataset with losing bids, a clearly specified simulation algorithm (Algorithm 1), a useful revenue decomposition, and a candid acknowledgment of limitations in footnote 5 and Section 9. The central quantitative claim, however, is currently overstated because the purported revenue optimum sits at the boundary of the policy space, and because the behavioral assumption underlying the simulation is only validated at the observed prior, not at the policies that would be needed to establish a true optimum. These issues are load-bearing for the headline 32% figure but are addressable through reframing and additional analysis.","major_comments":[{"comment":"The paper states that 'the revenue-maximizing prior mean is 0.1' and uses this to claim a 32% revenue gain over the efficiency-maximizing prior of 0.002. Footnote 5, however, explicitly says revenue is expected to increase further above 0.1 and that the range was truncated because sufficiently high priors would create bid-shading incentives. Thus 0.1 is not identified as an optimum; it is the largest prior for which the authors are willing to maintain Assumption 1. The 32% figure is therefore a lower bound on the gain at an arbitrary truncation point, not the gain from the revenue-optimal policy. This wording affects the abstract, Section 1, and Section 9. The authors should reframe the result as a restricted-policy lower bound or extend the model to cover the actual optimum.","section":"§6.1.1, Figure 8b, footnote 5"},{"comment":"Assumption 1 and Proposition 1 (truthful bidding, ba,k,t = va,k,t) are the core input to Algorithm 1, which feeds these bids directly into allocation and payment rules. The empirical support for this assumption (Table 2 and Appendix A) is estimated only under the observed prior mean of 0.1. Footnote 5 and Section 9 acknowledge that at sufficiently high priors incumbents could have dynamic incentives to shade bids, letting entrants win so their scores are revised downward. If the true revenue-optimal prior lies above 0.1, the simulated revenue at that optimum would be computed under exactly the behavioral regime the authors doubt. This is load-bearing because the headline revenue results are conditional on an untested behavioral invariance for the policies needed to establish the optimum. The paper should either provide evidence on bidding behavior under more exploratory priors or substan","section":"§5.2.1, §5.3 (Algorithm 1), §9"},{"comment":"The counterfactual simulation fixes the set of competing ads Ak,t as observed in the data. A higher entrant prior could change entry and participation decisions—attracting more entrants or altering incumbents' incentives to participate—which would feed back into the entrant-second revenue channel. Since the managerial recommendations in Section 8 are stated in terms of changing priors globally, the fixed-participation assumption should be explicitly discussed as a limitation or tested in a robustness check that varies the bidder set.","section":"§5.3, Algorithm 1"}],"minor_comments":[{"comment":"'installment' should be 'installation' (e.g., 'click/installment').","section":"§4.1.1"},{"comment":"The x-axis is on a log scale but the captions do not say so; add axis labels and a note about the log scale to improve readability.","section":"Figure 8"},{"comment":"Clarify whether t in log(t+1) is the global time index or a keyword-specific impression index; the exploration bonus's decay behavior depends on this.","section":"§7.1, Definition 3"},{"comment":"The abbreviation CPC is defined as 'average cost-per-conversion,' but CPC usually stands for cost-per-click; consider using a less ambiguous term such as CPA or spell it out.","section":"§5.2.1, Eq. (7)"},{"comment":"Auer et al. (2002) and Agrawal and Goyal (2012) are listed in the references but not cited in the body; either cite them in Section 2 or remove them.","section":"References"},{"comment":"The unbiasedness claim for the sample analogue estimator depends on random assignment; Appendix B shows IPS gives similar results, but this supporting evidence should be mentioned more prominently in the main text.","section":"§5.2.2"}],"recommendation":"major_revision","confidential_remarks":"This is an empirical marketing/economics paper deposited in cs.GT. The data and mechanism are interesting, and the limitation statements are unusually honest, but the current abstract and headline results overclaim an optimum that is actually a boundary of the admissible policy space. The boundary issue is fixable by reframing the 32% result as a lower bound or by extending the model of bidder behavior; the paper should not be accepted in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this is the first empirical study I've seen that cleanly documents how exploration in an auction-bandit setting acts as an implicit pricing tool. Using data from an Asian app store's sponsored-search auction (single slot, second-price, Thompson Sampling over quality scores), the authors show that the revenue-maximizing prior on entrant CVR is 0.1, two orders of magnitude larger than the efficiency-maximizing prior of 0.002, and that this gap is driven by the entrant landing in second place and pushing up the winner's price. That mechanism—exploration as price pressure, not just learning—is new in the empirical auction literature and is worth taking seriously.\n\nCredit where due: the data are unusually rich (they observe losing bidders, quality score draws, posterior parameters), the counterfactual simulation is clearly specified in Algorithm 1, and the authors do real work to test Assumption 1 (per-period utility maximization) against dynamic bidding: they run a lagged-CPC regression with an IV, and test entrants' overbidding incentives in the appendix. The paper also openly flags its own limitations. Footnote 5 admits the revenue-maximizing prior is at the boundary and revenue would keep rising above 0.1. Section 9 concedes that more aggressive exploration could trigger incumbent bid-shading that the current framework does not model.\n\nSoft spots, in proportion. The boundary issue is the most serious. The headline 32% revenue gain is not measured at an interior optimum; it's the gain at the largest prior the authors are comfortable defending under truthful bidding. That doesn't kill the qualitative claim—the revenue-efficiency divergence is clear across a wide range of priors—but it does mean the specific number is policy-dependent in a way the abstract doesn't convey. Second, there is zero uncertainty quantification on any counterfactual output. For a paper making managerial recommendations, that's a real gap, even if the simulation is deterministic given the estimated primitives. Third, the truthful-bidding test is performed under the status quo prior (mean 0.1), and the counterfactuals stay at or below that prior, so the authors are not extrapolating wildly—but they don't test behavior at priors below 0.1 either, which is where the efficiency-optimal policy sits. The proprietary data and lack of code mean the results can't be independently reproduced, so a referee should weigh that too.\n\nThe theory is a restatement of known special cases, but the empirical contribution stands on its own. I'd send this to a serious referee. The authors are thinking clearly and are unusually honest about what they haven't tested. A good referee should push them on confidence intervals, extending the grid with a bidding-behavior model, and releasing whatever code they can.","headline":"Fresh empirical result on exploration as a pricing tool in auction-bandit settings; the headline 32% gain rests on a boundary prior and an untested behavioral invariance, but the qualitative mechanism is solid.","tokens_in":23648,"tokens_out":2029,"would_cite":true,"duration_ms":19795,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91B26","62L05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to establish that in a real sponsored-search auction where conversion rates are learned online, the exploration policy that maximizes platform revenue is much more aggressive than the policy that maximizes efficiency, and t","keywords":["auctions","multi-armed bandits","Thompson Sampling","sponsored search","cold-start exploration","market thickness","revenue-efficiency tradeoff","quality score prior"],"falsifier":"Run or observe the same sponsored-search auction with the entrant prior set to 0.1 in a field A/B test, and measure whether incumbents' bids fall in response to the higher cost-per-conversion caused by inflated entrant scores; a negative bid response would mean the simulated 32% revenue gain is overstated and the revenue-optimal prior would be lower.","tokens_in":1734,"feed_emoji":"📈","tokens_out":2251,"duration_ms":63443,"temperature":0.7,"pith_summary":"This paper argues that in a combined auction-and-bandit problem, the platform's revenue-maximizing exploration policy is far more aggressive than the efficiency-maximizing policy, and that the gap is driven by the auction side, not by learning alone. Using impression-level data from an Asian app store's Thompson Sampling second-price auction, it compares revenue and efficiency across different priors on new advertisers' conversion rates and finds the revenue-optimal prior is 0.1 while the efficiency-optimal prior is 0.002, yielding about 32% more expected revenue. The reason is that a high entrant prior moves the entrant into second place, compressing the score gap and raising the price the winner pays. This 'entrant-second' channel is absent in pure bandit problems. The paper also finds that customizing the prior by keyword and using a UCB-style algorithm can improve revenue and ease the revenue-efficiency trade-off.","feed_headline":"High entrant priors lift ad-auction revenue by 32 percent","feed_subtitle":"Optimal exploration is a pricing lever: revenue-optimal prior is 0.1 versus 0.002 for efficiency.","key_machinery":"The central object is the Thompson Sampling Second-Price (TS-SP) auction mechanism, in which each advertiser's quality score q is drawn from a Beta distribution with prior parameters (α0, β0) for new entrants, updated only through observed wins and conversions, and in which the highest quality-adjusted bid b·q wins and pays the second-highest quality-adjusted bid per conversion. The mechanism's key working part is the entrant-second channel: an inflated entrant prior elevates the entrant's score to second place, compressing the gap between the top two scores and raising the payment extracted from the winner. This gap is operationalized as market thickness, defined as the average ratio of the","core_discovery":"The paper's central claim is that the revenue-maximizing prior mean for an entrant's conversion rate in a Thompson Sampling second-price auction is 0.1, while the efficiency-maximizing prior mean is 0.002, and that switching from the efficiency-optimal to the revenue-optimal prior raises expected revenue by roughly 32%. The efficiency-optimal prior corresponds to the solution of a weighted bandit problem, where exploration only serves to learn the entrants' true quality. The revenue-optimal prior is far higher because it does additional work: it pushes the entrant's quality-adjusted score into second place, closing the gap between the first and second scores and thereby increasing the second","pith_inferences":["If a platform adopted the revenue-optimal prior in practice, incumbents might eventually learn to shade bids in response to inflated entrant scores, a behavior the paper only tests under the observed prior; modeling that dynamic response would likely lower the optimal prior below 0.1.","The entrant-second channel should generalize to any auction market where the seller learns a quality attribute of new entrants, so cold-start priors in markets such as cloud computing, freelance labor, or content platforms could similarly act as hidden pricing tools.","The near-vertical parts of the customized Pareto frontiers suggest a testable 'free lunch' policy: identify ex ante the keywords with strong incumbents and large second-score gaps, and raise priors only there, capturing revenue with almost no efficiency cost.","A transparency regulation requiring platforms to disclose their quality-score prior parameters would let outsiders estimate how much of the prior is learning-driven versus revenue-driven; comparing disclosed priors to empirically estimated entrant conversion rates would provide a direct audit."],"forward_implications":["Platforms setting cold-start priors should expect the auction side to push optimal exploration well beyond what pure bandit learning would recommend, with the revenue-optimal prior roughly 50 times the efficiency-optimal prior.","The extra revenue from a higher entrant prior comes primarily from auctions where the entrant finishes second, not from entrants winning impressions, so revenue gains can be achieved without large misallocations in thin markets.","Customizing the prior by keyword shifts the revenue-efficiency Pareto frontier outward, yielding about a 5% revenue gain over the best uniform policy even in thick markets, and creating near-vertical segments of the frontier where revenue rises at negligible efficiency cost.","UCB-style exploration raises revenue relative to TS-SP at the same observed exploration level, because UCB maintains a positive quality-score bias for non-winners that raises second prices.","Because the exploration prior functions like an implicit first-price payment mechanism inside a declared second-price auction, high priors are a potential regulatory and transparency issue for auction platforms."],"supporting_citations":[{"why":"Introduces the pay-per-click auction-bandit problem and the price of truthfulness; supplies the theoretical setting this paper tests empirically.","marker":"Devanur and Kakade 2009"},{"why":"Establishes truthful mechanisms with implicit payment computation for auction-bandit settings, motivating the combined problem.","marker":"Babaioff et al. 2015"},{"why":"Provides improved online learning algorithms for CTR prediction in ad auctions and underpins the assumption that conversion rates are fixed over time.","marker":"Feng et al. 2023"},{"why":"Empirical evaluation showing Thompson Sampling performs well in practice, motivating the platform's choice of algorithm.","marker":"Chapelle and Li 2011"},{"why":"Connects market thickness to auction outcomes; the paper invokes this to interpret higher priors as effective thickening of the market.","marker":"Levin and Milgrom 2010"},{"why":"Credible-auctions trilemma result that supports the mechanism by which high entrant scores extract near-first-price payments in a declared second-price auction.","marker":"Akbarpour and Li 2020"},{"why":"Second-price truthful bidding result on which the paper's Proposition 1 (ba,k,t = va,k,t) rests.","marker":"Vickrey 1961"},{"why":"Structural auction estimation approach used to recover private valuations from observed bids.","marker":"Guerre et al. 2000"}],"fun_headline_variants":["Ad-auction revenue jumps 32% with higher entrant prior","Revenue-optimal prior beats efficiency prior by 32% in ad auctions","Thompson sampling priors: 0.1 vs 0.002, 32% revenue gap","Pricing lever: prior choice lifts ad revenue 32%","Higher prior, 32% more ad revenue in auction-bandit"],"cache_read_input_tokens":25344,"weakest_assumption_plain":"The load-bearing assumption is that advertisers maximize per-period utility and bid truthfully under every counterfactual prior, including the revenue-maximizing prior of 0.1, even though the paper's test for strategic bid shading only covers the observed prior of 0.1.","fun_headline_variants_meta":{"raw":{"variants":["Ad-auction revenue jumps 32% with higher entrant prior","Revenue-optimal prior beats efficiency prior by 32% in ad auctions","Thompson sampling priors: 0.1 vs 0.002, 32% revenue gap","Pricing lever: prior choice lifts ad revenue 32%","Higher prior, 32% more ad revenue in auction-bandit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1095,"prompt_tokens":725,"completion_tokens":370,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":286}},"tokens_in":469,"tokens_out":370,"duration_ms":4227,"temperature":1.0,"reasoning_tokens":286,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:31:55.192635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run or observe the same sponsored-search auction with the entrant prior set to 0.1 in a field A/B test, and measure whether incumbents' bids fall in response to the higher cost-per-conversion caused by inflated entrant scores; a negative bid response would mean the simulated 32% revenue gain is overstated and the revenue-optimal prior would be lower.","supporting_citations":[{"cited_title":"The price of truthfulness for pay-per-click auctions","cited_arxiv_id":null,"evidence_quote":"Introduces the pay-per-click auction-bandit problem and the price of truthfulness; supplies the theoretical setting this paper tests empirically."},{"cited_title":"Truthful mechanisms with implicit payment computation","cited_arxiv_id":null,"evidence_quote":"Establishes truthful mechanisms with implicit payment computation for auction-bandit settings, motivating the combined problem."},{"cited_title":"Improved online learning algorithms for ctr prediction in ad auctions","cited_arxiv_id":null,"evidence_quote":"Provides improved online learning algorithms for CTR prediction in ad auctions and underpins the assumption that conversion rates are fixed over time."},{"cited_title":"An empirical evaluation of thompson sampling","cited_arxiv_id":null,"evidence_quote":"Empirical evaluation showing Thompson Sampling performs well in practice, motivating the platform's choice of algorithm."},{"cited_title":"Online advertising: Heterogeneity and conflation in market design","cited_arxiv_id":null,"evidence_quote":"Connects market thickness to auction outcomes; the paper invokes this to interpret higher priors as effective thickening of the market."},{"cited_title":"Credible auctions: A trilemma","cited_arxiv_id":null,"evidence_quote":"Credible-auctions trilemma result that supports the mechanism by which high entrant scores extract near-first-price payments in a declared second-price auction."},{"cited_title":"Counterspeculation, auctions, and competitive sealed tenders","cited_arxiv_id":null,"evidence_quote":"Second-price truthful bidding result on which the paper's Proposition 1 (ba,k,t = va,k,t) rests."},{"cited_title":"Optimal nonparametric estimation of first-price auctions","cited_arxiv_id":null,"evidence_quote":"Structural auction estimation approach used to recover private valuations from observed bids."}],"review_version":1}