{"id":"a8f0b4a8-d881-42b2-8f85-7dceed6b42dd","arxiv_id":"2607.27686","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A persona-agent simulation generates click-intent labels for ads inside AI answers, a distilled evaluator reproduces those labels and beats zero-shot LLM judges on directional tests, and the same score drives a truthful per-click auction.","lead":"LLM-era search ads have no click data, so this paper manufactures a click-intent signal from simulated persona agents and an LLM feature scorer, then uses it to train an ad evaluator and to price ads truthfully. The evaluator passes directional checks and a small human study, but its core signal is not yet validated against real-world clicking.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 103/103 fictional-product test is not evidence of click-intent: Eq. (3) makes Q4 monotonically increasing in relevance, and the test varies only relevance while holding the ad fixed, so the result re-encodes the label's relevance gate.","rationale":"The reader's CONDITIONAL verdict correctly identifies the synthetic label formula in Section 3.1 as the load-bearing weakness. My concern sharpens this: the primary validation claimed for the 'meaningful intent signal' is not merely unvalidated against real clicks—it is internally circular. The fictional-product test holds the advertisement fixed and varies only topical relevance, and Eq. (3) makes the supervision target an increasing function of relevance. Thus 103/103 is close to a tautology given the training target, and the abstract's wording 'generalises without error to 103 fictional products—evidence that it captures semantic intent rather than lexical memorisation' overstates what is shown. The same issue weakens the cross-category swap test, which also primarily moves relevance. The human pairwise study (Section 4.4) does provide some independent evidence that the evaluator's preferences align with human judgments on generated ads, but it is small (five annotators, 100 pairs each) and likewise may reflect relevance and surface quality rather than click intent. The mechanism-design contribution (Section 5) is mathematically correct conditional on any monotone allocation, but its economic meaning depends on x(v) being a true click-probability signal, which the paper concedes it is not. The paper is transparent about these limits in Section 6, so the appropriate verdict remains CONDITIONAL: the framework is publishable as a reproducible ordinal surrogate, but the central claim of predicting click-through intent is not established by the reported experiments. No verdict change is needed; the condition is made more precise.","tokens_in":12338,"tokens_out":5964,"duration_ms":63473,"concrete_test":"Construct a relevance-matched variant of the fictional-product test: for each of 50 fabricated product categories, generate two placements of the same ad that are both contextually relevant (matched relevance scores) but differ in copy quality, novelty, or urgency cues—features that Eq. (3) says should also move Q4. If the evaluator's Q4 scores do not systematically follow the sign of these non-relevance feature differences, then the 103/103 result is fully explained by the relevance gate and the claim that the evaluator captures a multi-feature intent signal is unsupported. A complementary check: compute the partial correlation of evaluator Q4 predictions with the relevance head Q1 after controlling for Q3; if it is near 1, the CTI head is essentially a relevance copy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's strongest claim—that the evaluator 'has internalised a meaningful intent signal rather than a surface correlate'—rests heavily on the fictional-product generalisation test (Section 4.3, test 6). That test generates, for each of 103 fabricated product categories, 'one contextually relevant and one irrelevant placement of the same advertisement.' The Q4 label is defined in Eq. (3) as c = clip(copy·(relevance/5)·(1+λm_p), 0.2, 5.0). Since the advertisement is the same in both placements, copy and all other ad features are identical; the only varying input to the label is relevance. The label is therefore monotonically increasing in relevance by construction, so the evaluator need only learn the relevance gate to achieve 103/103. The test demonstrates semantic relevance sensitivity, not click-through intent. The cross-category swap test (test 1) similarly changes topical relevance, so it also cannot separate intent from relevance. The paper itself concedes in Section 6 that the labels are 'an ordinal surrogate for engagement, not a calibrated probability' and have 'not yet been validated against a large, independent human standard at scale.' But the abstract overstates what the experiments establish: the 103/103 result is largely a check that the evaluator learned the relevance gate baked into its own training target, not evidence that it internalised a meaningful intent signal beyond that gate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the absence of click-through intent (CTI) supervision for advertisements embedded in LLM-generated responses. It constructs synthetic labels via a persona-agent simulation in which an LLM scores six ad features and a hand-specified formula combines them with Big Five personality modulation and a relevance gate (Eqs. 1–3). These labels, together with NaiAD human labels for the other three quality dimensions, supervise a shared-bottleneck evaluator built on a frozen Qwen3-4B backbone with LoRA and an EMD loss. The authors report held-out Q4 Pearson correlation 0.82, a battery of six behavioral perturbation tests, 103/103 fictional-product relevance discrimination, and 86% mean per-annotator pairwise agreement with human judges in Best-of-N selection. They then use the evaluator output as the monotone allocation primitive in a Myerson-style truthful auction, deriving the payment identity, a best-of-k worked example, and extensions to ironing and ε-incentive compatibility. Section 6 is explicit that the agent labels are an uncalibrated ordinal surrogate not yet validated against a large independent human standard.","tokens_in":12658,"tokens_out":6083,"duration_ms":58612,"significance":"If the synthetic CTI label were a valid proxy for real user engagement, the paper would supply a genuinely useful primitive: a deterministic, differentiable, and cheap CTI estimator usable for both mechanism design and as a reward for generative ad optimization. The mechanism-design layer is standard but competently applied, and the persona-agent labelling framework is a reproducible methodological proposal. The paper is also unusually candid about its limitations. However, the current evidence does not establish that the evaluator measures click-through intent rather than the relevance-gated formula used to train it. The headline fictional-product and swap results are largely re-encodings of Eq. (3), and the human validation is small in scale. The contribution is therefore conditional: as a proposal for constructing synthetic intent supervision it is valuable; as a validated CTI evaluator it overclaims. I would like to see the claims scaled back or a substantially stronger external-validation study.","major_comments":[{"comment":"The 103/103 fictional-product result does not support the claim that the evaluator has internalised a meaningful intent signal beyond relevance. In this test the same advertisement is placed in a contextually relevant and an irrelevant context; because Eq. (3) defines the label as c = clip(copy·(relevance/5)·(1+λ m_p), 0.2, 5.0), for fixed copy and personality the label is monotonically increasing in relevance. An evaluator that merely learns the relevance-gate achieves 103/103. The cross-category swap test (test 1) has the same confound: replacing the ad with an off-category product changes only topical relevance. The sentence 'the discrimination is necessarily semantic' should read 'necessarily semantic-relevance.' To support the stronger intent claim, add perturbations that vary copy quality or personality while holding relevance constant, or that require combining relevance with the","section":"§4.3, test 6; Eq. (3)"},{"comment":"The entire downstream chain inherits the validity of the synthetic label formula, but the formula is a hand-assembled construct with free parameters (λ=0.5, trait–feature weights s_{t,f}, K=3, K_personas=30, clip bounds, affine transform (f−1)/4). Section 6 concedes that the agent-grounded labels have not been validated against a large independent human standard and are 'an ordinal surrogate for engagement, not a calibrated probability.' The Section 4.4 human study is small (100 pairs, five annotators, Fleiss κ=0.64) and measures judged click likelihood rather than real engagement. Consequently the abstract's claim of a validated click-intent signal is not supported. The authors should either provide independent validation linking the labels or evaluator to real or externally elicited click preferences, or explicitly restrict all claims to a 'synthetic intent proxy.'","section":"§3.1 and §6"},{"comment":"The sign-certain perturbation battery is internal consistency testing, not external validation. Because the training labels are constructed to be monotonically sensitive to relevance and copy quality, an evaluator that has learned the training formula will by construction lower scores on off-category swaps, filler corruption, and keyword stuffing. The comparison with zero-shot judges is informative about relative sensitivity, but it does not validate that the labels track human click intent. The paper's phrase 'intrinsically validated' is appropriate; the abstract's stronger claim that the evaluator has 'internalised a meaningful intent signal' goes beyond what these tests can show.","section":"§4.3, tests 1–5"},{"comment":"The auction results are mathematically correct for any monotone allocation function x(v), but their economic meaning depends on f being a calibrated or at least valid measure of click probability. Since f is an uncalibrated ordinal score (as acknowledged in Section 6), the per-click price π(v) is a price per synthetic intent unit, not per expected real click. The worked example notes this, but the abstract and Section 5's framing ('truthful pricing') should make equally explicit that truthfulness is with respect to the learned synthetic signal, pending calibration against live engagement data.","section":"§5, Eq. (6)"}],"minor_comments":[{"comment":"Report confidence intervals or significance tests for the 86%/92% agreement figures; with 100 pairs and five annotators the estimates are noisy. Also clarify whether the 86% is the mean of per-annotator agreement or a pooled agreement.","section":"§4.4"},{"comment":"A sensitivity analysis for the free parameters λ and s_{t,f} would strengthen the paper. The authors state the weights can be re-tuned post hoc, but do not show how the evaluator's behavior changes under plausible re-tunings.","section":"§3.1"},{"comment":"The phrase 'generalises without error to 103 fictional products' should be qualified as 'distinguishes relevant from irrelevant placements of the same ad in 103/103 fictional product categories.' As written, it invites an overly broad reading.","section":"Abstract and §4.3"},{"comment":"The explanation for Q1's low Spearman ('two-thirds of items fall within a 0.2-wide band') should state which quantity falls in that band (presumably the label values).","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written and honest paper whose main weakness is the external validity of the synthetic CTI label. I recommend major revision rather than rejection: the methodological core is coherent and reproducible, but the claims need to be either substantially tempered or supported by a larger human/behavioral validation study. The editor may also wish to consider whether a paper whose central measurement is explicitly uncalibrated fits the journal's expectations for empirical validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The real contribution here is the measurement layer: constructing click-intent supervision through persona-agent simulation with Big Five trait weighting, distilling it into a differentiable LoRA-adapted evaluator with an EMD loss, and validating it behaviorally. That combination is new, and the paper is transparent about its construction in a way the related work is not. The held-out fit (Q4 Pearson 0.82 on an agent-grounded target) is credible, the degradation and keyword-stuffing tests show the trained evaluator beats zero-shot judges on shape and monotonicity, and the blind pairwise human study (86% per-annotator agreement, 92% majority, rising with confidence) is decent initial evidence that the evaluator's orderings are sensible. The mechanism section is correct but standard: the payment identity is Myerson, and the best-of-k example is a clean demonstration, not a new theoretical result. The authors mostly say this themselves.\n\nThe soft spot is the one the reader flagged, and it is load-bearing. The Q4 label is defined by a hand-set formula over LLM feature scores with a relevance gate, so every downstream result inherits whatever that formula measures. The fictional-product test varies relevance while holding the ad fixed, and Eq. (3) makes the label monotonically increasing in relevance by construction. Getting 103/103 there means the evaluator learned the gate, not that it internalised click intent beyond relevance. The cross-category swap test has the same confound. The paper's own Section 6 concedes the labels are an ordinal surrogate and have not been validated against an independent human standard at scale, and that the scale is not calibrated to real CTR. None of this sinks the framework as a reproducible, useful ordinal signal for generation selection and mechanism illustration, but it does mean the abstract's 'predicts click-through intent' and 'internalised a meaningful intent signal' are stronger than what the experiments establish. The perturbation N is small (80 items, 15 swap pairs), which is a minor additional concern given how much weight rests on those tests.\n\nThe authors deserve credit for flagging the validation gap themselves, and the mechanism design is internally sound conditional on the evaluator. This paper should get a serious referee, not a desk reject. A referee should push for either independent human validation at scale or live A/B data, and should require the claims in the abstract to be softened to match the evidence. If it ships code and data, I would cite it as the cold-start CTI evaluation framework to beat. Bring it to reading group; it will spark a good argument.","headline":"A serious and unusually honest attempt to build a synthetic click-intent signal for LLM-native ads, with a genuinely new measurement construction; but the headline claims overstate what the evidence supports, since the 103/103 generalization test largely verifies that the evaluator learned the relevance gate baked into its own training label.","tokens_in":13223,"tokens_out":1183,"would_cite":true,"duration_ms":15297,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Synthetic persona-based intent signals can replace missing click logs for AI-embedded ads, enabling a differentiable evaluator and a truthful per-click auction.","keywords":["click-through intent","LLM-native advertising","persona-agent simulation","differentiable evaluator","truthful auction","Earth Mover's Distance","best-of-k allocation","generative ad optimisation"],"falsifier":"Run a live A/B test in which real users are shown a sample of the evaluated ad-embedded responses and compare observed per-item click-through rates against the evaluator's predicted intent rankings; if the ordering or the relative gaps diverge systematically—especially the relevance-gate assumption that off-topic ads are never clicked—the central claim fails. A cheaper pre-registered check: test whether the evaluator's 103/103 fictional-product discrimination survives when the same products receive human click-likelihood ratings.","tokens_in":12122,"feed_emoji":"💰","tokens_out":5934,"duration_ms":60018,"temperature":0.7,"pith_summary":"The paper tries to establish that click-through intent—how likely a user is to click an ad embedded inside an AI-generated answer—can be synthesised when there is no click data, no reliable human annotation, and no honest LLM judge. Its method: simulated personas with different Big Five personality profiles rate objective ad features, and those ratings are combined into a reproducible click-intent label. A compact neural evaluator trained on those labels predicts intent as a smooth, continuous score, passing behavioural tests where zero-shot LLM judges fail and generalising to invented products. The paper then uses that score as the allocation function in an auction and derives the unique payment rule under which truthful bidding is optimal. If true, this supplies the missing measurement layer that both pricing and generative ad optimisation presuppose.","feed_headline":"Persona-simulated intent powers truthful pricing for AI-embedded ads","feed_subtitle":"Trained on simulated user intent, the evaluator beats frontier judges and grounds an auction where honest bids win.","key_machinery":"The central object is the persona-agent click-intent label: an LLM scores six objective ad features, and click intent is computed as copy quality times relevance/5, modulated by a bounded Big Five personality term and clipped to [0.2, 5.0], then averaged over thirty personas. That label is distilled into a shared-bottleneck neural evaluator with four prediction heads—a single shared representation feeding dimension-specific heads—trained with an ordinal loss that penalises prediction errors by their distance on the 1–5 scale. The click-intent head supplies the continuous, deterministic mapping f(q,r); taking the expected maximum of f over k candidate responses makes the allocation x(k) monot","core_discovery":"Click-through intent for LLM-embedded ads can be created synthetically and then learned by a cheap, differentiable evaluator. The authors construct labels through a two-stage persona-agent simulation: an objective scorer rates six ad features, and Big Five personality traits modulate a relevance-gated base score; averaging over thirty sampled personas yields a reproducible 1–5 intent label. Training a shared-bottleneck evaluator with an ordinal score-distance loss on these labels produces a smooth expected-intent output that outperforms zero-shot frontier judges on relevance sensitivity (79% versus 60–67%), tracks dose-response degradation monotonically, generalises without error to 103 fict","pith_inferences":["Beyond the paper's own conclusions, the relevance-gated label construction implies that an ad generator optimised against this signal will be rewarded for being on-topic and well-written above all else; a natural stress test is whether the evaluator can be gamed by superficially relevant but manipulative copy.","The paper leaves implicit that its per-click prices are not real money until the ordinal score is calibrated against live click-through rates; until that calibration exists, the auction example shows the mechanism's structure, not realised revenue.","A transferable consequence: the same persona-modulation-plus-distillation recipe could be applied to other subjective qualities that LLM judges conflate with fluency, such as trustworthiness, humour, or persuasiveness.","A live experiment could test the central premise directly: show real users a sample of the evaluated ad-embedded responses, record their clicks, and check whether observed per-item click order matches the evaluator's predicted intent rankings."],"forward_implications":["Advertisers bidding in this mechanism can be charged per click under a payment rule where truthful bidding is optimal, and the per-click price is guaranteed never to exceed the declared valuation.","Because the evaluator is differentiable and fast, the same intent signal can be used directly as a reward for training ad-generation policies, closing the generative loop that currently lacks supervision.","The evaluator's zero-error generalisation to 103 fictional products indicates the signal tracks semantic intent rather than memorised product-name associations, so it should keep working for genuinely novel ads.","The mechanism prices any measurable allocation, not just best-of-k: ironing and ε-incentive-compatible prices handle learned policies whose allocation is non-monotone, so platforms need not retrain their generators."],"fun_headline_variants":["Simulated personas train AI ad evaluator to beat frontier judges","Synthetic user intent makes AI ad pricing honest and accurate","Persona simulations teach AI to judge ad relevance like humans","AI ad evaluator from simulated clicks enables truthful bidding","Cheap evaluator learns from synthetic users to price LLM ads"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the synthetic label formula—copy quality times relevance divided by five, modulated by Big Five personality weights and clipped—is a faithful stand-in for real users' click-through intent; if that proxy is wrong, the evaluator and the auction are pricing the wrong quantity even though the mechanism itself is internally truthful.","fun_headline_variants_meta":{"raw":{"variants":["Simulated personas train AI ad evaluator to beat frontier judges","Synthetic user intent makes AI ad pricing honest and accurate","Persona simulations teach AI to judge ad relevance like humans","AI ad evaluator from simulated clicks enables truthful bidding","Cheap evaluator learns from synthetic users to price LLM ads"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000578,"raw_usage":{"total_tokens":2568,"prompt_tokens":754,"completion_tokens":1814,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1745}},"tokens_in":498,"tokens_out":1814,"duration_ms":10799,"temperature":1.0,"reasoning_tokens":1745,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:16:59.434585+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a live A/B test in which real users are shown a sample of the evaluated ad-embedded responses and compare observed per-item click-through rates against the evaluator's predicted intent rankings; if the ordering or the relative gaps diverge systematically—especially the relevance-gate assumption that off-topic ads are never clicked—the central claim fails. A cheaper pre-registered check: test whether the evaluator's 103/103 fictional-product discrimination survives when the same products receive human click-likelihood ratings.","supporting_citations":[],"review_version":1}