REVIEW 4 major objections 4 minor 39 references
Evaluating and Pricing Advertisements in AI-Generated Responses
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Synthetic persona-based intent signals can replace missing click logs for AI-embedded ads, enabling a differentiable evaluator and a truthful per-click auction.
desk verdict A serious and unusually honest attempt to build a synthetic click-intent signal for LLM-native ads, with a genuinely new measurement construction; but the headline claims overstate what the evidence supports, since the 103/103 generalization test largely verifies that the evaluator learned the relevance gate baked into its own training label. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the persona-agent click-intent label: an LLM scores six objective ad features, and click intent is computed as copy quality times relevance/5, modulated by a bounded Big Five personality term and clipped to [0.2, 5.0], then averaged over thirty personas. That label is distilled into a shared-bottleneck neural evaluator with four prediction heads—a single shared representation feeding dimension-specific heads—trained with an ordinal loss that penalises prediction errors by their distance on the 1–5 scale. The click-intent head supplies the continuous, deterministic mapping f(q,r); taking the expected maximum of f over k candidate responses makes the allocation x(k) monot
What would settle it
Run a live A/B test in which real users are shown a sample of the evaluated ad-embedded responses and compare observed per-item click-through rates against the evaluator's predicted intent rankings; if the ordering or the relative gaps diverge systematically—especially the relevance-gate assumption that off-topic ads are never clicked—the central claim fails. A cheaper pre-registered check: test whether the evaluator's 103/103 fictional-product discrimination survives when the same products receive human click-likelihood ratings.
Extended reading notes
Core claim
Click-through intent for LLM-embedded ads can be created synthetically and then learned by a cheap, differentiable evaluator. The authors construct labels through a two-stage persona-agent simulation: an objective scorer rates six ad features, and Big Five personality traits modulate a relevance-gated base score; averaging over thirty sampled personas yields a reproducible 1–5 intent label. Training a shared-bottleneck evaluator with an ordinal score-distance loss on these labels produces a smooth expected-intent output that outperforms zero-shot frontier judges on relevance sensitivity (79% versus 60–67%), tracks dose-response degradation monotonically, generalises without error to 103 fict
Load-bearing premise
The load-bearing premise is that the synthetic label formula—copy quality times relevance divided by five, modulated by Big Five personality weights and clipped—is a faithful stand-in for real users' click-through intent; if that proxy is wrong, the evaluator and the auction are pricing the wrong quantity even though the mechanism itself is internally truthful.
Editorial extensions
If this is right
- Advertisers bidding in this mechanism can be charged per click under a payment rule where truthful bidding is optimal, and the per-click price is guaranteed never to exceed the declared valuation.
- Because the evaluator is differentiable and fast, the same intent signal can be used directly as a reward for training ad-generation policies, closing the generative loop that currently lacks supervision.
- The evaluator's zero-error generalisation to 103 fictional products indicates the signal tracks semantic intent rather than memorised product-name associations, so it should keep working for genuinely novel ads.
- The mechanism prices any measurable allocation, not just best-of-k: ironing and ε-incentive-compatible prices handle learned policies whose allocation is non-monotone, so platforms need not retrain their generators.
Reading between the lines
- Beyond the paper's own conclusions, the relevance-gated label construction implies that an ad generator optimised against this signal will be rewarded for being on-topic and well-written above all else; a natural stress test is whether the evaluator can be gamed by superficially relevant but manipulative copy.
- The paper leaves implicit that its per-click prices are not real money until the ordinal score is calibrated against live click-through rates; until that calibration exists, the auction example shows the mechanism's structure, not realised revenue.
- A transferable consequence: the same persona-modulation-plus-distillation recipe could be applied to other subjective qualities that LLM judges conflate with fluency, such as trustworthiness, humour, or persuasiveness.
- A live experiment could test the central premise directly: show real users a sample of the evaluated ad-embedded responses, record their clicks, and check whether observed per-item click order matches the evaluator's predicted intent rankings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses the absence of click-through intent (CTI) supervision for advertisements embedded in LLM-generated responses. It constructs synthetic labels via a persona-agent simulation in which an LLM scores six ad features and a hand-specified formula combines them with Big Five personality modulation and a relevance gate (Eqs. 1–3). These labels, together with NaiAD human labels for the other three quality dimensions, supervise a shared-bottleneck evaluator built on a frozen Qwen3-4B backbone with LoRA and an EMD loss. The authors report held-out Q4 Pearson correlation 0.82, a battery of six behavioral perturbation tests, 103/103 fictional-product relevance discrimination, and 86% mean per-annotator pairwise agreement with human judges in Best-of-N selection. They then use the evaluator output as the monotone allocation primitive in a Myerson-style truthful auction, deriving the payment identity, a best-of-k worked example, and extensions to ironing and ε-incentive compatibility. Section 6 is explicit that the agent labels are an uncalibrated ordinal surrogate not yet validated against a large independent human standard.
Significance. If the synthetic CTI label were a valid proxy for real user engagement, the paper would supply a genuinely useful primitive: a deterministic, differentiable, and cheap CTI estimator usable for both mechanism design and as a reward for generative ad optimization. The mechanism-design layer is standard but competently applied, and the persona-agent labelling framework is a reproducible methodological proposal. The paper is also unusually candid about its limitations. However, the current evidence does not establish that the evaluator measures click-through intent rather than the relevance-gated formula used to train it. The headline fictional-product and swap results are largely re-encodings of Eq. (3), and the human validation is small in scale. The contribution is therefore conditional: as a proposal for constructing synthetic intent supervision it is valuable; as a validated CTI evaluator it overclaims. I would like to see the claims scaled back or a substantially stronger external-validation study.
major comments (4)
- [§4.3, test 6; Eq. (3)] The 103/103 fictional-product result does not support the claim that the evaluator has internalised a meaningful intent signal beyond relevance. In this test the same advertisement is placed in a contextually relevant and an irrelevant context; because Eq. (3) defines the label as c = clip(copy·(relevance/5)·(1+λ m_p), 0.2, 5.0), for fixed copy and personality the label is monotonically increasing in relevance. An evaluator that merely learns the relevance-gate achieves 103/103. The cross-category swap test (test 1) has the same confound: replacing the ad with an off-category product changes only topical relevance. The sentence 'the discrimination is necessarily semantic' should read 'necessarily semantic-relevance.' To support the stronger intent claim, add perturbations that vary copy quality or personality while holding relevance constant, or that require combining relevance with the
- [§3.1 and §6] The entire downstream chain inherits the validity of the synthetic label formula, but the formula is a hand-assembled construct with free parameters (λ=0.5, trait–feature weights s_{t,f}, K=3, K_personas=30, clip bounds, affine transform (f−1)/4). Section 6 concedes that the agent-grounded labels have not been validated against a large independent human standard and are 'an ordinal surrogate for engagement, not a calibrated probability.' The Section 4.4 human study is small (100 pairs, five annotators, Fleiss κ=0.64) and measures judged click likelihood rather than real engagement. Consequently the abstract's claim of a validated click-intent signal is not supported. The authors should either provide independent validation linking the labels or evaluator to real or externally elicited click preferences, or explicitly restrict all claims to a 'synthetic intent proxy.'
- [§4.3, tests 1–5] The sign-certain perturbation battery is internal consistency testing, not external validation. Because the training labels are constructed to be monotonically sensitive to relevance and copy quality, an evaluator that has learned the training formula will by construction lower scores on off-category swaps, filler corruption, and keyword stuffing. The comparison with zero-shot judges is informative about relative sensitivity, but it does not validate that the labels track human click intent. The paper's phrase 'intrinsically validated' is appropriate; the abstract's stronger claim that the evaluator has 'internalised a meaningful intent signal' goes beyond what these tests can show.
- [§5, Eq. (6)] The auction results are mathematically correct for any monotone allocation function x(v), but their economic meaning depends on f being a calibrated or at least valid measure of click probability. Since f is an uncalibrated ordinal score (as acknowledged in Section 6), the per-click price π(v) is a price per synthetic intent unit, not per expected real click. The worked example notes this, but the abstract and Section 5's framing ('truthful pricing') should make equally explicit that truthfulness is with respect to the learned synthetic signal, pending calibration against live engagement data.
minor comments (4)
- [§4.4] Report confidence intervals or significance tests for the 86%/92% agreement figures; with 100 pairs and five annotators the estimates are noisy. Also clarify whether the 86% is the mean of per-annotator agreement or a pooled agreement.
- [§3.1] A sensitivity analysis for the free parameters λ and s_{t,f} would strengthen the paper. The authors state the weights can be re-tuned post hoc, but do not show how the evaluator's behavior changes under plausible re-tunings.
- [Abstract and §4.3] The phrase 'generalises without error to 103 fictional products' should be qualified as 'distinguishes relevant from irrelevant placements of the same ad in 103/103 fictional product categories.' As written, it invites an overly broad reading.
- [Table 1] The explanation for Q1's low Spearman ('two-thirds of items fall within a 0.2-wide band') should state which quantity falls in that band (presumably the label values).
Circularity Check
The 103/103 fictional-product and cross-category swap results are entailed by the relevance gate in the paper's own label formula; the 'internalised intent signal' claim is therefore partly a restatement of the training target, with only partial external support from the small pairwise human study.
-
self definitional
[Section 3.1, Eqs. (2)-(3); Section 4.3, Test 6 (Fictional-product generalisation)]
"basep = copy·(relevance/5), cp = clip(basep ·(1 + λ mp), 0.2, 5.0) ... The gate ensures that an irrelevant advertisement cannot earn high click-through intent regardless of copy quality or personality. ... for each of 103 fabricated product categories ... one contextually relevant and one irrelevant placement of the same advertisement were generated. The evaluator assigns the relevant placement the higher click-intent score in all 103 cases."
In Test 6 the advertisement is identical in the relevant and irrelevant placements, so copy, personality weights, and all non-relevance features are fixed. The only input that varies in the training label of Eq. (3) is relevance, and c_p is monotonically increasing in relevance by construction. Any model that fits its supervision must therefore rank the relevant placement higher; the 103/103 score and the 79% swap pass rate are checks that the evaluator learned the relevance gate built into its own labels, not independent evidence of click-through intent. The paper's Section 6 concession that the scores are 'an ordinal surrogate for engagement, not a calibrated probability' underlines that this test validates the synthetic target, not a real-world intent signal.
full rationale
The pricing derivation is not circular: x(k) = E[max f(r_j)] is monotone by construction as an expected maximum, and the payment identity p(v) = v x(v) - ∫ x(t)dt is the standard TIE characterization; this is a theorem, not an input. The circularity is concentrated in the validation claims about click-intent. With the ad fixed, the agent label c_p depends on relevance only through the product copy·(relevance/5), so the training target itself orders the relevant placement above the irrelevant one. The evaluator is trained on these labels with an EMD loss, so it can pass the fictional-product and cross-category swap tests by learning its own target's relevance gate. The paper even concedes in Section 6 that the evaluator 'tracks semantic relevance and text quality first' with psychological features as a bounded modulation. The only external anchor is the blind pairwise human study, which is small (100 pairs, five annotators) and is described by the authors as 'initial evidence' rather than scale validation. The NaiAD self-citation about uncalibratable human CTI is load-bearing for the choice of synthetic labels but is an empirical premise rather than a definitional reduction; it contributes to the weakness but does not alone create the circularity. Overall, some 'predictions' reduce by construction, but the mechanism-design result and the human pairwise anchor give the paper partial independent content, so the score is 6, not higher.
Assumptions & free parameters
free parameters (5)
- λ (personality modulation strength) =
0.5
- signed trait–feature weights s_{t,f} =
sign-fixed, magnitude unspecified (appears to be ±1)
- K=3 and K_personas=30 =
3, 30
- clip bounds 0.2 and 5.0 =
0.2, 5.0
- affine transform (f−1)/4
assumptions (7)
- domain assumption LLM feature scoring in Stage 1 is a reliable, consistent measure of objective ad quality features.
- ad hoc to paper Relevance gate: an off-topic ad cannot receive high click-through intent regardless of copy quality or personality.
- domain assumption Big Five trait–feature pairings and their signed effects apply to LLM-native advertising.
- domain assumption Normative Big Five population distributions from cited agent simulation work yield realistic personas.
- domain assumption Sign-certain perturbation directions are known a priori (foreign ad lower, keyword stuffing lower, degradation lower).
- domain assumption NaiAD VC-PPI human labels for Q1–Q3 are valid supervision.
- standard math Standard mechanism-design characterization: a single-parameter mechanism is TIE iff x(v) is monotone, with payment identity p(v)=vx(v)−∫x.
invented entities (2)
-
Persona-agent click-intent labels
-
Click-through intent (CTI) as a continuous latent score on 1–5
Cite this review
Pith. "Pith review of Evaluating and Pricing Advertisements in AI-Generated Responses." pith.science (2026). https://pith.science/paper/KNWVF2KI
@misc{pith2026260727686,
author = {Pith},
title = {Pith review of: Evaluating and Pricing Advertisements in AI-Generated Responses},
year = {2026},
howpublished = {\url{https://pith.science/paper/KNWVF2KI}},
note = {Machine review of arXiv:2607.27686}
}
read the original abstract
As search increasingly shifts toward LLM-driven answer engines, advertising is becoming embedded within the generated response itself and should therefore be evaluated for both user utility and commercial value. The key challenge is click-through intent: behavioural logs are unavailable, human annotation resists calibration, and frontier LLM judges conflate intent with linguistic fluency. These gaps compound, as principled pricing presupposes a continuous intent signal, while generating such a signal presupposes supervision that is currently unavailable. We construct the missing supervision through a psychologically grounded agent simulation framework, and distil it into a parameter-efficient evaluator that predicts click-through intent, together with the three companion dimensions of ad quality, as smooth, differentiable estimates. Validated through sign-certain behavioural perturbations, the evaluator surpasses frontier zero-shot judges on relevance sensitivity (79% versus 60-67%), tracks graded content degradation, generalises without error to 103 fictional products, and agrees with human preference in 86% of pairwise judgements across five annotators, with agreement rising in the evaluator's confidence. Upon its estimates we build the pricing layer directly, deriving the unique payment rule under which truthful bidding is optimal, demonstrating it on a best-of-k allocation, and extending the mechanism to non-monotone allocations. The same differentiable signal stands ready as a training objective for ad generation.
Figures
Reference graph
Works this paper leans on
-
[1]
Zhang, Yihang and Huang, Zimeng and Zhai, Ren and Kang, Yipeng and Wang, Tonghan , journal=
-
[2]
Influence of personality traits on generation
Saha, Partha and Sengupta, Angan and Gupta, Priya , journal=. Influence of personality traits on generation. 2024 , publisher=
2024
-
[3]
Simulating Prosocial Behavior and Social Contagion in
Zhou, Yujia and Wang, Hexi and Ai, Qingyao and Wu, Zhen and Liu, Yiqun , journal=. Simulating Prosocial Behavior and Social Contagion in
-
[4]
Scaling Personality Control in
Cho, Gunhee and Cheong, Yun-Gyung , journal=. Scaling Personality Control in
-
[5]
Online advertisements with
Feizi, Soheil and Hajiaghayi, MohammadTaghi and Rezaei, Keivan and Shin, Suho , journal=. Online advertisements with
-
[6]
Proceedings of the ACM Web Conference 2024 , pages=
Mechanism design for large language models , author=. Proceedings of the ACM Web Conference 2024 , pages=
2024
-
[7]
Truthful aggregation of
Soumalias, Ermis and Curry, Michael and Seuken, Sven , booktitle=. Truthful aggregation of
-
[8]
Ad auctions for
Hajiaghayi, MohammadTaghi and Lahaie, S. Ad auctions for. Advances in Neural Information Processing Systems , volume=
Show all 39 references
-
[9]
Auctions with
Dubey, Avinava and Feng, Zhe and Kidambi, Rahul and Mehta, Aranyak and Wang, Di , booktitle=. Auctions with
-
[10]
Hu, Silan and Zhang, Shiqi and Shi, Yimin and Xiao, Xiaokui , journal=
-
[11]
Ad Insertion in
Xu, Shengwei and Chen, Zhaohua and Deng, Xiaotie and Huang, Zhiyi and Schoenebeck, Grant , journal=. Ad Insertion in
-
[12]
Zhao, Chujie and Hu, Qun and Song, Shiping and Chen, Dagui and Zhu, Han and Xu, Jian and Zheng, Bo , journal=
-
[13]
Ads that Talk Back: Implications and Perceptions of Injecting Personalized Advertising into
Tang, Brian Jay and Sun, Kaiwen and Curran, Noah T and Schaub, Florian and Shin, Kang G , journal=. Ads that Talk Back: Implications and Perceptions of Injecting Personalized Advertising into
-
[14]
Proceedings of the 1st Workshop on Deep Learning for Recommender Systems , pages=
Wide & deep learning for recommender systems , author=. Proceedings of the 1st Workshop on Deep Learning for Recommender Systems , pages=
-
[15]
Guo, Huifeng and Tang, Ruiming and Ye, Yunming and Li, Zhenguo and He, Xiuqiang , journal=
-
[16]
Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , pages=
Deep interest network for click-through rate prediction , author=. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , pages=
-
[17]
2010 IEEE International Conference on Data Mining , pages=
Factorization machines , author=. 2010 IEEE International Conference on Data Mining , pages=. 2010 , organization=
2010
-
[18]
Ma, Luyi and Zhang, Wanjia Sherry and Fan, Zezhong and Thakur, Shubham and Zhao, Kai and Yao, Kehui and Agarwal, Ayush and Iyer, Rahul and Cho, Jason and Xu, Jianpeng and others , journal=
-
[19]
Proceedings of the 13th ACM Conference on Electronic Commerce , pages=
Social influence in social advertising: Evidence from field experiments , author=. Proceedings of the 13th ACM Conference on Electronic Commerce , pages=
-
[20]
Journal of Advertising , volume=
Explaining the impact of scarcity appeals in advertising: The mediating role of perceptions of susceptibility , author=. Journal of Advertising , volume=
-
[21]
Political Analysis , volume=
Out of one, many: Using language models to simulate human samples , author=. Political Analysis , volume=
-
[22]
Psychology & Marketing , volume=
Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines , author=. Psychology & Marketing , volume=
-
[23]
Human Psychometric Questionnaires Mischaracterize
Song, Woojung and Choi, Dongmin and Park, Yoonah and Han, Jongwook and Jo, Yohan , journal=. Human Psychometric Questionnaires Mischaracterize
-
[24]
and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle=
Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle=
-
[25]
International Conference on Learning Representations , year=
Flow matching for generative modeling , author=. International Conference on Learning Representations , year=
-
[26]
arXiv preprint arXiv:1611.05916 , year=
Squared earth mover's distance-based loss for training deep neural networks , author=. arXiv preprint arXiv:1611.05916 , year=
-
[27]
Beyond accuracy: Behavioral testing of
Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , booktitle=. Beyond accuracy: Behavioral testing of
-
[28]
IEEE Transactions on Software Engineering , volume=
A survey on metamorphic testing , author=. IEEE Transactions on Software Engineering , volume=
-
[29]
ACM Computing Surveys , volume=
Metamorphic testing: A review of challenges and opportunities , author=. ACM Computing Surveys , volume=
-
[30]
Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=
The ``problem'' of human label variation: On ground truth in data, modeling and evaluation , author=. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=
2022
-
[31]
Transactions of the Association for Computational Linguistics , volume=
Dealing with disagreements: Looking beyond the majority vote in subjective annotations , author=. Transactions of the Association for Computational Linguistics , volume=
-
[32]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Toward a perspectivist turn in ground truthing for predictive computing , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[33]
IEEE Transactions on Neural Networks and Learning Systems , volume=
Learning from noisy labels with deep neural networks: A survey , author=. IEEE Transactions on Neural Networks and Learning Systems , volume=
-
[34]
Psychological Science , volume=
Personalized persuasion: Tailoring persuasive appeals to recipients' personality traits , author=. Psychological Science , volume=
-
[35]
Proceedings of the
Psychological targeting as an effective approach to digital mass persuasion , author=. Proceedings of the
-
[36]
How does personality affect perception of advertising messages?
Grochowska, Alicja and M. How does personality affect perception of advertising messages?. International Journal of Advertising , volume=
-
[37]
Perla, Raphael and Maran, Thomas and Bagci, Beg. The (. Psychology & Marketing , volume=
-
[38]
Mathematics of Operations Research , volume=
Optimal auction design , author=. Mathematics of Operations Research , volume=
-
[39]
Zhang, Zhaowei and Fu, Yuhan and Zhang, Yihang and Liu, Xiaohan and Zhang, Ceyao and Zhang, Xiaoyuan and Kang, Yipeng and Wang, Tonghan and Yang, Yaodong , journal=
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.