{"id":"712e3385-cbb0-476c-ab1a-f5290e396881","arxiv_id":"2412.03577","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"OKG uses an LLM agent with web search, KPI memory, and adaptive wider/deeper expansion to generate sponsored-search keywords dynamically, and claims better KPI estimates than static baselines.","lead":"This paper presents OKG, an LLM agent that monitors live keyword performance and generates new search-ad keywords on the fly, plus a public dataset of real Japanese keyword KPIs. A generalist might read it to see whether real-time LLM agents can replace static keyword generation in paid search advertising.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All headline performance claims rest on KPI imputation from offline nearest neighbors, not on live delivery; if that surrogate is unvalidated, OKG's central 'real-time improvement' claim is unsupported.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: Section 5.1's nearest-neighbor KPI imputation is a surrogate for real delivery performance. I agree that this is the single most critical weakness. The entire quantitative case for OKG—including the adaptive feedback loop in Section 4.2, since the agent's decisions are conditioned on these imputed Pt values—collapses if the surrogate does not faithfully reflect actual keyword KPIs. The paper provides no validation of the surrogate, no error bars, and no live deployment evidence, so the central claim is not established. The dataset and system description may still be useful resources, but they do not rescue the empirical claim. My assessment therefore leaves the reader's REJECT verdict unchanged.","tokens_in":12602,"tokens_out":3469,"duration_ms":38479,"concrete_test":"Validate the Section 5.1 surrogate on held-out dataset keywords: for each keyword, hide its KPI, retrieve its nearest neighbor among remaining keywords using the same multilingual BERT cosine similarity with threshold 0.6, and compare predicted versus actual clicks, search volume, CPC, and competitor score. Report Pearson correlation and mean absolute error; if the correlation is below roughly 0.7 or the imputed values are not within a small multiple of true values, Table 1's comparisons are unreliable. As a stronger follow-up, run a small live Google Ads A/B test over T=3 steps for one product, comparing OKG-generated keywords against GPT-4 baselines with platform-reported clicks and CPC; if actual KPIs do not reproduce the same ranking, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that OKG significantly improves keyword performance and adaptability via real-time KPI monitoring—depends entirely on Table 1, whose KPIs are not measured from any live campaign. Section 5.1 states that for each generated keyword, the authors select the most similar keyword from offline data (highest BERT cosine similarity > 0.6) and use that neighbor's clicks, search volume, CPC, and competitor score as the generated keyword's KPIs. Thus every KPI used both to drive OKG's adaptive allocation (Section 4.2's Pt−1) and to compare against baselines is an imputed value from static offline data. If embedding similarity does not accurately predict real delivery performance, the reported improvements are artifacts of the imputation rule. The comparison is also non-independent: OKG retrieves from and is evaluated against the same offline vocabulary, while baselines are not, and no error bars, significance tests, or raw counts are reported. The claimed real-time adaptivity is never tested against actual ad-platform feedback; the public dataset may be a useful contribution, but it does not validate the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OKG, an LLM-agent framework for sponsored-search keyword generation that uses an online search tool, a vector memory module with retrieval-augmented generation, and a calculation tool to allocate new keywords between wider (new categories) and deeper (existing categories) expansion based on previously observed KPI data. The authors also release a public dataset of real Japanese keyword KPI records for ten Sony products. The experimental section compares OKG with GPT-4, Gemini-1.5-Pro, keyword extraction baselines, and Google Keyword Planner on imputed KPIs, relevance/coverage scores, and similarity to offline keywords, and reports that OKG outperforms all baselines. The paper concludes that OKG significantly improves keyword adaptability and responsiveness.","tokens_in":12870,"tokens_out":7447,"duration_ms":70840,"significance":"The proposed setting is timely: keyword generation in sponsored search usually relies on static offline data, and a system that continuously adapts keyword lists from performance feedback would be practically valuable. The public dataset and released code are a concrete contribution that could support future work. However, the empirical evidence for the headline claim is currently not adequate: the KPI evaluation is based on nearest-neighbor imputation from offline data rather than live campaign delivery, and the adaptive feedback loop is exercised only on those imputed values. The usefulness of the dataset and the overall architecture are independent of that evidence, but the claimed performance advantage is not established by the experiments as reported.","major_comments":[{"comment":"The central performance claim rests on KPI imputation, not on observed delivery. The paper states: 'For each generated keyword, we select the most similar keyword from offline data (highest similarity score and cosine similarity > 0.6) and use its KPIs to represent the generated keyword's KPIs.' Consequently, every Click, Search Volume, CPC, and Competitor Score in Table 1 is a copy of a historical keyword's KPI selected by embedding similarity. No evidence is given that BERT cosine similarity above 0.6 is predictive of actual advertising KPIs for a newly generated keyword. Because §4.2 uses Pt−1 (these imputed values) to compute pW_t and pD_t, the entire 'real-time adaptation' loop is simulated with synthetic feedback rather than measured against platform behavior. The headline claim that OKG 'significantly improves keyword performance' is therefore unsupported.","section":"§5.1 (Table 1)"},{"comment":"The evaluation is not an independent comparison. Generated keywords are scored against the same offline vocabulary that OKG's search and memory components are designed to exploit, and the KPI assignment rule gives higher scores to keywords that are embedding-neighbors of high-performing historical keywords. OKG is explicitly optimized to produce keywords similar to the offline data, so the comparison rewards matching the KPI database rather than delivering better campaign outcomes. A valid evaluation would require a held-out set of keywords with known KPIs, a validated surrogate, or a live A/B test with actual ad delivery; none of these is reported.","section":"§5.1–§5.3"},{"comment":"No statistical support is provided for 'significantly improves.' The experiments use T=3 with a single run, and Table 1 reports only normalized point values without error bars, confidence intervals, or significance tests. The ablation in Fig. 3 is performed on a single product (Sony TV) with no repeated trials, so it cannot support the claim that full OKG is reliably better than the ablations. The paper should report multiple runs, per-product results, and appropriate uncertainty or significance measures.","section":"§5, Implementation Details; Fig. 3"},{"comment":"The problem setting defines Pt as 'observed KPI' from the ad platform, and the objective is to maximize the sum of Pt over T. However, in the experiments Pt is never observed from a platform; it is taken from static offline data via nearest-neighbor retrieval. This mismatch between the formal setting and the evaluation means that the paper does not demonstrate the core claim of on-the-fly adaptation to real-time KPI changes. The authors either need to run the system against the Google Ads API (which they cite as a component) for a real campaign, or explicitly reframe the experiments as a simulation and validate the surrogate.","section":"§3 and §4.2"}],"minor_comments":[{"comment":"The reference list contains two entries with the same key 'Google. 2024' (Google Ads API and Google Keyword Planner), and the in-text citations use both 'Google, 2024a' and 'Google, 2024b'; please disambiguate the keys consistently.","section":"References"},{"comment":"The notation for the keyword set is inconsistent: the text alternates between lowercase 'kt' and uppercase 'Kt' for the same concept, and the definition 'K = union over t of Kt' makes the cumulative set and the per-step set easy to confuse in the objective function.","section":"Section 3"},{"comment":"The table header contains formatting typos ('Srch. V ol.', 'N.(0∼100)', 'Gemini1.5' missing a space, and inconsistent spacing in 'Kwd. Ext.'). Please clean up the table formatting.","section":"Table 1"},{"comment":"Because the examples are translated from Japanese, the original Japanese keyword outputs should be included alongside the English translations so that the claimed specificity and brand alignment can be independently checked.","section":"Appendix C"},{"comment":"The sentence 'Generated and original keywords are tokenized and embedded (using pooled embeddings) from a pretrained multilingual BERT model' should specify which pooling method and checkpoint were used, to support reproducibility.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful dataset and a plausible architecture, but the evaluation as it stands would not support publication in a strong venue. If the authors can validate the KPI surrogate or provide live-campaign results, I would be willing to reconsider. The paper might also be reframed as a dataset and system-description contribution with clearly labeled simulation results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my take on OKG. The headline claim—that the agent significantly improves keyword performance via real-time KPI feedback—is not supported by the experiments as run. The evaluation imputes every KPI from an offline nearest neighbor using BERT similarity (Section 5.1), so the comparison is a similarity-weighted copy of historical performance, not a measurement of live delivery. That is a load-bearing flaw: the system's own allocation rule (Pt−1) and the final Table 1 both run on this surrogate. With no error bars, significance tests, or multiple runs (T=3, one ablation product), the paper's central 'real-time improvement' claim is unproven.\n\nThat said, there is real substance here. The agentic loop—web search, KPI memory, adaptive wider/deeper expansion—is a sensible new architecture for a genuine problem. And the public dataset of real Japanese ad keywords with KPIs is the first open resource in this space; that alone is a useful contribution for follow-up work. The writing is clear, and the ablation, while limited, at least compares against fixed and one-direction baselines.\n\nThe soft spots beyond the imputation: the baselines are not evaluated under the same retrieval conditions; the similarity threshold (>0.6) and temperature are free parameters; and the 'online relevance' metrics (BLEU/ROUGE) are low across the board, which the paper acknowledges. None of these rescue the core issue.\n\nWho is this for? Researchers working on SSA keyword generation who want the dataset or a starting architecture. The performance claims should not be taken at face value.\n\nMy recommendation: this deserves a serious referee, but the evaluation needs to be rebuilt—at minimum a held-out validation of the imputation, error bars over multiple runs, and ideally a small live campaign test. I would probably reject in current form, but the dataset and idea merit revision.","headline":"Useful dataset and a sensible agent architecture, but the headline performance claim is unproven because all KPIs are imputed from offline nearest neighbors, not live delivery.","tokens_in":13334,"tokens_out":1764,"would_cite":true,"duration_ms":16978,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OKG claims that an LLM agent that monitors live sponsored-search KPIs and regenerates keywords each cycle outperforms static LLM, extraction, and keyword-planner baselines.","keywords":["sponsored search advertising","keyword generation","LLM agent","real-time KPI monitoring","adaptive keyword expansion","public dataset","retrieval-augmented generation"],"falsifier":"Run the exact keyword lists produced by OKG and by GPT-4 as separate live Google Ads campaigns for the same products over the same three time steps, then compare actual clicks, cost per click, and competitor scores; if OKG's advantage over GPT-4 disappears or reverses under real delivery data, the offline nearest-neighbor KPI surrogate, rather than the generation method, produced the reported gains.","tokens_in":12452,"feed_emoji":"🎯","tokens_out":10733,"duration_ms":88350,"temperature":0.7,"pith_summary":"This paper argues that sponsored-search keyword generation does not have to depend on large static datasets, which most advertisers cannot obtain. The proposed system, OKG, is an LLM agent that repeatedly searches for current product and market information, reads the live performance of its own generated keywords (clicks, cost per click, competitor score), and then produces the next keyword list by deepening winning categories and widening into new ones while keeping the total number of keywords fixed. The authors claim this on-the-fly loop yields higher clicks, lower cost per click, and lower competitor scores than GPT-4, Gemini-1.5-Pro, keyword extractors, and Google Keyword Planner in their evaluation, and that ablations show the adaptive split between deeper and wider expansion is what drives the improvement. The paper also releases a public dataset of real Japanese keyword records with KPIs across several product domains, which it says is the first such resource for the field.","feed_headline":"Live ad KPIs steer an LLM agent to stronger keywords","feed_subtitle":"OKG regenerates sponsored-search keywords each cycle from real-time clicks and costs, no big keyword dataset needed.","key_machinery":"The load-bearing mechanism is the OKG agent loop, a closed cycle that turns live performance feedback into the next keyword list. At time step $t$, the agent observes the current keyword set $k_t$ and its measured KPIs $P_t$, pulls real-time product and market information $S_t$ through a search tool, and retrieves historical keyword performance from a vector memory via RAG; a calculation tool then splits the fixed budget of $n$ keywords between wider expansion $W_t$ (new keyword categories) and deeper expansion $D_t$ (more specific keywords in proven categories). The next list is $k_{t+1}=W_t \\cup D_t$, with $|W_t| = \\lfloor p^W_t \\cdot n \\rfloor$ and $|D_t| = n - |W_t|$, where $p^W_t$ and $p^D_t$ are the previous step's KPI shares of the wider and deeper keywords. This feedback loop is the entirety of the argument: live KPI changes redirect the agent's keyword choices, which static offline-trained generators cannot do.","core_discovery":"The central discovery claimed in the paper is that keyword generation in sponsored search can be posed as an online decision problem—maximize total KPI over a time horizon T with a fixed per-step keyword budget n—and solved by an LLM agent that closes the loop between generation and live measurement. At each time step the agent observes the current keyword set and its KPIs, retrieves real-time product and market information, and computes the next keyword set as a union of wider expansions (new categories) and deeper expansions (more specific keywords in proven categories), with the split proportional to the previous step's KPI shares. The authors report that this approach outperforms LLM-based, keyword-extraction, and commercial-planner baselines on clicks, cost per click, and competitor score, and also produces higher relevance/coverage scores against online search results and higher similarity to offline real keywords. An ablation study is used to argue that neither wider-only nor deeper-only growth, nor a fixed split, nor a reflection-based variant matches the full adaptive mechanism, supporting the claim that real-time KPI feedback is the decisive component. The paper additionally contributes what it calls the first publicly accessible dataset of real keyword data with KPIs across diverse domains.","pith_inferences":["A test the paper leaves implicit is whether the offline nearest-neighbor KPI surrogate tracks real delivery: one could deploy the generated keyword lists in a live campaign and compare actual clicks and costs to the surrogate values, which would separate generation quality from evaluation artifact.","The wider/deeper proportional allocation rule is a generic exploration-exploitation schedule that could transfer to other fixed-budget list-generation settings, such as query suggestions, trending hashtags, or product recommendations, wherever per-step performance is observable.","If real-time feedback indeed matters more than reflection on past experience, a natural next experiment is to vary the feedback delay (daily vs weekly KPI updates) to map how quickly the agent loses its advantage as signals age.","The paper's evaluation covers only Japanese keywords and Sony products; whether OKG's gains persist across languages, product categories, and advertising platforms (e.g., Bing Ads or Yahoo) is an open question that the released dataset and code make testable."],"forward_implications":["If OKG is correct, an advertiser can start a sponsored-search campaign from just a product name and let the agent maintain and refresh the keyword list from live performance data, eliminating the need for a large pre-collected keyword dataset.","Keyword lists would become time-varying assets: under a fixed per-step budget, the agent shifts slots toward categories that recently earned clicks, so budget is reallocated away from underperforming keywords without manual intervention.","The released public dataset with real keyword KPIs would give future researchers a shared benchmark, making keyword-generation methods comparable on identical real-world performance data.","The ablation result that reflection on past steps does not improve quality would suggest that, for keyword selection, current live signals are more valuable than accumulated historical experience."],"supporting_citations":[{"why":"Supplies GPT-4, the LLM backbone used for OKG's planning and generation.","marker":"Achiam et al., 2023"},{"why":"Supplies Gemini-1.5-Pro, one of the two main LLM baselines OKG is compared against.","marker":"Reid et al., 2024"},{"why":"Provides LKG, the recent LLM-based keyword generation method that is the closest baseline in the study.","marker":"Wang et al., 2024"},{"why":"Supplies the GAN-based keyword expansion baseline line of work that OKG contrasts with.","marker":"Lee et al., 2018"},{"why":"Supplies the retrieval-augmented generation mechanism OKG uses to query its historical keyword-KPI memory.","marker":"Lewis et al., 2020"},{"why":"Provides the Google Ads API used for live KPI retrieval and the Google Keyword Planner baseline.","marker":"Google, 2024"},{"why":"Supplies the search tool OKG uses to gather real-time product, pricing, and market information S_t.","marker":"Serp, 2024"},{"why":"Supplies BLEU-2, one of the metrics used to measure how well generated keywords cover online search results.","marker":"Papineni et al., 2002"},{"why":"Supplies ROUGE-1, the recall-oriented metric used for keyword coverage against search results.","marker":"Lin, 2004"},{"why":"Supplies BERTScore, the semantic similarity metric used for both relevance and alignment with offline real keywords.","marker":"Zhang et al., 2019"}],"fun_headline_variants":["LLM agent regenerates ad keywords from live KPIs","On-the-fly ad keywords adapt to real-time KPIs","Real-time KPIs drive LLM keyword generation","Live KPI feedback steers LLM keyword regeneration","No static dataset: LLM ad keywords generated live"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported KPI comparisons assume that a generated keyword's real clicks, cost per click, and competitor score are accurately represented by the values of the most textually similar keyword in the offline dataset; if embedding similarity does not predict actual campaign delivery, the headline performance gains are measuring resemblance to old keywords rather than real advertising success.","fun_headline_variants_meta":{"raw":{"variants":["LLM agent regenerates ad keywords from live KPIs","On-the-fly ad keywords adapt to real-time KPIs","Real-time KPIs drive LLM keyword generation","Live KPI feedback steers LLM keyword regeneration","No static dataset: LLM ad keywords generated live"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001568,"raw_usage":{"total_tokens":6244,"prompt_tokens":913,"completion_tokens":5331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":5267}},"tokens_in":529,"tokens_out":5331,"duration_ms":32701,"temperature":1.0,"reasoning_tokens":5267,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:44:03.508189+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the exact keyword lists produced by OKG and by GPT-4 as separate live Google Ads campaigns for the same products over the same three time steps, then compare actual clicks, cost per click, and competitor scores; if OKG's advantage over GPT-4 disappears or reverses under real delivery data, the offline nearest-neighbor KPI surrogate, rather than the generation method, produced the reported gains.","supporting_citations":[{"cited_title":"https://developers","cited_arxiv_id":null,"evidence_quote":"Provides the Google Ads API used for live KPI retrieval and the Google Keyword Planner baseline."},{"cited_title":"https://serpapi.com/","cited_arxiv_id":null,"evidence_quote":"Supplies the search tool OKG uses to gather real-time product, pricing, and market information S_t."}],"review_version":1}