{"id":"a755a91d-d84c-48ef-ab53-c45c8d01bf63","arxiv_id":"2605.21812","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces a seed-guided contrastive framework that uses LLMs to generate realistic synthetic queries and topicality labels for cold-start natural language search, outperforming no-seed and InPars baselines on realism metrics and producing harder evaluation sets.","lead":"This paper describes an LLM-based system that creates synthetic user queries and relevance labels to train and test natural language search when no real user data exists yet. A smart generalist might read it to see how a large platform like Airbnb can launch AI search features without waiting for months of user activity.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No direct evidence that distribution-matched synthetic data improves retrieval/ranking on real user queries","rationale":"The reader's weakest assumption correctly isolates the missing link between proxy metrics and live transfer. Because the paper reports deployment but no quantitative real-query results, the central claim remains unverified rather than refuted; the proposed check directly tests whether the observed KL and hardness gains produce measurable gains on actual user data.","tokens_in":1763,"tokens_out":321,"duration_ms":19794,"concrete_test":"After real queries accumulate, train two embedding models—one on the seed-guided synthetic set, one on the InPars baseline set—then evaluate both on the same held-out real query set using recall@10 or MRR; if the seed-guided model does not outperform the InPars model by a statistically significant margin, the distributional improvements do not transfer to task performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline results (KL 0.66 on length, 0.04 on attributes, 79% vs 97% pairwise accuracy) are computed on generated queries against a real-user reference distribution. For the cold-start claim to hold, these surface statistics must imply that models trained or evaluated on the synthetic set will generalize to live queries. The manuscript provides no downstream experiment that trains an embedding retriever or ranker on the seed-guided synthetic data and measures recall, NDCG, or click metrics on a held-out set of real post-deployment queries, nor any A/B test comparing synthetic-trained vs. baseline models in production.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes an LLM-powered framework to generate synthetic queries and relevance labels for natural language search systems facing cold-start conditions, as at Airbnb. Query generation combines contrastive listing pairs from booking sessions with user-research seed queries to improve realism and diversity. Label generation uses contrastive methods to produce topicality labels by construction and Virtual Judge (VJ) labeling for broader coverage. The authors compare their seed-guided approach against a no-seed contrastive baseline and an InPars-style baseline, reporting a KL divergence of 0.66 on query length (7.5x better than InPars' 12.03), 0.04 on attribute distributions, and harder evaluation examples (79% pairwise accuracy vs. 97% for the no-seed baseline). They also describe deploying production pipelines that generate synthetic examples daily for embedding-based retrieval and ranking evaluation.","tokens_in":1900,"tokens_out":556,"duration_ms":51697,"significance":"If the synthetic data generalizes to live user queries, the work offers a practical, deployable solution for cold-start challenges in industry IR systems, with clear quantitative gains in distribution matching over strong baselines. The production deployment and focus on both query generation and labeling are practical strengths that could influence synthetic-data practices in search.","major_comments":[{"comment":"The central claim that the seed-guided synthetic data bridges the cold-start gap for retrieval and ranking rests on the assumption that distribution-matched synthetic queries and labels will improve models on real user queries. However, the manuscript reports no downstream experiment that trains an embedding retriever or ranker on the generated data and measures recall, NDCG, or click-through metrics on a held-out set of real post-deployment queries (or any A/B test in production). The reported KL divergences and pairwise accuracies are computed only against real-user reference distributions on the synthetic set itself.","section":"Experiments / Evaluation"}],"minor_comments":[{"comment":"The abstract introduces 'Virtual Judge (VJ) labeling' without a one-sentence definition; a brief parenthetical explanation on first use would aid readers.","section":"Abstract"},{"comment":"Clarify the exact difference in prompt construction or sampling between the 'seed-guided' method and the 'no-seed contrastive baseline' so that the 79% vs. 97% pairwise accuracy gap can be reproduced from the description alone.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":"The skeptic note correctly identifies the missing real-query downstream validation as the load-bearing gap; addressing it would move the paper from promising empirical matching to a stronger contribution for an IR venue."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment of the practical strengths of our work and for the constructive feedback on evaluation. We address the major comment below, clarifying the scope of our contributions while committing to revisions that strengthen the manuscript.","responses":[{"response":"We appreciate this observation and agree that direct downstream experiments—training an embedding retriever or ranker on the synthetic data and measuring recall, NDCG, or click-through rate on held-out real post-deployment queries—would provide stronger evidence for the framework's impact. Our current evaluation deliberately focuses on the fidelity of the generated data through distribution-matching metrics (KL divergence on query length and attribute distributions) and the discriminative power of the resulting evaluation sets (pairwise accuracy of 79% vs. 97% for the no-seed baseline). These proxies are standard in synthetic-data-for-IR literature and demonstrate that the seed-guided approach produces more realistic and harder examples than the InPars-style and no-seed baselines. The manuscript also reports the production deployment of daily synthetic-data pipelines for embedding-based retrieval and ranking evaluation, which supports iterative model improvement as real queries accumulate during the cold-to-warm transition. We will revise the manuscript to (1) explicitly state the scope of our claims, (2) add a dedicated limitations paragraph acknowledging the absence of end-to-end retrieval/ranking results on real queries, and (3) outline concrete plans for future A/B testing and offline evaluation once sufficient post-deployment real query data is available. This partial revision will address the referee's concern without changing the core technical contributions.","revision_made":"partial","referee_comment":"[Experiments / Evaluation] The central claim that the seed-guided synthetic data bridges the cold-start gap for retrieval and ranking rests on the assumption that distribution-matched synthetic queries and labels will improve models on real user queries. However, the manuscript reports no downstream experiment that trains an embedding retriever or ranker on the generated data and measures recall, NDCG, or click-through metrics on a held-out set of real post-deployment queries (or any A/B test in production). The reported KL divergences and pairwise accuracies are computed only against real-user reference distributions on the synthetic set itself."}],"tokens_in":1453,"tokens_out":465,"duration_ms":36954,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a concrete engineering pipeline for cold-start natural language search. They combine contrastive listing pairs from booking sessions with user-research seed queries to guide LLM generation, then add contrastive labels and a Virtual Judge step for broader coverage. This produces queries whose length and attribute distributions sit much closer to real users than the baselines they compare against.","headline":"The paper gives a practical LLM recipe for synthetic queries and labels that matches real Airbnb user distributions far better than InPars, but provides no direct test that this improves retrieval or ranking on actual live queries.","tokens_in":2423,"tokens_out":155,"would_cite":false,"duration_ms":24019,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"We combine contrastive listing pairs from booking sessions with seed queries from user research... KL divergence of 0.66 on query length... 0.04 on attribute distributions"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/ArithmeticFromLogic.lean","rs_theorem":"LogicNat","paper_passage":"contrastive generation that produces topicality labels by construction"}],"headline":"LLM synthetic query/label generation for Airbnb IR cold-start has no structural overlap with RS cost functions, phi-ladders or distinction-forcing","alignment":"orthogonal","rationale":"The paper's machinery (contrastive listing-pair sampling, seed-guided prompt variants, KL divergence on length/attribute distributions, Virtual Judge labeling, pairwise accuracy) is standard applied IR/ML work on data augmentation. It never invokes reciprocal costs, golden-ratio identities, 8-tick periodicity, or any parameter-free derivation. No passage parallels any RS theorem; the work is therefore orthogonal to the entire RS framework.","tokens_in":51091,"confidence":"high","tokens_out":290,"duration_ms":10686,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Seed-guided LLM synthesis matches real user query lengths and attributes closely enough to train and evaluate natural language search models from the first day.","keywords":["synthetic data generation","LLM query generation","cold-start problem","natural language search","relevance labeling","contrastive generation","Airbnb search"],"falsifier":"A side-by-side comparison of length and attribute distributions between the synthetic queries and the first large batch of real user queries collected after deployment, showing substantially higher KL divergence than the reported 0.66 and 0.04 values.","tokens_in":2678,"feed_emoji":"🔍","tokens_out":666,"duration_ms":31637,"temperature":0.7,"pith_summary":"The paper presents a framework that uses large language models to create synthetic queries and relevance labels when no real user data exists for a new natural language search system. It guides query generation by mixing contrastive pairs of listings drawn from actual booking sessions with seed queries collected from user research. This produces queries whose length distribution has a KL divergence of 0.66 from real queries and whose attribute-type distribution has a KL divergence of 0.04. Label generation relies on contrastive construction that defines topicality by design together with virtual-judge labeling for broader coverage. The resulting evaluation sets are harder than those from a no-seed baseline, showing 79 percent pairwise accuracy instead of 97 percent, and the method is already running in daily production pipelines.","feed_headline":"Seed-guided LLMs match real query lengths 7.5x better than baselines","feed_subtitle":"Synthetic queries and labels from booking sessions and research seeds enable immediate training and harder evaluation for natural language检索","key_machinery":"Seed-guided contrastive query generation that blends booking-session pairs with user-research seeds, paired with contrastive label generation and virtual-judge labeling.","core_discovery":"A seed-guided contrastive approach that feeds booking-session pairs and user-research queries into LLMs generates synthetic queries and labels whose length and attribute distributions are far closer to real users than either a no-seed contrastive baseline or an InPars-style baseline, while also creating more discriminative evaluation examples for ranking models.","pith_inferences":["The same seeding strategy could be tested on other search or recommendation platforms that face similar initial-data shortages.","Reducing the influence of the seeds over time as real data accumulates might further tighten the match to live user behavior.","The method could shorten the time between feature launch and usable model performance in additional domains that rely on natural language input."],"forward_implications":["Production pipelines can generate fresh synthetic examples every day to support embedding retrieval and ranking evaluation.","The same seeded data enables a gradual shift from cold-start training to warm-start training as real queries arrive.","Harder synthetic evaluation sets give clearer signals for iterative model improvement than easier no-seed sets."],"fun_headline_variants":["Seed-guided LLMs match real query lengths with 0.66 KL divergence","Contrastive generation produces labels by construction for topicality","Virtual Judge labeling offers broader coverage for search evaluation","Seed queries and booking pairs balance realism in LLM data generation"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That queries and labels produced by the LLM, even when guided by booking sessions and research seeds, will match the behavior of real users once the live system receives actual queries.","fun_headline_variants_meta":{"raw":{"variants":["Seed-guided LLMs match real query lengths with 0.66 KL divergence","Contrastive generation produces labels by construction for topicality","Virtual Judge labeling offers broader coverage for search evaluation","Seed queries and booking pairs balance realism in LLM data generation"]},"model":"grok-4.3","cost_usd":0.01474,"raw_usage":{"total_tokens":6267,"prompt_tokens":688,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":147403000,"prompt_tokens_details":{"text_tokens":688,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":5513,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":688,"tokens_out":66,"duration_ms":63298,"temperature":1.0,"reasoning_tokens":5513,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-22T07:59:23.375589+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A side-by-side comparison of length and attribute distributions between the synthetic queries and the first large batch of real user queries collected after deployment, showing substantially higher KL divergence than the reported 0.66 and 0.04 values.","supporting_citations":[],"review_version":1}