{"id":"6734138a-ec96-4f4a-b746-cd3444c076e7","arxiv_id":"2508.15486","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LongRetriever brings ultra-long user sequences into candidate retrieval and reports statistically significant, fully deployed gains in online A/B tests on a large e-commerce platform.","lead":"This paper introduces LongRetriever, a framework that uses a user's ultra-long behavior history during the candidate retrieval stage of a recommendation system, pairing in-context training with multi-context retrieval. It reports significant online A/B test gains and full deployment on a large e-commerce platform, which matters because most long-sequence work targets ranking rather than retrieval.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central efficacy claim rests on unverifiable A/B assertion; effect size, protocol, and latency feasibility are absent from available text.","rationale":"The reader correctly identified that the central claim rests on the abstract's undocumented A/B improvement and the feasibility of serving candidate-specific interactions within retrieval latency. The strongest load-bearing concern is the complete absence of experimental detail and the unreadable full text, which makes causal attribution impossible to verify. This is not a claim of fraud or error; it is a statement that the evidence provided is insufficient. Since the reader already marked the paper UNVERDICTED with low confidence, my analysis does not change the verdict—it reinforces it. The concrete test of obtaining the uncorrupted text and inspecting the A/B protocol would settle whether the concern lands by showing either a well-controlled experiment with effect sizes or a continuing lack of evidence. I agree with the reader's weakest assumption and find no reason to adjust the verdict.","tokens_in":17343,"tokens_out":2394,"duration_ms":26994,"concrete_test":"Obtain a readable copy of the full text (e.g., from arXiv source files or the authors) and locate the online A/B test write-up. Check whether the treatment arm differs from baseline only in the proposed modules—same candidate pool, same retrieval depth, same serving hardware/latency budget—and whether the paper reports a primary metric, relative/absolute effect size, and confidence interval. If the effect excludes zero with a clearly specified protocol and latency parity, the concern is resolved; if those numbers are missing, the central claim remains unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is only assessable via its abstract because the full text is corrupted and unreadable. The central claim—that LongRetriever produces 'statistically significant improvements' and is 'fully deployed'—is an assertion with no supporting experimental protocol. For this claim to be true, the observed A/B gain must be causally attributable to the proposed in-context training and multi-context retrieval, rather than to confounding changes (e.g., increased retrieval depth, larger candidate pool, or additional compute), and the candidate-specific interaction must be serveable within the retrieval stage's latency budget. Neither condition can be checked: the abstract reports no metrics, baselines, effect sizes, confidence intervals, traffic allocation, or latency measurements, and the corrupted body prevents inspection of the architecture or experiments. This is an evidence gap, not a demonstrated internal inconsistency, but it is load-bearing: without these details the central efficacy claim is unfalsifiable from the provided materials.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LongRetriever, a framework for incorporating ultra-long user sequences into the candidate retrieval stage of industrial recommendation systems. The two claimed technical ingredients are in-context training and multi-context retrieval, which are said to enable candidate-specific interaction between the user sequence and the candidate item while ensuring training-serving consistency under a search-based paradigm. The abstract asserts that extensive online A/B testing on a large-scale e-commerce platform produced statistically significant improvements, and that the system is fully deployed and impacts billions of users. The full text supplied for review is corrupted and almost entirely unreadable; only the abstract and a few garbled fragments could be assessed.","tokens_in":17576,"tokens_out":3622,"duration_ms":41798,"significance":"The motivating problem is timely and practically important: ultra-long user sequences are typically exploited only in the ranking stage, and moving them into retrieval could materially affect both efficiency and personalization at scale. If the proposed in-context training and multi-context retrieval mechanisms work as claimed, the paper would be a meaningful industrial contribution, especially because it explicitly targets the latency-sensitive retrieval stage. The deployment claim also suggests practical feasibility. However, the submission currently provides no verifiable technical content: no readable architecture, derivation, equations, algorithm, or experiment tables, and no quantitative detail about the A/B test. The paper does not ship machine-checked proofs, reproducible code, or parameter-free derivations. The significance is therefore conditional on a complete, readable manuscript with full experimental protocol.","major_comments":[{"comment":"The central efficacy claim is the sentence: 'Extensive online A/B testing conducted on a large-scale e-commerce platform demonstrates statistically significant improvements, confirming the framework's effectiveness.' This is the only evidence offered for the paper's main claim, yet it reports no metrics, baseline system, effect size, confidence interval, traffic allocation, experiment duration, or number of users. As the abstract is the only readable part of the submission, the claim is unfalsifiable from the supplied materials.","section":"Abstract"},{"comment":"The body of the manuscript is corrupted mojibake; no architecture description, loss functions, equations, algorithm pseudocode, or experimental tables are legible. Consequently the derivation of 'in-context training' and 'multi-context retrieval', and the claimed training-serving consistency under the search-based paradigm, cannot be checked. This is not a typographical issue: it blocks any technical evaluation of the proposed method.","section":"Full text (unreadable)"},{"comment":"A key practical claim is that candidate-specific interaction between the user sequence and candidate item can be served in the retrieval stage. No complexity analysis, latency measurements, percentile serving times, or infrastructure details are visible. Without such evidence, the production-deployment claim is unsupported; a reviewer cannot tell whether the method meets retrieval-stage latency budgets.","section":"Full text (serving feasibility)"},{"comment":"The abstract does not describe ablations or control conditions separating the contribution of in-context training and multi-context retrieval from possible confounding deployment changes, such as increased retrieval depth, a larger candidate pool, or additional compute. Even if full A/B results were reported, the observed improvement could not be attributed to the proposed mechanisms without such controls.","section":"Abstract / experimental methodology"}],"minor_comments":[{"comment":"The garbled full text contains 'arXiv:2508.15487v1 [cs.CL] 21 Aug 2025', while the submitted paper is labeled arXiv:2508.15486 (cs.IR). This mismatch should be resolved in a clean resubmission.","section":"Header/metadata"},{"comment":"The abstract states that current approaches focus on ranking stage and retrieval is under-explored, but gives no citations; a revised version should cite representative prior work on ultra-long-sequence recommender systems.","section":"Abstract"},{"comment":"The readable fragments contain repeated duplicated blocks and references to a different paper; the PDF clearly failed to compile. Authors must regenerate and verify the document before resubmission.","section":"Full text"}],"recommendation":"uncertain","confidential_remarks":"The supplied PDF is corrupted and appears to include text from another arXiv submission (2508.15487, cs.CL). I could not perform a technical review of the proposed method or experiments. I recommend that the editor ask the authors to submit a clean, compilable manuscript with full experimental details before any substantive refereeing. I found no evidence of misconduct, but the current version is not reviewable in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the abstract describes a plausible way to push ultra-long user sequences into the retrieval stage, but this submission cannot be evaluated: the PDF is corrupted, and the only readable text is an abstract promising statistically significant improvements with no numbers. The central efficacy claim is unfalsifiable from the provided materials.\n\nWhat is actually new: most ultra-long sequence work targets the ranking stage; moving it to retrieval, with candidate-specific interaction via in-context training and multi-context retrieval, is a sensible and under-explored direction. If the deployment claim is true, it is meaningful evidence of practical value.\n\nWhere the soft spots are: first, the full text is unreadable, so no method detail, architecture, or ablation can be checked. Second, the abstract reports no effect size, metric, baseline, variance, traffic allocation, or latency number. The reported A/B gain could in principle come from confounding deployment changes rather than the proposed mechanism, and the retrieval-stage latency budget is not addressed. These are evidence gaps, not demonstrated errors; there is no obvious internal contradiction in the abstract. The 'in-context training' idea could in principle be tuned to the very retrieval targets it later predicts, but that is a suspicion, not a finding.\n\nWho this is for: people working on industrial retrieval systems would likely get value from a clean version of this paper. As submitted, no one can read it, so no reader gets value from it.\n\nRecommendation: desk reject the corrupted file and invite the authors to resubmit a readable manuscript with real experimental detail. I would not spend referee time on this artifact. If a clean version appears with actual A/B numbers and a latency analysis, it deserves a serious referee.","headline":"A plausible industrial retrieval-stage framework whose only evidence is an unverifiable A/B claim; the corrupted full text makes this unpublishable in current form.","tokens_in":17991,"tokens_out":2179,"would_cite":false,"duration_ms":23384,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LongRetriever brings ultra-long user sequences into the candidate retrieval stage and reports statistically significant live gains on a large e-commerce platform.","keywords":["LongRetriever","candidate retrieval","ultra-long user sequences","in-context training","multi-context retrieval","recommender systems","online A/B testing","industrial deployment"],"falsifier":"Run an offline controlled comparison on a fixed corpus where LongRetriever's in-context training is ablated to a standard two-tower retriever with identical sequence length and serving budget; if candidate-specific interaction does not improve recall@k, or if the A/B conversion lift vanishes when the mechanism is disabled, the central causal claim fails.","tokens_in":17298,"feed_emoji":"🛒","tokens_out":2331,"duration_ms":24959,"temperature":0.7,"pith_summary":"This paper argues that candidate retrieval, not just ranking, can and should consume a user's ultra-long behavior sequence. It introduces LongRetriever, built around two ideas: in-context training, where the retriever learns to condition on the candidate item while reading the sequence, and multi-context retrieval, which keeps training and serving aligned in a search-based setup. If the reported online A/B results are what they appear to be, retrieval quality improves enough to be fully deployed on a large e-commerce platform serving billions of users.","feed_headline":"Ultra-long sequences now drive candidate retrieval, not just ranking","feed_subtitle":"LongRetriever adds candidate-specific interaction at the retrieval stage and reports live e-commerce gains.","key_machinery":"In-context training: the retriever is trained with the candidate item's representation in the same context as the user sequence, so the model learns interactions specific to each candidate. Multi-context retrieval: at serving time, the retrieval stage evaluates candidates under contextual representations built the same way as training, preserving training-serving consistency. Together they let a search-based retriever score candidates against ultra-long sequences rather than against a fixed user summary.","core_discovery":"LongRetriever's central claim is that ultra-long user sequences can be brought into the candidate retrieval stage without sacrificing latency, by making the sequence interact with the candidate item being scored instead of compressing the user into a fixed vector. The paper names the enabling mechanisms in-context training and multi-context retrieval: the first teaches the model to predict a candidate's relevance from the sequence with candidate information present at training time; the second ensures that at serving time, retrieval runs as a search over candidate-conditioned contexts rather than a single pass over a user embedding. The proof offered is the live deployment: statistically sig","pith_inferences":["If the deployment claim holds, candidate-conditioned retrieval could transfer to other search-based problems with large candidate spaces and long user histories, such as web search, video recommendation, or advertising.","In-context training implies the retriever learns a conditional representation per candidate rather than a fixed user vector, which may reduce the burden on the ranking stage to capture sequence-item interactions.","A testable extension would vary the sequence length and measure marginal recall gains to find the point where ultra-long history stops adding signal.","The approach's success could be checked by ablating in-context training and measuring whether the retrieval lift disappears, isolating the mechanism from other deployment changes."],"forward_implications":["Retrieval can use the full user sequence instead of a compressed summary, so candidates are scored against what the user actually did.","Training and serving use the same search-based procedure, so the model behaves at serving time as it did during training.","Candidate-specific interaction is feasible within retrieval latency, as demonstrated by the reported full deployment.","Billions of users see retrieval results conditioned on their ultra-long history, if the deployment claim is accurate."],"supporting_citations":[],"fun_headline_variants":["Ultra-long sequences now drive candidate retrieval","Candidate-specific interaction enables ultra-long retrieval","Live deployment shows gains in ultra-long sequence retrieval","Retrieval stage leverages ultra-long user sequences","LongRetriever: multi-context retrieval for ultra-long sequences"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The paper's central claim depends on the assumption that the measured online A/B improvement comes from in-context training and multi-context retrieval themselves, rather than from other deployment changes, and that the added candidate-specific computation fits the retrieval stage's latency budget.","fun_headline_variants_meta":{"raw":{"variants":["Ultra-long sequences now drive candidate retrieval","Candidate-specific interaction enables ultra-long retrieval","Live deployment shows gains in ultra-long sequence retrieval","Retrieval stage leverages ultra-long user sequences","LongRetriever: multi-context retrieval for ultra-long sequences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000446,"raw_usage":{"total_tokens":2031,"prompt_tokens":627,"completion_tokens":1404,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":371,"completion_tokens_details":{"reasoning_tokens":1346}},"tokens_in":371,"tokens_out":1404,"duration_ms":14579,"temperature":1.0,"reasoning_tokens":1346,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:51:19.989188+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an offline controlled comparison on a fixed corpus where LongRetriever's in-context training is ablated to a standard two-tower retriever with identical sequence length and serving budget; if candidate-specific interaction does not improve recall@k, or if the A/B conversion lift vanishes when the mechanism is disabled, the central causal claim fails.","supporting_citations":[],"review_version":1}