{"id":"79340d32-c7e2-48b8-8651-9a0967d03add","arxiv_id":"2603.14864","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A new long-horizon preference-grounded shopping benchmark shows SOTA LLMs below 70% success, while a 4B model fine-tuned with tool-wise process rewards beats stronger baselines.","lead":"The paper introduces Shopping Companion Bench, a long-horizon e-commerce agent benchmark with two preference-memory tasks over 1.2M real products, and trains agents with annotation-free tool-wise rewards. It matters because shopping agents fail when they forget or invent user preferences across sessions, and this work targets that gap with a hard public-style evaluation and denser training signal.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified on the Shopping Companion claim; the provided full text is a different paper (EnSF data assimilation), so the abstract claim cannot be stress-tested against methods or results.","rationale":"The reader correctly treated this as abstract-only for Shopping Companion and set UNVERDICTED / LOW confidence because the full text is a different paper. That is the right call: without the Shopping Companion methods, reward definitions, baselines, or ablations, one cannot locate a soft spot in the preference-memory or tool-wise-reward argument, nor credit independent support. Manufacturing a critique of EnSF/Lorenz-96/KdV would be off-target for the stated paper_id and strongest_claim. Agreement with the reader is therefore full on the mismatch and on the weakest assumption (benchmark/reward validity vs artifacts). Verdict remains UNVERDICTED until the correct manuscript is supplied; no adjustment to ACCEPT/CONDITIONAL/REJECT is justified from the current materials.","tokens_in":21360,"tokens_out":550,"duration_ms":5573,"concrete_test":"Replace the CACHEABLE body with the actual Shopping Companion manuscript (or its arXiv PDF/source for 2603.14864). Re-run the stress test on: (i) how synthetic user trajectories and the 1.2M product pool are constructed, (ii) exact definition of the two cross-session tasks and success metrics, (iii) the tool-wise reward formulas and whether they can be gamed without true preference capture, and (iv) whether the 4B gains hold under held-out product/user splits. If those checks pass, the abstract claim stands; if not, the reward/benchmark validity concern lands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption is well-posed for arXiv:2603.14864 (Shopping Companion Bench + tool-wise rewards), but the CACHEABLE full manuscript is arXiv:2603.14863 (score-filter-enhanced data assimilation for ML dynamical systems: LSTM/R-DeepONet + EnSF on Lorenz-96 and KdV). No sections, equations, tables, or experiments from Shopping Companion are present to audit. Therefore no load-bearing technical flaw in the shopping-agent claim (preference hallucination, attribute verification, annotation-free tool-wise rewards, or 4B vs GPT-5 <70%) can be identified or refuted from the supplied text. The mismatch itself is the only concrete issue: the central claim of 2603.14864 is unsupported by the body that was provided, leaving the abstract uncheckable rather than internally inconsistent.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The submission is presented under the title and abstract of “Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks” (arXiv:2603.14864, cs.CL). That abstract claims a new long-horizon preference-grounded shopping benchmark (Shopping Companion Bench) over a 1.2M-item product pool, a failure-mode analysis (preference hallucination and attribute verification), annotation-free tool-wise process rewards, and experimental results in which GPT-5 stays below 70% success while a fine-tuned 4B model outperforms strong baselines. The body that was supplied, however, is an entirely different manuscript: “A Score Filter Enhanced Data Assimilation Framework for Data-Driven Dynamical Systems” (content matching arXiv:2603.14863), which develops an Ensemble Score Filter (EnSF) hybrid with LSTM and R-DeepONet surrogates and evaluates it on Lorenz–96 and the KdV equation. No Shopping Companion methods, tasks, rewards, or results appear in the body.","tokens_in":21587,"tokens_out":929,"duration_ms":15976,"significance":"If the abstract’s claims were supported by a matching manuscript, a hard long-horizon preference-memory shopping benchmark with process-level, annotation-free tool rewards and a competitive lightweight agent would be of clear interest to the LLM-agent and e-commerce communities. As submitted, that contribution cannot be assessed: the body contains no benchmark definition, no agent training pipeline, and no e-commerce experiments. The EnSF–ML assimilation work in the body is a separate, technically coherent line of research on uncertainty reduction for data-driven dynamical systems, but it is not the paper announced by the title and abstract. Significance of the claimed Shopping Companion result is therefore indeterminate from this package.","major_comments":[{"comment":"Title/abstract vs. body mismatch: the abstract and paper_id announce Shopping Companion Bench, tool-wise rewards, GPT-5 <70% success, and a fine-tuned 4B agent, but Sections 1–5 and all figures/tables develop EnSF-enhanced LSTM/R-DeepONet forecasting on Lorenz–96 and KdV. There is no section defining the two shopping tasks, the 1.2M product pool, preference memory, tool-wise rewards, or any LLM-agent experiment. The central claims of 2603.14864 are therefore unsupported by the supplied manuscript and cannot be audited for soundness, baselines, or ablations.","section":"Title, Abstract, and full body (Sections 1–5)"},{"comment":"Because the body is a different paper, load-bearing Shopping Companion claims (preference hallucination as a cascading failure mode; annotation-free tool-wise rewards alleviating reward sparsity; 4B outperformance of strong baselines; GPT-5 success below 70%) have no methods, equations, tables, or error bars against which to check them. Evaluation of those claims is not possible without the correct manuscript body.","section":"Abstract claims vs. Sections 4.1–4.2 (which report only Lorenz–96/KdV RMSE)"}],"minor_comments":[{"comment":"The supplied body (EnSF paper) itself has ordinary presentation issues—e.g., garbled math glyphs in the LSTM equations (Section 2.1), inconsistent notation for observation operators, and figure captions that sometimes restate results without quantitative detail—but these are secondary to the abstract/body identity failure and would only matter if that paper were the submission under review.","section":"Section 2.1, Eq. (1); Figures 3–15"}],"recommendation":"reject","confidential_remarks":"This appears to be a packaging or cache error: paper_id/title/abstract for 2603.14864 (cs.CL Shopping Companion) were attached to the full text of 2603.14863 (math.DS EnSF data assimilation). I cannot fairly accept, revise, or scientifically reject the Shopping Companion contribution itself—only the mismatched submission package. Recommend returning to the authors (or the pipeline) for the correct PDF before any scientific review of either paper. If the journal intended review of the EnSF manuscript alone, that should be resubmitted under its own title and abstract."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline first: the manuscript body in the cache is not Shopping Companion. It is a score-filter / EnSF data-assimilation paper on Lorenz-96 and KdV. For arXiv:2603.14864 we effectively have the abstract only. Treat every number below as unverified.\n\nFrom that abstract alone, the intended contribution is clear and useful within agents/e-commerce: a long-horizon preference-memory shopping benchmark on a large real catalog (~1.2M items), a failure-mode diagnosis (preference hallucination cascading, weak attribute checks), and annotation-free tool-wise process rewards aimed at sparse long-horizon credit. Claiming GPT-class models stay under 70% and that a fine-tuned 4B beats strong baselines is the kind of result people in the area would actually use if it holds.\n\nWhat I cannot check: task construction and synthetic-user validity, whether “success” is gamed by the tool interface, ablations that separate the reward design from the benchmark, baselines, variance, and any leakage between reward signals and evaluation. The circularity risk is structural—same group defines tasks, rewards, and metrics—not proof of bad faith, but it is the main soft spot until independent reimplementation or external validation exists.\n\nThe EnSF body that was supplied is a different, serious-looking dynamical-systems paper; it does not support or refute the shopping claims. So this is not a soft spot in Shopping Companion’s math; it is a packaging/mismatch problem that blocks review of the titled work.\n\nWho it is for: people building shopping or long-horizon tool agents who need harder preference-memory evals. Value is real if the full paper and artifacts match the abstract. I would send a complete, consistent manuscript to peer review. I would not cite or run a reading group on the abstract alone. Get the correct PDF and code before spending more time.","headline":"We only have the Shopping Companion abstract; the attached full text is a different paper (EnSF data assimilation), so the shopping claims cannot be audited.","tokens_in":22186,"tokens_out":484,"would_cite":false,"duration_ms":15646,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Annotation-free tool-wise rewards let a fine-tuned 4B model outperform strong baselines on long-horizon preference-grounded shopping, where even frontier models stay under 70% success.","keywords":["LLM agents","e-commerce","preference memory","process supervision","tool-wise rewards","shopping benchmark","long-horizon tasks"],"falsifier":"An independent evaluation with real human shoppers or a held-out product pool in which the 4B model no longer beats baselines, or in which tool-wise rewards fail to reduce preference hallucination relative to terminal-only rewards, would falsify the central claim.","tokens_in":22250,"feed_emoji":"🛒","tokens_out":803,"duration_ms":19456,"temperature":0.7,"pith_summary":"LLM shopping agents must track user preferences across long, multi-session conversations, yet the field has lacked both hard evaluation suites and dense training signals. This paper introduces Shopping Companion Bench: two tasks that require cross-session preference memory over a pool of more than 1.2 million real products. Failure analysis isolates two main error sources—cascading preference hallucination and weak verification of product attributes against stated requirements. To fix them, the authors design annotation-free rewards that score each tool call, giving process supervision instead of a single sparse terminal signal. A lightweight 4B model trained with these rewards consistently beats strong baselines on preference capture and task success, while even state-of-the-art models remain below 70% success, showing both the benchmark’s difficulty and the reward design’s value.","feed_headline":"4B shopping agent beats larger models with tool rewards","feed_subtitle":"New 1.2M-product benchmark leaves even top models under 70%; dense process rewards close the gap.","key_machinery":"Shopping Companion Bench plus annotation-free tool-wise rewards that assign process supervision to every tool call; these rewards specifically target preference hallucination and attribute-verification failures so credit is no longer delayed until task end.","core_discovery":"Annotation-free, tool-wise process rewards alleviate reward sparsity on long-horizon preference-grounded shopping tasks, enabling a fine-tuned 4B model to outperform strong baselines on Shopping Companion Bench—a new two-task suite requiring cross-session preference memory over 1.2 million real items—where even models such as GPT-5 stay below 70% success.","pith_inferences":["The same tool-wise process rewards could improve agents for multi-session preference tasks outside shopping, such as travel planning or medical intake.","Larger models may close or reverse the gap if they are also trained with the same dense tool rewards rather than only terminal success.","Reported gains will largely depend on how faithfully the synthetic user trajectories match real shopper behavior."],"forward_implications":["Long-horizon shopping agents can be trained without expensive human process annotations.","Lightweight models become competitive for preference-grounded e-commerce once dense tool-level rewards are available.","Preference hallucination and attribute verification become first-class failure modes that future agents must explicitly address.","Cross-session memory benchmarks over large real catalogs will stay hard for frontier models until process supervision improves.","Tool-level reward design can transfer to other multi-step agent domains that suffer from sparse terminal rewards."],"fun_headline_variants":["4B agent beats larger models on shopping bench via tool rewards","Tool-wise rewards lift 4B model past SOTA on preference shopping tasks","Shopping Companion Bench: 4B with process rewards tops GPT-5 sub-70%","Dense tool rewards fix sparsity so 4B shopping agent leads baselines","Cross-session shopping bench: fine-tuned 4B outperforms stronger models"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That success on the two constructed tasks under the authors’ tool-wise rewards measures real long-horizon shopping competence rather than artifacts of synthetic user trajectories, product-pool construction, or reward hacking of the tool interface.","fun_headline_variants_meta":{"raw":{"variants":["4B agent beats larger models on shopping bench via tool rewards","Tool-wise rewards lift 4B model past SOTA on preference shopping tasks","Shopping Companion Bench: 4B with process rewards tops GPT-5 sub-70%","Dense tool rewards fix sparsity so 4B shopping agent leads baselines","Cross-session shopping bench: fine-tuned 4B outperforms stronger models"]},"model":"grok-4.5","effort":"low","cost_usd":0.003708,"raw_usage":{"total_tokens":1206,"prompt_tokens":786,"num_sources_used":0,"completion_tokens":101,"cost_in_usd_ticks":37080000,"prompt_tokens_details":{"text_tokens":786,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":319,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":786,"tokens_out":101,"duration_ms":3828,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T20:51:26.325698+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"An independent evaluation with real human shoppers or a held-out product pool in which the 4B model no longer beats baselines, or in which tool-wise rewards fail to reduce preference hallucination relative to terminal-only rewards, would falsify the central claim.","supporting_citations":[],"review_version":1}