{"id":"5699e498-4b29-4379-bfb5-7a321a0f284e","arxiv_id":"2508.05398","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"Simulated exposure conditions on fully observed preference data show that the reliability of offline recommender evaluation depends jointly on the logging exposure model and the sampling strategy, and the paper maps which strategies stay faithful and robust.","lead":"This paper tests how the choice of sampled items in offline recommender evaluation changes the reliability of the evaluation itself, using a fully observed preference dataset as ground truth under simulated exposure biases. The payoff is practical guidance on when sampling distorts model rankings, so evaluation setups can be chosen to survive real exposure conditions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Potential circularity between exposure simulation and sampling strategies may predetermine the reported ranking; unreadable body prevents checking.","rationale":"The reader's verdict is UNVERDICTED because the full text is unreadable and the header mismatch prevents verification. My concern is more specific: even if the text were readable, the design may contain a circularity whereby the exposure simulation and the sampling strategies share the same item-selection distribution, making the 'predictive power' comparison partially tautological. This is a load-bearing concern because the paper's practical guidance for selecting sampling strategies depends on the ranking being meaningful. However, this is a testable empirical concern, not a proven flaw. The concrete test I propose would settle it by using an exposure model that breaks the shared distribution. I agree with the reader that the submission as-is lacks sufficient information to assess soundness, but I emphasize the specific mechanism that should be checked. The reader's weakest_assumption about ground-truth validity overlaps with my concern, but they did not identify the shared-distribution channel explicitly, hence 'partial'. Since my concern does not change the verdict—UNVERDICTED remains appropriate given the unreadable body—I set verdict_should_be to UNCHANGED.","tokens_in":24221,"tokens_out":2639,"duration_ms":28491,"concrete_test":"Independently reimplement the ground-truth dataset and exposure simulation. Rerun the four-dimension evaluation with an alternative exposure mechanism that is deliberately uncorrelated with the item-popularity distribution used by the sampling strategies (e.g., exposure driven by item recency or random noise). If the ranking of sampling strategies under 'predictive power' changes materially, the original ranking is likely an artifact of shared popularity assumptions; if the ranking persists, the concern is mitigated. Also verify that the readable body header matches arXiv:2508.05398.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that specific combinations of logging exposure and item-sampling strategies affect offline evaluation reliability, with some combinations preserving true preference rankings and others reversing them. For this to be a meaningful empirical finding, the ground-truth dataset and the simulated exposure process must be independent of the sampling strategies being tested. The abstract says 'Using a fully observed dataset as ground truth, we systematically simulate diverse exposure biases.' If that dataset is synthetic, its generative preference model is an unstated free choice, and every reliability measurement inherits it. More specifically, if the simulated exposure probability is a function of item popularity and the sampling strategies under test also select items based on popularity or a similar marginal distribution, then the 'predictive power' dimension (alignment with ground truth) may be partly circular: a sampling strategy will appear to recover true preferences simply because it mirrors the same item-selection distribution used to generate the logged data. In that case the ranking of strategies is not a general result about sampling bias but an artifact of shared distributional assumptions. The supplied full text is mojibake and the header carries a different arXiv identifier (2508.05395v2, astro-ph.CO) than the stated submission (2508.05398, cs.IR), so the actual methodology—dataset source, exposure families, parameter ranges, and whether any independent validation was performed—cannot be audited. This makes the circularity concern unresolvable from the current submission.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims that the reliability of offline recommender evaluation is determined by the interaction between the logging exposure model and the item-sampling strategy, rather than by either factor alone. Using a 'fully observed dataset' as ground truth, the authors say they simulate diverse exposure biases and evaluate common sampling strategies on four dimensions: sampling resolution, fidelity, robustness, and predictive power. The stated contribution is practical guidance for selecting sampling strategies that yield faithful and robust offline comparisons. However, the supplied full text is unreadable mojibake, and the visible header bears a different arXiv identifier (2508.05395v2, astro-ph.CO) than the cited submission (2508.05398, cs.IR). Consequently, no methods, dataset description, parameterization, equations, tables, or numerical results can be verified from the manuscript as provided.","tokens_in":24388,"tokens_out":2876,"duration_ms":31637,"significance":"If the claimed interaction between exposure bias and sampling strategy were properly established with a transparent ground-truth dataset, the paper could offer a useful practical contribution to offline recommender evaluation. The four-dimensional evaluation scheme is a reasonable organizing framework, and the central claim that some sampling strategies reverse preference rankings under particular exposure conditions is concrete and falsifiable. However, the current submission provides no readable evidence for these claims. The abstract reports findings without numbers, error bars, significance tests, dataset identity, or simulation parameterization, and the body is unreadable. The potential practical significance is therefore entirely contingent on a version of the paper that can actually be checked.","major_comments":[{"comment":"The supplied full text is unreadable mojibake. Sections, equations, tables, and results cannot be recovered. No methodology, dataset source, exposure-bias family, parameter ranges, or statistical analyses can be inspected. Because the paper's central claim is an empirical one, this is not a cosmetic issue; the manuscript as submitted cannot support any of its stated findings.","section":"Full text (entire body)"},{"comment":"The visible header reads 'arXiv:2508.05395v2 [astro-ph.CO] 24 Jun 2026,' which does not match the stated submission 'arXiv:2508.05398 (cs.IR).' This mismatch, combined with the mojibake, raises a question about whether the supplied text is the actual manuscript under review. The authors must provide a clean, correctly identified version.","section":"arXiv header / document identity"},{"comment":"The abstract states that the study uses 'a fully observed dataset as ground truth' but does not name the dataset or state whether it is real or synthetic. It also does not specify the exposure-bias families, the set of sampling strategies, or the parameter ranges simulated. Without these details, the claimed ranking of sampling strategies cannot be interpreted or reproduced. If the dataset is synthetic, the unstated generative preference model defines what 'true preference' means and every reliability measurement inherits it.","section":"Abstract"},{"comment":"The reader's concern about potential circularity between the exposure simulator and the sampling strategies cannot be adjudicated from the readable portions. To rule it out, the paper must show that the simulated exposure process and the sampling strategies under test are not driven by the same item-selection distribution. If both depend on item popularity in the same way, the 'predictive power' dimension may partly reward strategies that mirror the logging distribution. This issue is load-bearing and must be addressed explicitly in a readable version.","section":"Abstract / predictive power definition"}],"minor_comments":[{"comment":"The four dimensions (sampling resolution, fidelity, robustness, predictive power) are listed but not defined in the readable portion. Formal definitions are needed.","section":"Abstract"},{"comment":"The abstract reports 'findings' without any quantitative summary. Adding a small set of headline numbers or a pointer to a results table would help the reader assess the contribution.","section":"Abstract"},{"comment":"The document suffers from an encoding corruption that makes the body unreadable. A correctly encoded PDF or LaTeX source must be provided.","section":"Document preparation"},{"comment":"The paper should state whether code and data will be released, including the ground-truth preference scores and the exposure simulation code, to allow independent verification.","section":"Reproducibility"}],"recommendation":"uncertain","confidential_remarks":"The core issue is that the supplied manuscript is not readable, and the header does not match the stated paper. I cannot certify soundness or any specific scientific claim. The circularity concern raised by the stress-test is plausible but unverifiable from this text. If the authors can submit a clean, correctly identified version with full methodological details, I would be willing to re-review. Otherwise, a desk rejection may be appropriate because the current artifact cannot be evaluated as a research paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know before anything else. First, the research question is a good one: prior work tests sampling-bias corrections on fixed logged data; this paper proposes to vary the exposure model while holding a fully observed dataset as ground truth, and to rank sampling strategies along resolution, fidelity, robustness, and predictive power. That is a legitimate extension, and the four dimensions give practitioners a useful vocabulary for reporting evaluation reliability. Second, the supplied full text is unreadable mojibake and its header carries arXiv:2508.05395v2 [astro-ph.CO], not the stated cs.IR submission. I could not read a single sentence of the body. Everything I say below is based on the abstract alone.\n\nWhat the paper does well, based on what I can see: the design avoids the most obvious circularity. “Predictive power” compares sampled evaluations to ground-truth preferences; “fidelity” compares them to full-log evaluation. Those are distinct definitions. The claim that logging and sampling interact is plausible and consistent with what we know about selection bias. If the body backs up the abstract, this is a solid contribution to the evaluation-methodology subfield.\n\nSoft spots. The abstract gives no numbers, no dataset identity, no exposure family, no parameter ranges. The stress-test worry about circularity — that if the ground-truth data is synthetic and the exposure simulation shares item-selection assumptions with the sampling strategies under test, the ranking could be partly baked in — is legitimate, but it is not established by anything in the abstract. There is no sign that the sampling strategies are derived from the same distribution as the exposure simulation. I would treat that as a reviewer question, not a known flaw. The unreadable body is the real problem: it prevents checking the methodology, the dataset, the error bars, everything. The header mismatch compounds it; we cannot even confirm that the body belongs to this submission. Whether this is a file-encoding accident or something more serious, the editor needs a readable copy before any verdict.\n\nBottom line: the thinking behind the abstract is clear and the contribution, if realized, is worth referee time. But as submitted, the paper is unverdictable. I would ask the authors to resubmit a readable PDF, ideally with code and data, and then send it to peer review. Don’t desk-reject the idea; desk-reject this artifact.","headline":"The abstract poses a genuinely useful question about sampling and exposure in offline recommender evaluation, but the supplied full text is unreadable mojibake with a mismatched arXiv header, so the body cannot be audited and the claims are unverifiable.","tokens_in":24996,"tokens_out":2736,"would_cite":false,"duration_ms":26622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Offline recommender evaluation can be distorted by the interaction between how items were logged and how they are sampled, not by either choice alone.","keywords":["offline evaluation","recommender systems","exposure bias","sampling bias","evaluation reliability","item sampling strategies","model comparison","ground truth"],"falsifier":"Let another group run the same exposure-and-sampling protocol on a different fully observed dataset (or a real logged dataset with known interventions) and check whether the relative ranking of sampling strategies on the four dimensions remains the same; if the ranking flips, the paper's guidance is dataset-specific rather than general.","tokens_in":23981,"feed_emoji":"🎯","tokens_out":2917,"duration_ms":27554,"temperature":0.7,"pith_summary":"The paper tries to establish when and how item-sampling strategies distort offline recommender evaluation. It treats user interactions as logged under controllable exposure biases and compares sampling strategies against a fully observed ground truth of preferences. The central claim is that reliability has four separable dimensions: resolution, fidelity, robustness, and predictive power, and that no single sampling strategy dominates on all four. A sympathetic reader should care because offline benchmarks are the standard way to compare recommenders when online testing is risky, and a sampling choice that reverses model rankings would invalidate those comparisons.","feed_headline":"Sampling choice can flip offline recommender rankings","feed_subtitle":"Simulated exposure biases show when item sampling stays faithful and when it inverts true preferences.","key_machinery":"The evaluation harness is a fully observed preference dataset used as ground truth, into which exposure biases are injected to generate logged interactions; common sampling strategies are then scored on four dimensions. The four-dimension profile is the central object, because it separates the distinct failure modes that a single accuracy number would hide.","core_discovery":"Using a fully observed dataset as ground truth, the authors simulate several exposure biases and ask whether common item-sampling strategies still allow correct model comparisons. They assess reliability along four dimensions: sampling resolution (how well the sample separates competing recommenders), fidelity (agreement with evaluating on the full logged set), robustness (stability of the evaluation as exposure bias changes), and predictive power (alignment with ground-truth preferences). The paper finds that these dimensions behave differently across sampling strategies, so the safest choice depends on the logging exposure model in place; in some combinations sampling preserves the true ra","pith_inferences":["The ranking of sampling strategies is probably conditional on the choice of fully observed ground-truth dataset; if that dataset's preference distribution changes, the recommended strategy may change.","If the ground-truth dataset is synthetic, the generative preference model is a free parameter that every reliability measurement inherits; disclosing it would let readers check whether exposure simulation and sampling share assumptions.","The four-dimension profile could be used as a template for designing new sampling strategies: a strategy would be an improvement only if it shifts the profile outward on at least one dimension without degrading the others.","One testable extension would be to run the same protocol on multiple fully observed datasets with different long-tail and popularity structures, to see which ranking of sampling strategies persists."],"forward_implications":["Practitioners can choose a sampling strategy by which failure mode matters most: separability, agreement with full evaluation, stability under exposure, or agreement with true preferences.","Results obtained under one exposure model should not be assumed to transfer to another; logging and sampling must be considered together.","A sampling strategy that scores high on fidelity can still have low predictive power, so agreement with the full log is not evidence that the evaluation reflects true preferences.","Reporting all four dimensions would make offline recommender comparisons more honest and easier to reproduce."],"supporting_citations":[],"fun_headline_variants":["Sampling can distort offline recommender comparisons","Exposure bias decides if item sampling stays reliable","Sampling strategy can invert true recommender preferences","Offline eval: item sampling may flip model rankings","How exposure bias changes what sampling tells us"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The fully observed dataset really does contain the users' true preferences, and the simulated exposure biases do not secretly share assumptions with the sampling strategies being tested.","fun_headline_variants_meta":{"raw":{"variants":["Sampling can distort offline recommender comparisons","Exposure bias decides if item sampling stays reliable","Sampling strategy can invert true recommender preferences","Offline eval: item sampling may flip model rankings","How exposure bias changes what sampling tells us"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000621,"raw_usage":{"total_tokens":2685,"prompt_tokens":680,"completion_tokens":2005,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":1935}},"tokens_in":424,"tokens_out":2005,"duration_ms":14493,"temperature":1.0,"reasoning_tokens":1935,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:22:08.163338+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Let another group run the same exposure-and-sampling protocol on a different fully observed dataset (or a real logged dataset with known interventions) and check whether the relative ranking of sampling strategies on the four dimensions remains the same; if the ranking flips, the paper's guidance is dataset-specific rather than general.","supporting_citations":[],"review_version":1}