{"id":"01d2bdfa-2347-444f-9401-f9d1ee70e743","arxiv_id":"2508.16487","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract claims a sample-optimal, O(KL^2) preference exploration algorithm, but the supplied full text is an unrelated medical imaging manuscript.","lead":"The abstract of this submission announces FraPPE, a fast algorithm for preference-based pure exploration in multi-objective bandits, claiming optimal sample complexity and O(KL^2) runtime. The full text supplied is a different paper, a deep learning study on histology and transcriptomics for cancer diagnosis, so the announced results cannot be checked from this manuscript.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim unsupported: supplied full text is an unrelated medical-imaging paper; verify actual arXiv manuscript before any technical assessment.","rationale":"The review rule requires treating all supplied manuscript text as in-scope evidence. Here the supplied full text is an unrelated medical-imaging paper, so the submission as provided cannot support the FraPPE central claim. The correct verdict remains UNVERDICTED, as the reader concluded. My concern does not move the verdict; it reinforces it. I partially agree with the reader's weakest assumption: the Frank-Wolfe global-optimality condition is a plausible fragile point, but it is secondary to the more fundamental issue that the actual FraPPE manuscript is absent. The proposed test—fetching the real arXiv record—would settle whether the missing content exists and whether the technical concern can be evaluated.","tokens_in":21149,"tokens_out":3754,"duration_ms":46033,"concrete_test":"Retrieve the actual arXiv:2508.16487 record (via arXiv API or PDF download) and check whether the full text contains the FraPPE algorithm, the three structural properties, the Frank-Wolfe-based maxmin solution with O(KL^2) complexity, and the asymptotic optimality proof. If the full text does not match the metadata or lacks these components, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that FraPPE achieves optimal sample complexity with O(KL^2) runtime—depends entirely on derivations, proofs, and experiments in the manuscript body. The supplied full text is not that manuscript: it is an IEEE TMI article on disentangled histology/transcriptomics learning, with a different title, authors, and arXiv identifier (2508.16479v2 vs. 2508.16487). None of the three structural properties, the Frank-Wolfe optimization details, the sample-complexity theorem, or the experimental comparisons are present. The reader's weakest assumption—that Frank-Wolfe must provably reach a global optimum—cannot even be checked without the actual text. Thus the evidence base for the central claim is empty: no algorithm, no proof, no runtime analysis, no experiments. This is the load-bearing gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission is indexed as arXiv:2508.16487 (cs.LG), titled 'FraPPE: Fast and Efficient Preference-based Pure Exploration.' The abstract claims a new algorithm, FraPPE, for preference-based pure exploration in vector-valued bandits, with three structural properties of a lower bound enabling a tractable minimisation reduction, a Frank-Wolfe optimiser accelerating the maximisation, an overall O(KL^2) time for solving the maxmin problem, and asymptotic optimal sample complexity matching an existing lower bound for arbitrary preference cones. However, the supplied full text is not this manuscript. It is an IEEE Transactions on Medical Imaging paper titled 'Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization' (arXiv:2508.16479v2), with a different abstract, methods, experiments, and references. Consequently, none of the claimed algorithmic contributions, proofs, runtime analyses, or experimental comparisons for FraPPE are present in the submitted text.","tokens_in":21305,"tokens_out":1943,"duration_ms":23339,"significance":"If the claims in the abstract were substantiated, the work would address a genuine open problem: no existing PrePEx algorithm is both sample-optimal and computationally efficient for arbitrary preference cones. The claimed O(KL^2) maxmin solver and asymptotic sample optimality would be a notable advance for multi-objective pure exploration. However, the submitted manuscript body contains no algorithm statement, no theoretical derivation, no proofs, and no experiments for FraPPE. The central claims are therefore entirely unsupported by the supplied text, and their significance cannot be assessed. This is not a case of a defensible claim with local gaps; the evidence base for every load-bearing assertion is absent.","major_comments":[{"comment":"The supplied full text is an unrelated medical-imaging paper: 'Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization,' with footer identifier arXiv:2508.16479v2, while the submission is arXiv:2508.16487. It contains no definition of FraPPE, no preference-cone bandit setup, no lower-bound equations, no Frank-Wolfe analysis, and no sample-complexity theorem. The abstract's central claims—asymptotic optimal sample complexity and O(KL^2) maxmin solving—are thus unsupported by any derivable evidence in the manuscript. This is a load-bearing defect that invalidates the submission as a reviewer artifact.","section":"Full Text (entire body)"},{"comment":"The abstract asserts 'three structural properties of the lower bound' and a Frank-Wolfe-based maximisation acceleration, but none of these properties, the resulting reduction, or the convergence conditions for Frank-Wolfe are presented anywhere in the submitted text. In particular, the claimed O(KL^2) runtime depends on the unstated convexity/smoothness and global-optimality conditions of the maxmin problem. Since the body is a different paper, not even the notation K, L, and C is defined. These are not presentation issues; they are the core technical content needed to verify the claims.","section":"Abstract"}],"minor_comments":[{"comment":"The manuscript footer shows 'arXiv:2508.16479v2 [eess.IV] 28 Feb 2026,' inconsistent with the stated submission ID 2508.16487. This mismatch corroborates that the wrong full text was provided.","section":"Full Text footer"},{"comment":"The title and abstract refer to FraPPE and preference-based pure exploration, whereas the body is titled and written about histology/transcriptomics multi-modal learning. The author names on the supplied body also do not match those implied by the FraPPE submission.","section":"Title/Abstract"}],"recommendation":"reject","confidential_remarks":"To the editor: this submission appears to have been assembled with the wrong full text: the body is an unrelated IEEE TMI paper (arXiv:2508.16479v2). No technical content for FraPPE is available for review. If this is a pipeline or submission error, the correct manuscript should be requested; however, as submitted, the paper cannot be reviewed and the central claims are unverifiable. I recommend a desk rejection with an opportunity to resubmit the correct manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: you cannot review what isn't there. The metadata and abstract describe FraPPE, a preference-based pure exploration algorithm with a claimed O(KL^2) maxmin solve and asymptotic optimality. The full text we're given is an IEEE TMI paper on disentangled histology and transcriptomics, with a different title, authors, and arXiv number (2508.16479v2, not 2508.16487). That is not a minor formatting error. It means every technical claim—the three structural properties, the Frank-Wolfe step, the sample-complexity theorem, the experiments—is completely absent from the document under review.\n\nWhat deserve credit: the abstract itself is plainly written and the claimed gap is concrete. If FraPPE really is the first computationally efficient, sample-optimal PrePEx algorithm for arbitrary preference cones, that's a real result for the multi-objective bandit community. The O(KL^2) improvement over existing maxmin solvers would be the kind of acceleration people care about. But an abstract is not a paper.\n\nSoft spots, in order: first, the document inconsistency is fatal at the desk level. Second, the reader's flagged concern about Frank-Wolfe convergence is reasonable but currently uncheckable—we don't even have the objective function, let alone its convexity or smoothness. That's a 'we cannot assess' rather than a detected flaw. Third, there's no circularity visible because there are no equations; the claimed upper-bound/lower-bound match is plausible as stated.\n\nMy read: this is a submission-error, not a science-review case. The right move is to bounce it back to the authors to confirm the arXiv ID and provide the actual FraPPE manuscript. If the real paper contains the derivations the abstract promises, it may well deserve a careful referee. But the document on the table does not. I would not cite it, would not bring it to reading group, and would not send it to peer review as-is.","headline":"The abstract promises a genuinely useful bandit result, but the supplied full text is an unrelated histology paper, so there is no manuscript to review.","tokens_in":21788,"tokens_out":2195,"would_cite":false,"duration_ms":22718,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L05","90C29","90C52"],"pacs":[],"model":"deepseek-v4-flash","headline":"A new bandit algorithm identifies Pareto-optimal arms sample-optimally for any preference cone, with O(KL²) computation time.","keywords":["preference-based pure exploration","multi-objective bandits","Pareto optimal set","preference cone","sample complexity lower bound","Frank-Wolfe optimization","maxmin optimization","pure exploration"],"falsifier":"Run FraPPE on a preference cone and arm reward distributions where the lower-bound maximization is known to be non-convex; if the optimizer stops at a local optimum and the sample complexity exceeds the lower bound, the claim fails. Concretely, a synthetic instance with K=3, L=2 and a pointed cone that creates a saddle in the maxmin objective is a direct test of the O(KL²) and asymptotic optimality claims.","tokens_in":20977,"feed_emoji":"🎯","tokens_out":5980,"duration_ms":60748,"temperature":0.7,"pith_summary":"Preference-based Pure Exploration (PrePEx) asks which arms of a multi-objective bandit are Pareto-optimal under a given preference cone. This paper proposes FraPPE, an algorithm that asymptotically matches the known lower bound on the number of samples needed to answer that question with confidence, for any arbitrary preference cone. The central technical move is to make the lower bound's maxmin optimization tractable: three structural properties reduce the minimization side, and a Frank-Wolfe optimizer accelerates the maximization side. Together they solve the maxmin problem in O(KL²) time for K arms and L reward dimensions, a speedup over earlier algorithms. If FraPPE is correct, sample-optimal Pareto-set identification is no longer computationally out of reach.","feed_headline":"Frank-Wolfe cuts Pareto bandit search cost to O(KL²)","feed_subtitle":"First PrePEx algorithm to match the sample lower bound for arbitrary preference cones","key_machinery":"The central object is the maxmin optimization inside the PrePEx lower bound: minimize over plausible Pareto-set configurations and maximize over reward distributions consistent with the instance. The machinery is a two-part attack on this problem. Three structural properties of the lower bound turn the minimization into a computationally tractable reduction, while a Frank-Wolfe optimizer, a conditional-gradient method suited to convex and smooth objectives with sparse iterates, accelerates the maximization. The combination is what brings the per-iteration cost to O(KL²) and lets the algorithm track the lower bound asymptotically.","core_discovery":"FraPPE, a PrePEx algorithm, is the first computationally efficient algorithm that achieves asymptotic sample optimality for arbitrary preference cones. It tracks the existing lower bound by solving the lower bound's maxmin problem in O(KL²) time: the minimization over candidate Pareto sets is reduced to a tractable form using three structural properties of the lower bound, and the maximization over reward distributions is handled by a Frank-Wolfe optimizer. The paper proves that FraPPE asymptotically achieves the optimal sample complexity, and experiments on synthetic and real datasets report the lowest sample complexities among existing PrePEx algorithms for exact Pareto-set identification.","pith_inferences":["Frank-Wolfe's convergence depends on the lower-bound maximization objective being convex and smooth; the abstract does not state those regularity conditions, so the O(KL²) guarantee and the asymptotic optimality are conditional on them.","The three structural properties of the lower bound may transfer to other active-learning or ranking problems whose objective can be written as a maxmin over a structured family, enabling similar speedups.","A testable extension is to benchmark FraPPE against the lower bound on synthetic instances with a non-polyhedral cone, such as a Lorentz cone, where the Frank-Wolfe acceleration may require more iterations or stall at a local optimum.","If Frank-Wolfe's iterates are sparse, FraPPE might effectively identify not only the Pareto set but also a small set of arms and reward directions that dominate the sample cost, an interpretability byproduct the authors do not advertise."],"forward_implications":["Sample-optimal Pareto-set identification becomes feasible for bandits with many arms and reward dimensions, because the per-decision cost no longer grows intractably.","Arbitrary preference cones, not just the positive orthant, can be handled without a computational penalty, widening applicability to lexicographic, polyhedral, or other dominance orders.","The O(KL²) maxmin solver can be reused as a subroutine in other pure-exploration algorithms whose lower bounds share the same maxmin structure.","FraPPE's sample complexity matches the lower bound asymptotically, so further gains must come from constant factors or a different problem formulation."],"supporting_citations":[],"fun_headline_variants":["First PrePEx algorithm to hit sample lower bound","FraPPE cracks arbitrary-cone Pareto bandits in O(KL²)","Optimal sample complexity for any preference cone","Frank-Wolfe powers fastest Pareto set search"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The runtime and sample-optimality guarantees both depend on the Frank-Wolfe optimizer reliably reaching the global maximum of the lower-bound problem; if the optimizer can get stuck at a local maximum, the speed and the sample guarantee do not follow.","fun_headline_variants_meta":{"raw":{"variants":["First PrePEx algorithm to hit sample lower bound","FraPPE cracks arbitrary-cone Pareto bandits in O(KL²)","Optimal sample complexity for any preference cone","Frank-Wolfe powers fastest Pareto set search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000121,"raw_usage":{"total_tokens":925,"prompt_tokens":734,"completion_tokens":191,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":126}},"tokens_in":478,"tokens_out":191,"duration_ms":2683,"temperature":1.0,"reasoning_tokens":126,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:15:20.955806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run FraPPE on a preference cone and arm reward distributions where the lower-bound maximization is known to be non-convex; if the optimizer stops at a local optimum and the sample complexity exceeds the lower bound, the claim fails. Concretely, a synthetic instance with K=3, L=2 and a pointed cone that creates a saddle in the maxmin objective is a direct test of the O(KL²) and asymptotic optimality claims.","supporting_citations":[],"review_version":1}