{"id":"12843e06-9e1b-4862-bd65-522e1da9a079","arxiv_id":"2508.17083","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The metadata announces hashing and cancer-genomics methods, but the full text is a different neural SDE paper, making the listed paper's claims unverifiable.","lead":"The submission's metadata and abstract describe 'Learning ON Large Datasets Using Bit-String Trees', a thesis covering hashing, classification, and cancer genomics. The supplied full text, however, is an unrelated ProbML 2026 paper on neural stochastic differential equations for suicide risk modeling (arXiv 2508.17090v4), so the listed paper's claims cannot be verified from this submission.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The submitted full text is an unrelated paper, so no claim in the abstract can be checked; the verdict must remain UNVERDICTED until the matching text is supplied.","rationale":"The reader's verdict correctly identifies that this submission cannot be reviewed as a research preprint because the supplied full text does not match the metadata. The abstract of arXiv:2508.17083 describes ComBI, GRAF, uGRAF, and CRCS for large-scale hashing, classification, and cancer genomics, while the full text is 'Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling' from arXiv:2508.17090v4. None of the abstract's algorithmic or experimental claims can be checked against this text. There is also an internal inconsistency between the title ('Bit-String Trees') and the abstract ('Binary Search Trees'). Under the review rules, an unusual inserted passage must be flagged and weighed; here it is decisive. The verdict is UNVERDICTED with UNKNOWN confidence. The abstract-level scores are placeholders that reflect the absence of verifiable evidence, not a judgment on the underlying science. The authors should correct the submission by attaching the full text that matches the abstract.","tokens_in":37708,"tokens_out":1297,"duration_ms":13516,"concrete_test":"Download the manuscript actually associated with arXiv:2508.17083 from the arXiv API (or ask the authors for the matching full text) and confirm that its title, abstract, and body consistently describe ComBI, GRAF, uGRAF, and CRCS. Then check that the benchmark section reports dataset sizes and sources, train/test splits, the precision-recall curve or recall level at which 0.90 precision is measured, the hardware/software environment, and the hyperparameters and search configurations used for Multi-Index Hashing and Cellfishing.jl. If the matching text is unobtainable or omits any of these items, the 4X-296X and 2X-13X speed-up claims remain unverifiable.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim is the abstract's empirical assertion that ComBI achieves 0.90 precision with 4X-296X speed-ups over Multi-Index Hashing at up to one billion samples, and 2X-13X gains over Cellfishing.jl on single-cell RNA-seq. Every load-bearing premise for that claim—the existence and description of the billion-sample datasets, the experimental protocol, the precision-recall operating point, the baseline configurations, and the runtime methodology—lives in the missing full text. The supplied full text, arXiv:2508.17090v4, is a paper on constrained neural stochastic differential equations for suicide-risk modeling; it contains no mention of ComBI, GRAF, uGRAF, CRCS, bit-string trees, Multi-Index Hashing, Cellfishing.jl, or any of the abstract's experiments. Under the review rule requiring in-scope treatment of inserted or mismatched passages, this is decisive: there is no object to stress-test. No internal inconsistency in the abstract itself (beyond the title/abstract terminology mismatch between 'Bit-String Trees' and 'Binary Search Trees') can substitute for the missing evidence, nor can the SDE paper's own theorems and experiments support the abstract's claims. The abstract-level numbers are placeholders; the correctness_risk is unknown because no substantiating material is present. This is not a disagreement with consensus or a methodological quibble; it is the absence of the artifact under review.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission arXiv:2508.17083 consists of an abstract announcing four methods (ComBI, GRAF, uGRAF, CRCS) for similarity-preserving hashing, classification, and cancer genomics, with quantitative claims including 0.90 precision and 4X-296X speed-ups over Multi-Index Hashing on datasets of up to one billion samples. The supplied full text, however, is arXiv:2508.17090v4, a paper on neural stochastic differential equations on compact state spaces applied to suicide-risk modeling. That full text contains no mention of ComBI, GRAF, uGRAF, CRCS, bit-string trees, Multi-Index Hashing, Cellfishing.jl, or any of the abstract's experiments. The stress-test concern is confirmed: as submitted, the manuscript contains only an abstract with no matching technical content that can be checked.","tokens_in":37855,"tokens_out":4661,"duration_ms":43300,"significance":"If the abstract's claims were substantiated, the work would be significant: an order-of-magnitude faster approximate nearest-neighbor search at billion-sample scale with high precision, plus competitive classifiers and cancer-genomics tools, would be useful to the machine learning and biomedical communities. However, the submitted material provides no derivations, algorithms, datasets, experimental protocols, code, or proofs for any of these methods. The supplied full text is an unrelated paper, so the abstract's claims cannot be verified, reproduced, or placed in context. No strength of the claimed contributions can be assessed from the submission as it stands.","major_comments":[{"comment":"The submitted full text is a different paper. It is titled 'Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling' and its abstract, theorems, experiments, and appendices concern viability of SDEs on compact polyhedra and EMA suicide-risk data. It contains no occurrence of ComBI, GRAF, uGRAF, CRCS, bit-string trees, Multi-Index Hashing, Cellfishing.jl, or any of the benchmarks described in the submission's abstract. Because the central claims of the manuscript reside entirely in the abstract, and because the supplied full text does not address them, there is no object to review.","section":"Full Text (supplied manuscript, arXiv:2508.17090v4)"},{"comment":"The abstract's empirical claims are unsupported by any experimental detail. The sentence reporting '0.90 precision with 4X-296X speed-ups over Multi-Index Hashing' and '2X-13X gains' over Cellfishing.jl gives no precision-recall operating point, no dataset construction or size verification, no baseline configuration, no runtime measurement methodology, and no error bars or statistical tests. Even if a matching full text were supplied, these details would be required to evaluate whether the reported numbers are meaningful; in the present submission they are entirely absent.","section":"Abstract"},{"comment":"The only code link and the only experimental tables in the supplied full text belong to the SDE paper, not to the abstract's methods. For example, Table 1 and Appendix G.3 describe EMA forecasting experiments for WSP-based latent neural SDEs, and the GitHub repository linked in the introduction is the WSP demo. These artifacts cannot provide support for the abstract's hashing, classification, or cancer-genomics claims, and no corresponding artifacts for ComBI, GRAF, uGRAF, or CRCS are present.","section":"Full Text / Appendices"}],"minor_comments":[{"comment":"The title uses 'Bit-String Trees' while the abstract defines the approach in terms of 'Binary Search Trees' (BSTs); the relationship between these terms should be clarified or made consistent.","section":"Title / Abstract"},{"comment":"The abstract combines four distinct contributions (ComBI, GRAF, uGRAF, CRCS) in a single submission without indicating how they relate methodologically beyond a shared hashing or tree-based theme; the intended narrative connection should be stated explicitly.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission-integrity issue rather than a scientific disagreement: the attached full text is an unrelated paper. I recommend that the editor desk-reject the current version and invite the authors to resubmit with the correct manuscript. Until the matching full text is provided, no substantive review of the abstract's claims is possible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this one is broken. The submitted full text, arXiv:2508.17090v4, is a paper on constrained neural SDEs for suicide-risk modeling by Lu et al., while the title and abstract describe bit-string trees, ComBI, GRAF, uGRAF, and CRCS for large-scale hashing and cancer genomics. They share no content. No claim in the abstract can be verified against the supplied text. That is the load-bearing problem, and it is decisive.\n\nI can't credit any of the abstract's empirical numbers: there are no datasets, no baselines, no protocols, no code. The abstract's 4X–296X speed-ups and 0.90 precision are just assertions. The title/abstract mismatch ('Bit-String Trees' vs. 'Binary Search Trees') is minor by comparison.\n\nWhat the paper does well: nothing in the claimed work is present. The attached SDE paper is a different contribution, and it looks like real work—it proves constraints for SDEs on compact polyhedra, introduces a parameterization (WSP), and reports experiments on EMA data. But it is not the paper described by the metadata, and its authors don't match the listed author. So I can't treat it as evidence for the abstract's claims.\n\nThe soft spot is the one already stated: there is no object to review. If the abstract is a thesis summary, the thesis itself is missing. If the SDE paper was uploaded by mistake, the author needs to resubmit. Either way, a referee cannot evaluate precision, speed-ups, or survival-prediction results without the matching text. This isn't a methodological quibble; it's an absence of the artifact under review.\n\nMy recommendation: desk reject. Don't send this to referees. Ask the author to correct the submission so the full text matches the abstract and title. If a matching version exists, it can be resubmitted and reviewed on its own merits.","headline":"Not reviewable as submitted: the full text is an entirely different paper, so none of the abstract's claims can be checked.","tokens_in":38477,"tokens_out":2186,"would_cite":false,"duration_ms":22996,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The thesis claims that ComBI, a compressed BST of inverted hash tables, delivers fast approximate nearest-neighbor search at billion-sample scale with 0.90 precision and 4–296× speed-ups over Multi-Index Hashing, along with companion…","keywords":["approximate nearest neighbor search","bit-string trees","compressed BST of inverted hash tables","similarity-preserving hashing","guided random forest","single-cell RNA-seq","cancer genomics","codon switch representation"],"falsifier":"Re-run ComBI against Multi-Index Hashing and Cellfishing.jl on a publicly documented billion-vector dataset at matched recall levels; if the speed-ups fall below the claimed ranges or precision drops well below 0.90 at the standard operating point, the central claim fails.","tokens_in":37394,"feed_emoji":"🔎","tokens_out":5689,"duration_ms":55023,"temperature":0.7,"pith_summary":"The thesis abstract proposes ComBI, a compressed BST of inverted hash tables, as a replacement for standard space-partitioning hashing in approximate nearest-neighbor search. The claimed payoff is fast, memory-reduced search on datasets up to one billion samples, with 0.90 precision and 4–296× speed-ups over Multi-Index Hashing and 2–13× gains over Cellfishing.jl on single-cell RNA-seq searches. The same hashing ideas are extended into GRAF, a guided random forest, and uGRAF, its unsupervised variant, which together with ComBI are said to estimate per-sample classifiability for scalable cancer-patient survival prediction. A third component, CRCS, embeds codon switches into numerical vectors for somatic mutation identification, driver-gene discovery, and tumor mutation scoring. The submitted full text is a different manuscript, so these are abstract-level claims that the supplied text does not document.","feed_headline":"ComBI: compressed bit-string tree targets 4–296× faster search","feed_subtitle":"The thesis reports 0.90 precision on billion-sample nearest-neighbor retrieval, plus gains on single-cell RNA-seq.","key_machinery":"The load-bearing objects are three. ComBI is a compressed BST of inverted hash tables, meaning each internal node stores an inverted index rather than a plain split; it is the mechanism claimed to cut memory and search time while keeping precision high. GRAF (and uGRAF) is a tree-ensemble classifier that combines global and local partitioning, bridging decision trees and boosting. CRCS is a deep embedding that maps codon switches to continuous vectors, letting mutations be scored without matched normal samples. Together they are claimed to form a single pipeline from hashing to classification to genomic prediction.","core_discovery":"On its own terms, the paper's central discovery is that a compressed binary search tree made of inverted hash tables (ComBI) can preserve the indexing benefits of space-partitioning hashing while avoiding the exponential growth and sparsity that make ordinary BST-based hashes inefficient on large data. The abstract further claims that this structure yields 0.90 precision at up to one billion samples, outrunning Multi-Index Hashing by 4–296× and Cellfishing.jl by 2–13× on single-cell RNA-seq searches. The paper also claims that a guided random forest (GRAF), an unsupervised variant (uGRAF), and a continuous representation of codon switches (CRCS) extend the same hashing-and-partitioning ideas to competitive classification across 115 datasets and to cancer genomics, including survival prediction in bladder, liver, and brain cancers.","pith_inferences":["The abstract does not state the recall level at which 0.90 precision is measured; a natural test is to re-run the comparison at matched recall values, since speed-ups can change sharply with operating point.","If the compression in ComBI is what drives the gains, the same inverted-hash compressed tree should transfer to metric nearest-neighbor search beyond bit-string spaces; that is an extension the abstract does not claim.","The submitted full text is a different paper, so the empirical numbers should be treated as unverified until the experiments appear; this is a reader-side caution, not a verdict on the methods.","GRAF and ComBI's claimed ability to estimate per-sample classifiability could be tested directly by comparing its ranking of patients against standard survival-risk scores on the same cohorts."],"forward_implications":["If ComBI's speed and precision hold, billion-sample similarity search could move from cluster-scale hashing to a single machine without much accuracy loss.","The 2–13× gain over Cellfishing.jl would make single-cell RNA-seq marker searches interactive on large cell atlases.","GRAF's reported accuracy on 115 datasets would make guided tree ensembles a competitive default for tabular classification.","Per-sample classifiability from ComBI/GRAF could let survival models be trained and evaluated on very large cancer cohorts without matched normal tissue.","CRCS, if valid, would expand somatic-mutation discovery to tumor samples where matched normals are unavailable."],"supporting_citations":[],"fun_headline_variants":["ComBI: 4–296× faster billion-scale nearest-neighbor search","0.90 precision at 1B samples with ComBI search","ComBI flips BST sparsity into 4–296× speed-ups","ComBI compresses bit-string trees for 1B-sample search"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The abstract's speed and precision numbers stand on the assumptions that the billion-sample benchmarks are representative, the baselines are configured competitively, precision is reported at a standard operating point, and the submitted full text actually contains these experiments; none of this can be checked from the supplied manuscript.","fun_headline_variants_meta":{"raw":{"variants":["ComBI: 4–296× faster billion-scale nearest-neighbor search","0.90 precision at 1B samples with ComBI search","ComBI flips BST sparsity into 4–296× speed-ups","ComBI compresses bit-string trees for 1B-sample search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001056,"raw_usage":{"total_tokens":4462,"prompt_tokens":1005,"completion_tokens":3457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":3378}},"tokens_in":621,"tokens_out":3457,"duration_ms":24748,"temperature":1.0,"reasoning_tokens":3378,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:09:38.815593+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run ComBI against Multi-Index Hashing and Cellfishing.jl on a publicly documented billion-vector dataset at matched recall levels; if the speed-ups fall below the claimed ranges or precision drops well below 0.90 at the standard operating point, the central claim fails.","supporting_citations":[],"review_version":2}