{"id":"7a09e38b-e0ae-4375-a541-d95f81532263","arxiv_id":"2305.00578","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"NAC clusters high-dimensional data by scoring candidate partitions on a directed kNN graph with two permutation-standardized statistics that capture within-cluster enrichment and asymmetric separation.","lead":"The paper introduces NAC, a clustering algorithm that builds a directed k-nearest-neighbor graph and scores partitions with two permutation-standardized statistics to exploit asymmetric neighbor patterns caused by differing dispersion. A smart generalist might read it for a practical alternative to location-based clustering in high-dimensional settings such as genomics.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption correctly isolates the population-level justification as the key unverified element. Because the supplied text contains no contradictory derivation or simulation detail that would falsify the claim, the UNVERDICTED status is left unchanged.","tokens_in":1682,"tokens_out":201,"duration_ms":15642,"concrete_test":"Derive the limiting values of the weighted within-edge and contrast statistics under a two-component scale-mixture model (different variances, equal means) and check whether their standardized combination remains informative when the components are not separated in Euclidean distance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract describes a population-level analysis in which the directed kNN graph encodes separation via asymmetric neighbor counts and the two permutation-standardized statistics target complementary patterns. No internal inconsistency, hidden assumption, or unsupported step is visible in the given description of the objective or the claimed adaptation across regimes.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the NAC method for high-dimensional clustering. It constructs a directed k-nearest-neighbor graph from the data and evaluates candidate partitions via an objective that combines two permutation-standardized statistics: a weighted within-edge statistic capturing within-cluster enrichment and a contrast statistic capturing asymmetric separation. A population-level analysis is provided to show that these statistics target complementary nearest-neighbor patterns. Simulations across mean, scale, and combined location-scale separations, together with gene-expression applications, are used to demonstrate competitiveness under location separation and particular effectiveness when nearest-neighbor asymmetry is present.","tokens_in":1754,"tokens_out":375,"duration_ms":25148,"significance":"If the central claims hold, the work is significant because it supplies a model-free clustering principle that exploits dispersion-induced asymmetry in high dimensions without requiring a mixture model or low-dimensional embedding. The explicit population-level analysis demonstrating complementary targeting by the two statistics is a clear strength, as are the simulation studies that systematically vary separation regimes and the real-data illustrations. These elements together position NAC as a practical addition for small-sample, high-dimensional settings such as genomics. The permutation-standardization concern raised in the stress-test note does not appear to land as a load-bearing inconsistency on the basis of the provided description, because the standardization is presented as a device to isolate the asymmetry signal rather than as a data-dependent tuning step.","major_comments":[],"minor_comments":[{"comment":"The abstract packs the description of the two statistics and the objective into a single long sentence; splitting this material would improve readability without altering content.","section":null},{"comment":"Notation for the two standardized statistics (e.g., symbols for the weighted within-edge and contrast quantities) should be introduced once in the main text and used consistently thereafter.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and positive evaluation of our manuscript. The summary and significance assessment accurately reflect the paper's contributions. No major comments were raised in the report.","responses":[],"tokens_in":1278,"tokens_out":55,"duration_ms":5603,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that this paper introduces NAC for high-dimensional clustering by exploiting asymmetry in nearest-neighbor counts from dispersion differences, rather than the usual location or similarity assumptions. They construct a directed kNN graph and define two permutation-standardized statistics: a weighted within-edge statistic for within-cluster enrichment and a contrast statistic for asymmetric separation. These are combined in the objective to adapt to different regimes. The population analysis explains the complementary targeting, simulations across mean, scale, and combined differences show competitive or better performance, and gene-expression applications illustrate the use case in small-sample high-dim data. This is new relative to the geometric and local-similarity methods, and it does well by providing a clear principle and evidence from analysis and sims without requiring mixture models. On the soft spots, the fact that permutation standardization happens on the same data as the graph construction could lead to some circularity if k or weights are tuned on those realizations, though the standardization is intended to isolate the signal. The abstract claims support, but verifying the high-dim validity would require the full derivations. Nothing looks broken, just needs the details checked. This paper is for statisticians focused on clustering methods for genomics and high-dim data where standard assumptions fail. Readers interested in alternative separation principles would get value. It deserves a serious referee because the contribution is distinct and the supporting work appears solid enough to review.","headline":"NAC uses directed kNN asymmetry to cluster high-dim data where dispersion differences matter more than location.","tokens_in":2221,"tokens_out":338,"would_cite":false,"duration_ms":36746,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Statistical kNN-asymmetry clustering; no RS cost or ratio machinery","alignment":"orthogonal","rationale":"Paper's core (directed kNN graph, Z_w/Z_d permutation statistics, M_κ objective) is a nonparametric clustering criterion exploiting high-d dispersion asymmetry. It invokes no J-cost, reciprocal symmetry, φ-ladder, 8-tick periodicity, or parameter-free constant derivations. Domain (stat.ME) lies outside RS forcing chain; no theorem in the supplied corpus is engaged or contradicted.","tokens_in":54858,"confidence":"high","tokens_out":126,"duration_ms":5611,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Clustering high-dimensional data succeeds by scoring partitions on asymmetric neighbor counts in a directed k-nearest-neighbor graph.","keywords":["high-dimensional clustering","nearest-neighbor graph","asymmetric separation","dispersion differences","permutation standardization","directed graph clustering","scale separation","gene expression clustering"],"falsifier":"Generate data from two components that differ only in dispersion; if the NAC objective selects the true labels no more often than a random partition across repeated draws, the central claim fails.","tokens_in":2579,"feed_emoji":"","tokens_out":631,"duration_ms":11945,"temperature":0.7,"pith_summary":"The paper shows that groups in high dimensions can separate through differences in dispersion rather than location, producing one-way neighbor patterns that a directed graph records. NAC builds this graph once, then scores candidate partitions with two permutation-standardized quantities: a weighted count of within-cluster edges and a contrast that measures directed imbalance between clusters. Their combination adapts automatically to mean separation, scale separation, or both, without fitting a mixture model or projecting to low dimensions. Simulations confirm competitive performance on location differences and clear gains when asymmetry dominates; gene-expression examples illustrate the same behavior on real small-sample data.","feed_headline":"Asymmetric neighbors find clusters in high dimensions","feed_subtitle":"Directed kNN counts and two permutation tests adapt to location or dispersion separation without mixture models.","key_machinery":"The directed k-nearest-neighbor graph together with its two permutation-standardized statistics (weighted within-edge count and directed contrast) that are summed to form the clustering objective.","core_discovery":"The NAC objective, formed by adding a standardized within-edge enrichment statistic and a standardized contrast statistic on the directed kNN graph, recovers the correct partition by targeting complementary nearest-neighbor patterns that arise under both location and scale separation; population analysis shows each statistic isolates its target signal after permutation centering.","pith_inferences":["The permutation-based standardization may allow the same machinery to flag other forms of local heterogeneity beyond scale, such as differing tail behavior.","Because the graph is built once, NAC could serve as a fast post-processing step after any initial embedding that preserves local distances.","In regimes where multiple separation mechanisms coexist, the relative weight of the two statistics could be tuned by cross-validation on held-out edge counts."],"forward_implications":["NAC remains competitive with standard methods when clusters differ only in location.","NAC gains clear advantage precisely when nearest-neighbor asymmetry induced by scale differences is the main signal.","The method requires neither a parametric mixture model nor an explicit low-dimensional embedding.","The same two standardized statistics apply directly to small-sample high-dimensional problems such as gene-expression matrices."],"fun_headline_variants":["Asymmetric nearest neighbors cluster high-dimensional data","Directed kNN finds clusters via neighbor asymmetry","NAC combines kNN stats for location and scale separation","Neighbor asymmetry separates clusters in high dimensions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The directed kNN graph must carry the dominant separation signal in its asymmetric neighbor counts, and permutation standardization must remove finite-sample noise without needing Euclidean well-separation.","fun_headline_variants_meta":{"raw":{"variants":["Asymmetric nearest neighbors cluster high-dimensional data","Directed kNN finds clusters via neighbor asymmetry","NAC combines kNN stats for location and scale separation","Neighbor asymmetry separates clusters in high dimensions"]},"model":"grok-4.3","cost_usd":0.004504,"raw_usage":{"total_tokens":2216,"prompt_tokens":614,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":45037000,"prompt_tokens_details":{"text_tokens":614,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1548,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":614,"tokens_out":54,"duration_ms":8898,"temperature":1.0,"reasoning_tokens":1548,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-24T08:37:05.941029+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Generate data from two components that differ only in dispersion; if the NAC objective selects the true labels no more often than a random partition across repeated draws, the central claim fails.","supporting_citations":[],"review_version":1}