{"id":"1bd64d4c-0719-4edb-908d-23706bbd59c5","arxiv_id":"2508.14604","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A point cloud video action recognition model is claimed (UST-SSM), but the submitted full text is an unrelated energy disaggregation paper, so the central claim is unverifiable.","lead":"The title and abstract describe UST-SSM, a state space model for recognizing human actions in point cloud videos. The full text attached is a different paper about home energy disaggregation, so the method and its claimed results cannot be checked.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The supplied full text is a different paper (DualNILM); the UST-SSM central claim is unsupported by any evidence in the submission, so the verdict should remain UNVERDICTED.","rationale":"The reader's verdict is UNVERDICTED, with the rationale explicitly noting that the supplied full text is a different manuscript (DualNILM) and that the UST-SSM claims are supported only by the abstract. My independent read agrees: the strongest claim cannot be assessed because the evidence that would support it is absent from the submission. The reader's stated weakest_assumption focuses on the prompt-guided clustering in STSS as the most load-bearing technical premise. That is a reasonable forward-looking concern if the actual UST-SSM paper were present, but it is not the most immediate load-bearing concern. Before evaluating whether the clustering places useful correspondences within the SSM's receptive field, one must first establish that the manuscript actually describes STSS and the experiments. The submission does not. Therefore the decisive concern is the manuscript-text mismatch, which makes the central claim unverifiable. This is why agreement_with_reader is 'partial': the reader identified the mismatch in the rationale and final verdict, but their formal weakest_assumption points at a technical detail that is premature given the evidentiary gap. My recommendation to keep the verdict UNCHANGED (UNVERDICTED) is not a rejection of the paper's possible merits; it is an honest statement that the claim is currently unsupported by any in-scope evidence. The concrete test—checking the actual arXiv record—is cheap and would settle whether this is a misattachment (in which case the proper UST-SSM text must be reviewed) or a genuine absence of supporting content. No ad hominem is intended; the issue is structural, not personal. Following the reviewing rule that all manuscript text is in-scope evidence, the DualNILM full text itself demonstrates the incoherence of the submission as a vehicle for the UST-SSM claims, independent of any judgment about DualNILM's own quality.","tokens_in":43855,"tokens_out":1845,"duration_ms":23337,"concrete_test":"Retrieve the official arXiv record for 2508.14604 and compare its full text against the submitted abstract. If the record matches the provided DualNILM text, the submission is misidentified and no UST-SSM evidence exists. If the record instead contains the UST-SSM paper, verify that the STSS prompt-guided clustering mechanism and the reported accuracies on MSR-Action3D, NTU RGB+D, and Synthia 4D are actually present and reproducible. This single check determines whether any UST-SSM claim is auditable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that a selective SSM with Spatial-Temporal Selection Scanning, Spatio-Temporal Structure Aggregation, and Temporal Interaction Sampling achieves strong action recognition on point cloud videos—requires the actual UST-SSM manuscript to audit. The full text provided is arXiv:2508.14600v4, a NILM paper (DualNILM) with different title, authors, equations, and experiments. None of the UST-SSM components (STSS, STSA, TIS), the prompt-guided clustering mechanism, the 1D-scan reordering, or the reported results on MSR-Action3D, NTU RGB+D, and Synthia 4D appear anywhere in the submitted text. The reader's identified weakest assumption—that prompt-guided clustering places semantically similar but spatio-temporally distant points adjacent within the SSM's receptive field—is structurally correct but moot: the supplied manuscript does not even describe the clustering, let alone provide evidence for it. The load-bearing issue is thus evidentiary absence: the submission's content does not bear on the abstract's claims. This is a manuscript-integrity problem, not a technical flaw in the proposed method. If the real UST-SSM paper exists but was misattached, the claim is unverifiable from this submission; if the attached paper is what was actually submitted, then the UST-SSM claim is entirely unsupported. Either way, the verdict remains UNVERDICTED; there is no basis to assess correctness, novelty, or reproducibility.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission carries the title and abstract of a paper titled \"UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling,\" which claims a selective state space model for point cloud video action recognition with components STSS, STSA, TIS, prompt-guided clustering, and experiments on MSR-Action3D, NTU RGB+D, and Synthia 4D. However, the full text supplied is a different paper, \"Energy Injection Identification enabled Disaggregation with Deep Multi-Task Learning\" (DualNILM), a NILM paper with a Transformer-based multi-task architecture, equations for aggregate power with behind-the-meter injection, and experiments on laboratory/REDD/UK-DALE datasets. None of the UST-SSM components, the point cloud video formulation, the prompt-guided clustering mechanism, the 1D-scan reordering, or the claimed benchmarks appear anywhere in the supplied manuscript.","tokens_in":44019,"tokens_out":1968,"duration_ms":23815,"significance":"If the UST-SSM method as described in the abstract were fully developed and validated, it could be a meaningful contribution: extending selective SSMs to point cloud videos with linear complexity is an interesting direction, and the proposed components address a real difficulty (the spatio-temporal disorder of point clouds for unidirectional scanning). However, the submitted manuscript contains no technical exposition, no equations, no algorithm, no experimental protocol, no results, and no ablations for UST-SSM. The full text is entirely a separate NILM paper. Consequently, the significance of the claimed contribution cannot be assessed: there is no verifiable content to support the abstract's claims. No strengths such as machine-checked proofs, reproducible code, or parameter-free derivations for UST-SSM are present in the submission.","major_comments":[{"comment":"The central claim of the paper, as stated in the abstract, is that UST-SSM achieves strong action recognition on point cloud videos via Spatial-Temporal Selection Scanning, Spatio-Temporal Structure Aggregation, and Temporal Interaction Sampling. The full text, however, is a different paper titled 'Energy Injection Identification enabled Disaggregation with Deep Multi-Task Learning' (DualNILM). The full text contains no mention of STSS, STSA, TIS, prompt-guided clustering, 1D scanning, MSR-Action3D, NTU RGB+D, or Synthia 4D. The method, equations (e.g., Eqs. 1-20), experiments (e.g., Tables 3-17), and ablation studies all pertain to NILM. The submitted content therefore provides no evidence whatsoever for the abstract's claims. This is a load-bearing evidentiary absence: the manuscript does not contain the paper it purports to be, so the claimed contribution is entirely unsupported.","section":"Abstract vs. Full Text (whole manuscript)"},{"comment":"Even taking the abstract in isolation, the sentence 'Experimental results on the MSR-Action3D, NTU RGB+D, and Synthia 4D datasets validate the effectiveness of our method' is made without any accompanying protocol, statistics, baselines, or ablations. No numbers are reported, no evaluation metric is defined, and no comparison methods are listed. Because the full text does not elaborate on these experiments, the claim is not checkable. The reader cannot determine whether the reported validation would be significant, whether the method outperforms existing state of the art, or whether the experiments are conducted under standard protocols for these datasets.","section":"Abstract (experimental claim)"},{"comment":"The full text's architecture and experiments are entirely for DualNILM. Section 4 describes CNN encoders, Transformer encoders/decoders, and task-specific projections for appliance state recognition and energy injection disaggregation. Section 5 describes NILM datasets and PV simulation. None of these sections correspond to the UST-SSM components named in the abstract. Since the abstract's mechanism (prompt-guided clustering to reorganize unordered points into semantic-aware sequences) is the load-bearing idea that would justify the SSM's unidirectional scanning, its complete absence from the manuscript means the core technical proposal cannot be evaluated.","section":"Section 4 (architecture) and Section 5 (experiments)"}],"minor_comments":[{"comment":"The abstract states 'Our code is available at https://github.com/wangzy01/UST-SSM.' The full text does not mention this repository or provide any code listings. If the submission is intended to be the UST-SSM paper, the repository link should be supplemented with a version of the code or a detailed appendix; as submitted, the link is not verifiable.","section":"Abstract (code link)"}],"recommendation":"reject","confidential_remarks":"To the editor: This appears to be a case of a mismatched submission: the abstract describes UST-SSM, a point cloud video action recognition method, while the full text is an unrelated NILM paper (DualNILM). This is not a matter of missing details or presentation errors that could be fixed with a revision; the entire technical content, all equations, experiments, and references pertain to a different research area. I recommend rejection, possibly with an administrative note to the authors to verify the correct file was submitted. If a corrected UST-SSM manuscript exists, it would need to be submitted anew and undergo a fresh review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plain English: the submission is not a coherent paper. The abstract describes UST-SSM, a point cloud video action recognition method using selective state space models with prompt-guided clustering. The full text is a different manuscript: 'Energy Injection Identification enabled Disaggregation with Deep Multi-Task Learning' (DualNILM), arXiv:2508.14600v4, different authors and topic. So the UST-SSM claims rest on the abstract alone. There is no way to audit correctness, novelty, or reproducibility. Verdict: UNVERDICTED.\n\nWhat is actually new: the abstract's design idea—reorganizing unordered point cloud video into 1D sequences via prompt-guided clustering so the SSM's unidirectional scan can exploit spatio-temporally distant similarities—is a plausible research direction. If implemented cleanly, it could be a useful within-subfield contribution for efficient 4D action recognition. But the supplied text contains none of the machinery: no equations for STSS, STSA, or TIS, no clustering details, no experimental protocol or numbers for MSR-Action3D, NTU RGB+D, or Synthia 4D. So I can't confirm or refute it.\n\nWhat the attached DualNILM paper does well, in fairness: it's an internally coherent NILM paper with a clear problem formulation, a reasonable Transformer-based multi-task architecture, experiments on lab and synthetic datasets, ablations, and open data. It's just not the paper under review. Minor sloppiness: Appendix D.1 has placeholder strings where XGBoost hyperparameters should be.\n\nThe decisive problem is the mismatch itself. A serious editor should desk-reject this version and ask the authors to resubmit the correct PDF. If UST-SSM exists, it deserves a normal peer review; this artifact does not. I would not send this to reviewers or cite it.","headline":"The submission's full text is a different paper (DualNILM); the UST-SSM claims are unsupported, so this version should be returned, not peer-reviewed.","tokens_in":44670,"tokens_out":2480,"would_cite":false,"duration_ms":30720,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Point cloud video action recognition via semantic-aware state space scanning","keywords":["point cloud video","state space models","selective scanning","action recognition","prompt-guided clustering","spatio-temporal modeling","linear complexity"],"falsifier":"Run UST-SSM on any of the three datasets with STSS replaced by (a) frame-by-frame raw order and (b) a fixed random permutation of points; if accuracy stays about the same, the semantic ordering is not what drives the result. Also, evaluate on action classes not seen during training to check for prompt overfitting.","tokens_in":43591,"feed_emoji":"🎥","tokens_out":2650,"duration_ms":29454,"temperature":0.7,"pith_summary":"The paper tries to establish that selective state space models (SSMs), which handle long sequences at linear cost, can be made to work on point cloud videos. The obstacle is that point clouds are unordered in both space and time, so a unidirectional scan sees little structure. The paper's answer is to reorder points into semantic-aware sequences through prompt-guided clustering, then add modules that restore missing geometry and motion details and sharpen temporal links. If correct, action recognition on 4D point cloud video becomes a linear-complexity sequence modeling problem rather than a dense 3D one.","feed_headline":"SSMs read point cloud video after semantic reordering","feed_subtitle":"Prompt-guided clustering lines up distant similar points so a 1D scan can track human actions in 3D.","key_machinery":"Spatial-Temporal Selection Scanning (STSS) is the load-bearing component: it reorganizes unordered point cloud video frames into a 1D sequence through prompt-guided clustering, so that a selective SSM can treat the video as a sequence. STSA aggregates spatio-temporal features to compensate for missing 4D geometry and motion; TIS enhances temporal interaction using non-anchor frames and expanded receptive fields.","core_discovery":"The central claim is that the spatio-temporal disorder of point cloud videos—not the lack of temporal information—is what blocks SSMs from modeling them. UST-SSM removes that block with Spatial-Temporal Selection Scanning (STSS), which uses prompt-guided clustering to arrange unordered points so that similar points that are far apart in space or time become neighbors in the 1D scan. With that ordering, the SSM's unidirectional state can propagate information between those distant-but-similar points. Two supporting components, Spatio-Temporal Structure Aggregation (STSA) and Temporal Interaction Sampling (TIS), recover 4D geometric and motion detail and improve fine-grained temporal dependenc","pith_inferences":["If the prompts used for clustering are learned from action categories, the ordering may be biased toward training classes; testing on unseen actions would reveal whether the semantic ordering generalizes.","The same reordering idea could transfer to other unordered sequence problems (LiDAR sweeps, unordered graph node sets) that are currently fed to SSMs in raw order.","A head-to-head with transformer-based 4D action recognition on the same datasets would show whether the linear-cost claim costs accuracy."],"forward_implications":["Action recognition from point cloud video can be formulated as linear-complexity sequence modeling rather than dense 3D convolution.","Semantic reordering of points can make a unidirectional state machine reach spatially and temporally distant but similar points.","Compensating missing 4D details through aggregation improves recognition when geometry or motion is sparse.","Sampling non-anchor frames strengthens fine-grained temporal dependencies."],"supporting_citations":[],"fun_headline_variants":["Reordering point clouds unlocks SSMs for 3D video analysis","Semantic clustering fixes point cloud disorder for SSM video modeling","UST-SSM: SSMs work on point cloud videos after reordering","Distant similar points become neighbors to aid SSM in 3D video","Prompt-guided clustering enables SSMs to process point cloud videos"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The method's advantage depends on prompt-guided clustering really putting spatially and temporally distant but similar points next to each other in the scan; if the ordering is no better than a fixed or random order, the SSM gains nothing.","fun_headline_variants_meta":{"raw":{"variants":["Reordering point clouds unlocks SSMs for 3D video analysis","Semantic clustering fixes point cloud disorder for SSM video modeling","UST-SSM: SSMs work on point cloud videos after reordering","Distant similar points become neighbors to aid SSM in 3D video","Prompt-guided clustering enables SSMs to process point cloud videos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1136,"prompt_tokens":776,"completion_tokens":360,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":269}},"tokens_in":520,"tokens_out":360,"duration_ms":4959,"temperature":1.0,"reasoning_tokens":269,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:24:26.509097+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run UST-SSM on any of the three datasets with STSS replaced by (a) frame-by-frame raw order and (b) a fixed random permutation of points; if accuracy stays about the same, the semantic ordering is not what drives the result. Also, evaluate on action classes not seen during training to check for prompt overfitting.","supporting_citations":[],"review_version":1}