{"id":"1a6ee35a-aba5-4066-9eb5-26e2b88c2817","arxiv_id":"2506.00915","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A task-oriented review of skeleton-based action recognition that reorganizes known methods along a data processing pipeline and contains no new experimental result.","lead":"This preprint is a survey of 3D skeleton-based action recognition, organized by processing stages (modalities, augmentation, representation, feature extraction, modeling) instead of by model family. As submitted, its usefulness is undercut by large verbatim excerpts from other papers inside figure captions and by a promised datasets section that never appears.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Verbatim first-person text from cited papers in figure captions undermines the review's reliability, so the claimed comprehensive task-oriented roadmap is unsupported.","rationale":"The reader's verdict is REJECT with high confidence, and I agree. However, the reader's stated weakest assumption was the comparability of SOTA accuracy numbers in Tables 2 and 3; while that is a genuine concern, I find the more load-bearing threat to the central claim in the manuscript's verbatim reuse of first-person text from cited papers in figure captions and body text. The central claim is that this review provides a comprehensive, task-oriented, accessible roadmap; a necessary condition is that the contained descriptions are the authors' own reliable syntheses. The copied passages, such as Figure 5's caption containing the source paper's \"III. PROPOSED APPROACH\" section and Figure 8's caption containing \"our multi-stream CNN model,\" directly violate that condition. This is not a matter of style: it means that sections of the review are not original summaries and that the authors' own framework is not being cleanly presented. Even if the task-oriented taxonomy is conceptually sound, the submitted text does not support the claim that it is a trustworthy roadmap until these passages are rewritten, the duplicated paragraph in Section 6.1 is removed, and the promised dataset section is restored. The SOTA comparability issue is real and also merits attention, but it is secondary to the fundamental reliability defect. Hence I endorse the REJECT verdict while identifying a different load-bearing concern than the reader's formal weakest assumption.","tokens_in":54396,"tokens_out":6566,"duration_ms":69116,"concrete_test":"Run a systematic text-overlap check (e.g., iThenticate) of the manuscript against the references most likely to be the original sources of the flagged passages: the contrastive learning paper (Rao et al.), the shape-motion paper (Li et al.), and HD-GCN. If any figure caption or body paragraph contains the source paper's section headings, first-person phrases such as \"we propose\" or \"our model,\" or equations that are not accompanied by quotation marks or explicit attribution, the manuscript fails the reliability requirement. Separately, verify the actual section headings against the roadmap in Section 1: if Section 8 is titled \"Future Work\" rather than a datasets section, the promised content is structurally missing.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim of this review is that it provides a comprehensive, task-oriented framework and a trustworthy structured roadmap for 3D skeleton-based action recognition. A necessary condition for that claim is that the descriptions of methods, figures, and results are the authors' own careful syntheses of the literature. The submitted text violates this condition pervasively. Figure 5's caption contains a full section (\"III. PROPOSED APPROACH\") and first-person methodology text from a contrastive learning paper, including sentences like \"We do NOT require feature engineering...\" that are not the review authors' statements. Figure 8's caption contains the shape-motion paper's experimental section, including \"our multi-stream CNN model\" and implementation details. Figure 9's caption contains the HD-GCN paper's methodology and equations. These are not isolated slips: they show that the review's own descriptions are not reliable, original summaries, so a reader cannot trust the roadmap's contents. Additionally, Section 6.1 repeats the same ST-GCN/temporal-graph paragraph twice, and the introduction promises a datasets section as Section 8 while the delivered Section 8 is Future Work, contradicting the abstract's promise of a comprehensive dataset overview. This is a correctness/reliability threat to the central claim itself: if the text is compiled from other papers without genuine synthesis, the review does not provide the fundamental and accessible understanding it advertises.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of 3D skeleton-based action recognition. Its stated contribution is a task-oriented framework that organizes the field by pipeline stages—skeleton modalities, data augmentation, data representation, feature extraction, and spatio-temporal modeling—rather than by network architecture, and that additionally covers recent hybrid, Mamba, LLM, and generative approaches together with benchmark datasets and state-of-the-art accuracy tables. The central claim is that this organization provides a more intrinsic and complete roadmap than prior model-centric surveys. The paper contains no new derivations; its evidentiary basis is the completeness and accuracy of the literature coverage.","tokens_in":54499,"tokens_out":5263,"duration_ms":56275,"significance":"A well-executed survey with this task-oriented organization would be useful: it would complement architecture-centric surveys, highlight preprocessing and representation choices that are often underweighted, and bring recent LLM/Mamba/generative work into a unified discussion. The authors also deserve credit for attempting broad coverage across many methods and datasets and for including explicit equations for the four common skeleton modalities. However, the value of a survey depends on the reliability and original synthesis of its descriptions. As submitted, the manuscript contains substantial unattributed source text inside figure captions, a promised datasets section that is absent, a duplicated paragraph in Section 6.1, and SOTA tables without stated comparability conditions. These problems directly undercut the central claim of a trustworthy, comprehensive, task-oriented roadmap. For this reason, the current version does not meet the standard for publication.","major_comments":[{"comment":"Several figure captions contain substantial verbatim text from the cited papers, including first-person methodology and experimental description. For example, the caption of Figure 5 reproduces an entire section titled \"III. PROPOSED APPROACH\" from a contrastive-learning paper, including the sentence \"We do NOT require feature engineering like [5], [6], [28] or designing task-specific models...\" and that paper's internal figure reference \"Take Figure 5: Visualization of normal augmentations.\" Similarly, the caption of Figure 8 includes a full experimental section with \"our multi-stream CNN model\" and implementation details from the shape-motion paper, and the caption of Figure 9 includes the HD-GCN paper's methodology and equations. These are not the authors' own syntheses, and they are not presented as quotations. Because the review's central claim is that it provides a reliable, accessible structured roadmap, this pervasive source-text contamination undermines the trustworthiness of the review's descriptions and therefore its central claim.","section":"Figure 5 caption (also Figures 6, 8, 9)"},{"comment":"The Introduction states that \"Section 8 lists commonly used datasets and the performance of state-of-the-art models on these datasets,\" and the abstract promises \"a comprehensive overview of public 3D skeleton datasets.\" In the delivered manuscript, Section 8 is actually titled \"Future Work,\" while the dataset descriptions and Tables 1–3 are placed at the end of Section 7, after the Mamba paragraph, without a section heading. A core component promised in the abstract and introduction is thus missing from its announced location, which contradicts the paper's own structure and weakens the claimed comprehensiveness.","section":"Section 1, Section 8"},{"comment":"The same content is presented twice in near-verbatim form. The paragraph beginning \"Graph Convolutional Networks (GCNs) have been shown to excel in modeling skeleton space...\" is followed, after a short intervening paragraph, by a second paragraph beginning \"Graph Convolutional Networks (GCNs) are highly effective in modeling skeleton structures...\" that repeats the ST-GCN, temporal graph router, and Shift-GCN descriptions. This duplication is a concrete editorial defect that further reduces confidence in the accuracy of the review's content.","section":"Section 6.1, \"Temporal Feature Extraction Using Temporal Graphs\""},{"comment":"The state-of-the-art comparison is not supported by stated comparability conditions. Tables 2 and 3 report accuracy numbers from different papers without specifying the selection criteria for entries or verifying that the compared numbers were obtained under the same evaluation protocols, input streams, ensemble settings, and preprocessing choices. On NTU RGB+D and Kinetics benchmarks, such details are known to differ across papers and can change rankings. Without explicit caveats or a consistent protocol audit, the comparative conclusions drawn from these tables are not reliable. Since the abstract advertises \"an analysis of state-of-the-art algorithms evaluated on these benchmarks,\" this is a load-bearing issue.","section":"Tables 2 and 3"},{"comment":"The paper categorizes spatio-temporal modeling into serial, parallel, and fusion structures, but gives no operational rule for assigning a method to one of the three categories. Some methods described in Section 7.3, such as MS-G3D's \"unified spatiotemporal graph convolution module\" and STSF-GCN's SlowFast-style dual pathways, appear to blend or transcend these categories. If the taxonomy is intended to be a core part of the proposed task-oriented framework, the criteria for the trichotomy and its exhaustiveness need to be stated explicitly.","section":"Section 7 and Figure 11"}],"minor_comments":[{"comment":"The caption contains an IEEE Xplore watermark line, \"Authorized licensed use limited to: SUN YAT-SEN UNIVERSITY...\", which should be removed.","section":"Figure 4 caption"},{"comment":"The header \"Other Msethods\" should read \"Other Methods,\" and the row \"SLnL-rFA [227] 2019 36.6\" appears twice.","section":"Table 2"},{"comment":"In the CTR-GCN row, the NTU RGB+D 120 Cross-Setup entry is truncated as \"90.\"; a digit or digits appear to be missing.","section":"Table 3"},{"comment":"Several mathematical symbols render as placeholder glyphs (for example, the script symbols in the definitions of the augmentation and representation functions), making the formalization difficult to read.","section":"Section 2, Eqs. (1)–(13)"},{"comment":"The token-based representation panel contains the stray text \"???6,7\", which appears to be a leftover annotation from a source figure.","section":"Figure 7 (d)"}],"recommendation":"reject","confidential_remarks":"This is a review paper, so its central claim rests on the reliability and completeness of its coverage. The presence of substantial unattributed first-person source text in multiple figure captions is not a cosmetic issue; it indicates that the review's descriptions have not been independently verified and rewritten. Even if the text were cleaned up, a full audit of every figure caption and quoted passage would be required before the manuscript could be considered further. That goes beyond what a normal revision can reasonably address, and I recommend rejection in the current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe headline: this survey has a sensible organizing idea, but the submitted text is not trustworthy as a review. The task-oriented pipeline (modalities, augmentation, representation, extraction, modeling) is a reasonable way to structure the literature, and the coverage of Mamba, LLM, and generative methods is timely. If cleaned up, it could serve as a decent entry point for newcomers.\n\nWhat is actually new: the emphasis on preprocessing steps—modality derivation and augmentation—gets more space than in most prior surveys, and the four-modality schema (Joint/Bone/Joint-Motion/Bone-Motion) is neatly formalized in Section 2. The three broad spatio-temporal structures (serial/parallel/fusion) are a reasonable taxonomy. The SOTA tables are typical of the field.\n\nThe problems are serious. Several figure captions contain verbatim first-person text from the papers being reviewed: Figure 5 includes a full section from a contrastive learning paper, including 'We do NOT require feature engineering'; Figure 8 contains another paper's experimental section with 'our multi-stream CNN model'; Figure 9 carries the HD-GCN methodology. Figure 4 even retains an IEEE Xplore download watermark. This is not a minor editing slip. It means the authors' descriptions are not consistently their own, and a reader cannot rely on the roadmap's content. There is also a duplicated paragraph in Section 6.1 (the ST-GCN/temporal-graph passage is repeated verbatim), and the introduction promises a datasets section as Section 8, but the actual Section 8 is Future Work—the promised comprehensive dataset overview is not delivered in the place advertised.\n\nThe framework itself is not circular and does not depend on any derived claims; the authors cite their own prior survey [51] transparently, and the overlap is expected. But the textual integrity issues undermine the central claim of a comprehensive, reliable review. This needs a full cleanup: rewrite the captions, remove borrowed text, fix the duplication, restore the dataset section, and remove the watermark. Until then, the paper cannot function as a trustworthy reference.\n\nWho this is for: a reader who wants a task-structured map of the field might find value after these fixes. As submitted, I wouldn't cite it.\n\nRecommendation: desk reject in current form; if the authors resubmit a cleaned version, send it to peer review.\n\nBest,","headline":"A useful task-oriented organizing idea, but the manuscript's copied figure captions and structural errors make it untrustworthy as a review.","tokens_in":55175,"tokens_out":2727,"would_cite":false,"duration_ms":27310,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that 3D skeleton-based action recognition is best organized as a pipeline of task stages—modality derivation, augmentation, representation, feature extraction, and spatio-temporal modeling—rather than by model…","keywords":["action recognition","3D skeleton","task-oriented review","skeleton data augmentation","skeleton representation","spatio-temporal modeling","state-of-the-art benchmarks","Mamba and LLM methods"],"falsifier":"Re-run the top entries of Tables 2 and 3 under one fixed protocol—the same NTU RGB+D 120 split, the same joint-and-bone input streams, and no ensemble averaging—and check whether the reported ranking and margins survive; any large drop in a table-topping figure would falsify the comparative state-of-the-art claims. The completeness of the taxonomy could also be falsified by exhibiting a published method whose spatio-temporal coupling fits none of the serial, parallel, or fusion structures.","tokens_in":54041,"feed_emoji":"🦴","tokens_out":7601,"duration_ms":65931,"temperature":0.7,"pith_summary":"This review's central claim is that 3D skeleton-based action recognition has been surveyed the wrong way around: organizing the literature by model architecture—RNNs, CNNs, GCNs, Transformers—hides the task's essential structure. The authors propose a task-oriented framework that decomposes the field into pipeline stages—derived modalities, data augmentation, data representation, feature extraction, and spatio-temporal modeling—and argue that preprocessing steps deserve the same attention as network design. If the framework is right, it gives researchers a more intrinsic roadmap: each stage can be studied, reported, and optimized independently, and emerging approaches such as hybrid networks, Mamba, large language models, and generative models fit naturally into the same stages. The review also gathers public datasets and state-of-the-art benchmark numbers as evidence for how the staged pipeline plays out in practice. A sympathetic reader would take it as the claim that how you organize a survey shapes what questions the field asks, and that the pipeline-stage organization is the more fundamental one.","feed_headline":"Skeleton action surveys should follow the task, not the model","feed_subtitle":"A staged view—modalities, augmentation, representation, modeling—shows what architecture-based reviews leave out.","key_machinery":"The organizing machinery is the pipeline-stage framework itself (Figure 1 of the paper), formalized with a small set of equations. Raw input is a skeleton sequence $S = \\{X_t \\mid t=1,\\dots,T\\}$ with $X_t \\in \\mathbb{R}^{J\\times D}$; the derived modalities are defined by coordinate differences: $X^{\\text{bone}}_t = X^{\\text{joint1}}_t - X^{\\text{joint2}}_t$, $X^{\\text{jm}}_t = X^{\\text{joint}}_t - X^{\\text{joint}}_{t-1}$, and $X^{\\text{bm}}_t = X^{\\text{bone}}_t - X^{\\text{bone}}_{t-1}$. The four representations map the sequence onto the input formats of the four network families, and spatio-temporal modeling is reduced to three structural forms—serial ($F = \\Phi_{\\text{temporal}}(\\Phi_{\\text{spatial}}(\\mathcal{S}))$), parallel ($F = \\Phi_{\\text{temporal}}(\\mathcal{S}) + \\Phi_{\\text{spatial}}(\\mathcal{S})$), and fusion ($F = \\text{Fusion}(\\Phi_{\\text{temporal}}(\\mathcal{S}), \\Phi_{\\text{spatial}}(\\mathcal{S}))$)—which together let every reviewed method be placed at a stage and in a structural slot. This taxonomy is what carries the argument that the field can be surveyed completely without classifying by backbone architecture.","core_discovery":"The paper's central claim is that a task-oriented decomposition is a more essential and more complete way to understand 3D skeleton-based action recognition than the architecture-based classification used by previous surveys. It organizes the field along the actual processing chain: deriving the four generalized skeleton modalities (Joint, Bone, Joint-Motion, Bone-Motion) from raw coordinates, augmenting skeleton sequences with normal, extreme, mixing, and viewpoint-invariant transformations, converting the unstructured data into sequential, pseudo-image, graph, or token representations, extracting spatial and temporal features, and finally co-modeling space and time through serial, parallel, or fusion structures. The paper holds that preprocessing and representation stages strongly influence final accuracy, so reviews that skip straight to modeling miss a large part of what determines performance. It further claims that recent advances—hybrid architectures, Mamba-based state-space models, LLM-assisted feature extraction, and generative pre-training—are best understood within the same staged workflow rather than as new architecture families. The benchmark tables and dataset summaries are presented as the evidence that the staged pipeline, applied across the field, is where the measured progress lives.","pith_inferences":["The framework implicitly predicts a measurable claim the paper does not test: on a fixed benchmark and protocol, variation in modality design and augmentation strategy accounts for a comparable share of accuracy spread as backbone choice; a controlled study holding the network fixed while varying preprocessing could test it.","If the benchmark numbers in Tables 2 and 3 are not protocol-comparable, the review's SOTA rankings are fragile; an unstated corollary is that the field would benefit from a shared evaluation harness that standardizes splits, input streams, and ensemble policy.","The serial/parallel/fusion trichotomy resembles a general design pattern that could be exported beyond skeletons, for instance to other structured time-series recognition tasks, where the same three ways of coupling spatial and temporal modules recur.","The taxonomy suggests a natural organization for a future leaderboard or benchmark suite, with accuracy reported per stage configuration rather than per model name."],"forward_implications":["If the framework is right, reporting norms should change: papers would describe modality sets and augmentation schedules as first-class methodological choices on par with network depth, because both stages are claimed to shape accuracy.","Researchers new to the field get a staged roadmap; each stage can be optimized and evaluated in isolation before integration, which the review frames as the intrinsic structure of the task.","The framework absorbs the newest architectures—hybrid networks, Mamba state-space models, LLM prompting, and generative pre-training—as instances of assisted feature extraction or spatio-temporal co-modeling rather than as a new category of model.","The benchmark analysis identifies NTU RGB+D 120 as the most demanding current protocol, so progress claims in the field should be judged against it, not merely against NTU 60 or Kinetics numbers.","The completeness claim implies that any future method can be located in the framework; the survey thereby positions itself as the reference map for subsequent work."],"supporting_citations":[{"why":"The comprehensive model-oriented survey of 3D skeleton-based action recognition that the paper explicitly positions against; it supplies the prior organization being replaced.","marker":"[51]"},{"why":"The multi-modality action recognition review that covers skeleton methods only briefly, grounding the paper's claim that existing reviews under-treat skeleton data.","marker":"[52]"},{"why":"The Transformer-focused skeleton review, cited as an example of architecture-limited coverage that misses hybrid, Mamba, and LLM approaches.","marker":"[41]"},{"why":"ST-GCN, the foundational spatio-temporal graph model on which the fixed-topology discussion and the serial-structure GCN lineage rest.","marker":"[35]"},{"why":"The two-stream adaptive GCN that defines the source-to-target bone vector convention and adaptive adjacency matrices supporting the generalized-modalities account.","marker":"[66]"},{"why":"NTU RGB+D, the primary dataset behind the first set of state-of-the-art comparisons in Table 3.","marker":"[107]"},{"why":"NTU RGB+D 120, the extended benchmark the paper calls the most demanding current protocol.","marker":"[219]"},{"why":"Kinetics-400, the dataset behind the Table 2 comparison of skeleton-based methods.","marker":"[214]"},{"why":"The unified cross-spacetime graph convolution that anchors the spatio-temporal fusion-structure category.","marker":"[69]"}],"fun_headline_variants":["Task-first survey reshapes skeleton action recognition","From raw joints to LLMs: skeleton action pipeline reviewed","Preprocessing matters: task-based review of skeleton action","Mamba, LLMs, and more in task-oriented skeleton survey","Beyond models: a roadmap for skeleton action recognition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The state-of-the-art tables assume that accuracy figures taken from different papers were produced under the same evaluation protocol, the same input modalities, and the same training or ensembling conditions, so that the numbers can be compared as if they came from one experiment.","fun_headline_variants_meta":{"raw":{"variants":["Task-first survey reshapes skeleton action recognition","From raw joints to LLMs: skeleton action pipeline reviewed","Preprocessing matters: task-based review of skeleton action","Mamba, LLMs, and more in task-oriented skeleton survey","Beyond models: a roadmap for skeleton action recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1447,"prompt_tokens":998,"completion_tokens":449,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":372}},"tokens_in":614,"tokens_out":449,"duration_ms":4934,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:55:06.081074+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the top entries of Tables 2 and 3 under one fixed protocol—the same NTU RGB+D 120 split, the same joint-and-bone input streams, and no ensemble averaging—and check whether the reported ranking and margins survive; any large drop in a table-topping figure would falsify the comparative state-of-the-art claims. The completeness of the taxonomy could also be falsified by exhibiting a published method whose spatio-temporal coupling fits none of the serial, parallel, or fusion structures.","supporting_citations":[],"review_version":1}