{"id":"5723527d-dc7b-4a5d-a6fc-c8cc8e7af37e","arxiv_id":"2603.05963","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"S2I turns skeleton sequences into body-part-arranged image-like tensors, enabling vision-pretrained models for self-supervised skeleton representation learning across heterogeneous formats.","lead":"The paper proposes Skeleton-to-Image Encoding (S2I), which rearranges 3D joint sequences into fixed-size image-like maps so large vision-pretrained models can learn skeleton features. If it works as claimed, action systems could reuse existing vision backbones without new skeleton-specific branches and handle mixed skeleton formats in one pipeline.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Manuscript-ID mismatch: full text is CR-QAT (2603.05964), not S2I (2603.05963); central claim cannot be stress-tested.","rationale":"The reader already flagged the manuscript mismatch and correctly issued UNVERDICTED with low confidence. My pass confirms the same hard blocker: the supplied full text is an entirely different work (CR-QAT). No technical claim of S2I can be examined, so no stronger or weaker verdict is warranted. The concrete test simply restores the correct source; until that is done the review cannot progress. Agreement with the reader is therefore complete, and the verdict stays UNVERDICTED.","tokens_in":14444,"tokens_out":465,"duration_ms":5601,"concrete_test":"Retrieve the true PDF/source of arXiv:2603.05963 (or the authors’ camera-ready). Confirm that its title, abstract, and experimental tables match the S2I abstract; if the retrieved document is still CR-QAT or otherwise unrelated, the present review remains invalid and must be discarded.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The review target is arXiv 2603.05963 (Skeleton-to-Image Encoding). The CACHEABLE PAPER SOURCE CONTEXT, however, supplies the complete manuscript of a different paper—CR-QAT for open-vocabulary object detection (arXiv 2603.05964v2). No section, equation, figure, table, or experiment of the S2I paper is present. Consequently the reader’s strongest claim (“S2I enables, for the first time, vision-pretrained models for self-supervised skeleton representation learning …”) and its weakest assumption (body-part semantic partitioning + resize preserves kinematics for RGB-pretrained models) rest solely on the abstract. Without the actual method description, ablations, or cross-format results, it is impossible to verify whether the encoding preserves temporal dynamics, whether the “first-time” claim holds against prior pose-as-image work, or whether the reported NTU/PKU gains are real. The load-bearing failure is therefore not an internal flaw of S2I but the absence of the paper itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The submission is titled and abstracted as Skeleton-to-Image Encoding (S2I), a method that partitions skeleton joints by body-part semantics, arranges them into image-like tensors, and resizes them so that vision-pretrained models can be used for self-supervised skeleton representation learning, with claimed gains on NTU-60/120 and PKU-MMD including cross-format settings. The body of the provided manuscript, however, is an entirely different paper (CR-QAT: Curriculum Relational Quantization-Aware Training for Open-Vocabulary Object Detection, arXiv 2603.05964). No method description, equations, figures, tables, or experimental results for S2I appear in the full text. Consequently the central claims of S2I cannot be examined beyond the abstract.","tokens_in":14655,"tokens_out":689,"duration_ms":14353,"significance":"If the abstract claims were substantiated, a unified image-like encoding that lets large vision-pretrained models transfer to heterogeneous skeleton data without new architectural branches would be a useful contribution to multi-modal action recognition and skeleton SSL, especially under data scarcity. Because the manuscript body does not contain the S2I method or any of its experiments, significance of the claimed result cannot be assessed from the supplied document.","major_comments":[{"comment":"Manuscript identity mismatch: the title, abstract, and paper_id (2603.05963) describe Skeleton-to-Image Encoding, yet every section from the Introduction through Conclusion, all figures, tables, equations, and references belong to CR-QAT (open-vocabulary detection quantization). No S2I encoding definition, training protocol, ablation, or NTU/PKU result is present. The central claim therefore cannot be verified or stress-tested.","section":null},{"comment":"Load-bearing premise of S2I (body-part semantic joint partitioning + resize preserves temporal/kinematic structure for RGB-pretrained models) is asserted only in the abstract and is unsupported by any analysis, ablation, or comparison in the supplied text. Without the actual method section this premise remains unexamined.","section":null},{"comment":"Claims of 'first-time' use of vision-pretrained models for skeleton SSL and of successful cross-format evaluation on NTU-60/120 and PKU-MMD cannot be checked: the experimental sections, baselines, and tables in the document address LVIS/COCO quantization of YOLO-World, not skeleton representation learning.","section":null}],"minor_comments":[{"comment":"Even if the correct S2I manuscript were substituted, the abstract alone leaves notation for the partitioning/arrangement operator and the precise self-supervised objective unspecified; those would need clear definitions in any resubmission.","section":null}],"recommendation":"uncertain","confidential_remarks":"The CACHEABLE PAPER SOURCE CONTEXT and the abstract belong to two different arXiv IDs (2603.05963 vs 2603.05964). This appears to be a data-pipeline or packaging error rather than an authorial flaw of either paper. I cannot produce a substantive technical review of S2I until the correct full manuscript is supplied. Recommend returning the submission for correction of the source document before any further refereeing."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The first thing to know: the review target is Skeleton-to-Image Encoding (2603.05963), but the manuscript block we were given is CR-QAT for open-vocabulary detection (2603.05964). So this is effectively an abstract-only read of S2I. I cannot verify method details, ablations, or the NTU/PKU numbers.\n\nFrom the abstract alone, the pitch is clear. S2I turns skeleton sequences into fixed-size image-like tensors by grouping joints by body-part semantics, then feeds them to vision-pretrained models for self-supervised skeleton learning. The second claim is format unification: one image-like layout for heterogeneous skeletons so you do not need a new branch per joint set. If that works, it is useful for action recognition and multi-modal setups that want to reuse large vision backbones without building skeleton-specific towers from scratch.\n\nWhat is actually new is narrower than the “for the first time” line. Pose-as-image / heatmap / pseudo-RGB pipelines are old. The interesting slice is the specific body-part layout plus vision-pretrained SSL plus cross-format evaluation. That could still be a solid methods contribution if the encoding preserves temporal and kinematic structure well enough for RGB-pretrained features to transfer. That premise is load-bearing and, here, only asserted.\n\nSoft spots, in proportion: (1) manuscript mismatch means soundness is uncheckable; (2) novelty framing is aggressive relative to the broader pose-as-image literature; (3) no inspectable baselines, protocols, or cross-format protocol details. Nothing in the abstract looks circular or incoherent—just under-evidenced from what we have.\n\nWho it is for: people doing skeleton SSL, multi-modal action, or anyone trying to avoid training large skeleton models from scratch. A serious editor would send a real full paper with those experiments to referees; I would not green-light peer review on this package as delivered. I would not cite or bring it to reading group until the correct PDF is in hand. If the real paper matches the abstract and the cross-format results hold, it is worth a careful look then—not before.","headline":"We only have the S2I abstract; the supplied “full text” is a different paper (CR-QAT), so the central claims cannot be checked.","tokens_in":15338,"tokens_out":551,"would_cite":false,"duration_ms":12277,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Turning skeleton joint sequences into image-like tensors lets vision-pretrained models do self-supervised skeleton representation learning for the first time.","keywords":["skeleton representation learning","Skeleton-to-Image Encoding","vision-pretrained models","self-supervised learning","action recognition","heterogeneous skeletons","cross-format evaluation"],"falsifier":"Replace the body-part-semantic layout with a random or purely geometric joint arrangement, retrain the same vision-pretrained backbone under identical self-supervised objectives, and check whether the large gains on NTU-60/120 and the cross-format setting disappear.","tokens_in":15282,"feed_emoji":"🦴","tokens_out":584,"duration_ms":16339,"temperature":0.7,"pith_summary":"Large vision models work well on images and multi-modal tasks, but cannot be applied directly to 3D human skeletons because the data formats differ and large skeleton datasets are scarce. This paper introduces Skeleton-to-Image Encoding (S2I), which partitions joints by body-part semantics, arranges them into a 2-D layout, and resizes the result to standard image dimensions. The resulting image-like sequences can be fed straight into existing vision-pretrained networks for self-supervised pretraining, transferring visual knowledge into the skeleton domain. The same encoding also supplies a single, uniform format that can absorb heterogeneous skeleton layouts from different sensors or datasets. Experiments on NTU-60, NTU-120 and PKU-MMD, including cross-format transfer, show that the approach yields competitive self-supervised skeleton representations without any skeleton-specific architecture or large-scale skeleton pretraining corpus.","feed_headline":"Skeletons become images so vision models can learn them","feed_subtitle":"A body-part layout lets large pretrained vision nets do self-supervised skeleton learning without extra branches","key_machinery":"Skeleton-to-Image Encoding (S2I): the body-part-semantic joint partitioning and subsequent resize-to-image step that converts raw skeleton sequences into inputs consumable by ordinary vision backbones.","core_discovery":"A simple, deterministic encoding that maps any skeleton sequence into a fixed-size image-like tensor—by grouping joints according to body-part semantics and resizing—is sufficient to unlock large vision-pretrained models for self-supervised skeleton representation learning and to unify previously incompatible skeleton formats under one pipeline.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Body-part maps turn skeletons into images for vision pretrains","Skeleton sequences become fixed images so vision models learn them","Joint grouping and resize unlock vision nets on skeleton data","S2I encoding unifies skeleton formats for pretrained vision SSL","Any skeleton maps to image tensors enabling vision self-supervision"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That arranging joints by body-part semantics and stretching the result into an ordinary image still keeps the temporal and kinematic cues that action recognition needs, even though the models were pretrained only on natural photographs.","fun_headline_variants_meta":{"raw":{"variants":["Body-part maps turn skeletons into images for vision pretrains","Skeleton sequences become fixed images so vision models learn them","Joint grouping and resize unlock vision nets on skeleton data","S2I encoding unifies skeleton formats for pretrained vision SSL","Any skeleton maps to image tensors enabling vision self-supervision"]},"model":"grok-4.5","effort":"low","cost_usd":0.006512,"raw_usage":{"total_tokens":1634,"prompt_tokens":775,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":65120000,"prompt_tokens_details":{"text_tokens":775,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":775,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":775,"tokens_out":84,"duration_ms":7962,"temperature":1.0,"reasoning_tokens":775,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T14:06:05.138811+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Replace the body-part-semantic layout with a random or purely geometric joint arrangement, retrain the same vision-pretrained backbone under identical self-supervised objectives, and check whether the large gains on NTU-60/120 and the cross-format setting disappear.","supporting_citations":[],"review_version":1}