{"id":"351bc010-b076-4b41-9428-f00cf90ca9be","arxiv_id":"2504.18977","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A low-parameter, biologically inspired 3D pyramidal network with a linear-SVM feature-fusion extension matches or exceeds several prior video-classification baselines on Weizmann, KTH, YUPENN, and Maryland datasets.","lead":"This paper proposes a 3D pyramidal neural network, called 3DPyraNet, and a feature-fusion variant that sends learned video features to a linear SVM. The authors report strong accuracy on human-action and dynamic-scene benchmarks while using far fewer parameters than deep 3D CNNs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SOTA claim is underdetermined because evaluation protocols are non-equivalent and the paper's own baseline numbers conflict across Tables 2-4.","rationale":"The reader's verdict is REJECT, and I agree with the core objection: the empirical claims cannot be verified because the experimental protocol is under-specified and the comparisons are made against non-equivalent baselines. I add one concrete internal inconsistency that strengthens the concern: the same baseline, C3D on Maryland, is reported as 78 in Table 2 and 87.7 in Table 4, and the claimed state-of-the-art margin is computed from the higher number. Similarly, YUPENN baseline values differ between Tables 3 and 4. This is not a novelty objection; the 3D pyramidal partial-weight-sharing architecture is a plausible research direction, and the parameter-count advantage is credible. The problem is that the central claim is empirical, and the supporting evidence is self-contradictory and insufficient. A revision that releases code, specifies exact splits, reports error bars, and reconciles Tables 2-4 could make the claim testable, but as submitted the central claim is not supported. I therefore keep the reader's reject verdict.","tokens_in":20746,"tokens_out":5978,"duration_ms":61376,"concrete_test":"Run 3DPyraNet-F under a standard published protocol, e.g. the original KTH split with 16 training and 9 test persons, report per-class means over five runs with standard deviations, and compare against the corrected baseline values obtained by recomputing overall accuracies from the per-class rows in Tables 2 and 3. If the KTH accuracy falls below 93.42, or if reconciling Tables 2-4 changes the 7.17% and 5.33% margins, the central claim requires revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the claim of outperforming state-of-the-art on three benchmarks, the accuracy numbers in Table 4 must measure the same quantity as the baselines. The paper explicitly disclaims that equivalence. Section 5.2 says KTH results use a random set of samples from the full KTH datasets with only background subtraction, whereas the cited 3D-ConvNet baseline split KTH into KTH1/KTH2 by complexity and used a voting scheme. Section 5.3.2 concedes that Feichtenhofer's better result is based on majority voting while we did individual clip classification. No class-wise splits, number of runs, or standard deviations are given for Weizmann, KTH, YUPENN, or Maryland, so the claimed margins are not reproducible. The internal numbers also contradict each other: Table 2 gives C3D Maryland overall as 78 while Table 4 gives C3D Maryland as 87.7 and uses that value for the reported 7.17% gain; Table 3 gives C3D YUPENN overall as 97 while Table 4 gives 98.1; Section 5.3.2 calls 94.83% state-of-the-art on Maryland while Table 2 shows 95 as the overall value. These conflicts are load-bearing because the SOTA margins are computed directly from those baseline numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes 3DPyraNet, a three-hidden-layer 3D pyramidal neural network whose partially shared position-oriented weight scheme is intended to learn spatio-temporal features from video, and 3DPyraNet-F, which fuses the last-layer features and classifies them with a linear SVM. Experiments are reported on Weizmann, KTH, YUPENN, and Maryland for action and dynamic scene recognition. The paper reports 98.99% on Weizmann, 93.42% on KTH, 93.67% on YUPENN, and 94.83% on Maryland for 3DPyraNet-F, with 0.83M parameters versus 17.5M for C3D, and claims state-of-the-art performance on three of the four datasets. Back-propagation equations for the 3D pyramidal, pooling, and fully connected layers are given in Section 3.3.","tokens_in":21020,"tokens_out":6851,"duration_ms":60099,"significance":"The architectural idea of a biologically motivated pyramidal 3D weight-sharing scheme is coherent, and the parameter budget (0.83M) relative to C3D is attractive if the accuracies can be reproduced. The derivation of the back-propagation equations in Section 3.3 is one of the more concrete parts of the paper. However, the empirical contribution is the main selling point, and it is currently underdetermined by the evaluation protocol: split definitions, sampling procedures, standard deviations, and clip-level versus video-level metrics are missing, and several baseline numbers in Tables 2-4 conflict. The SOTA margins are therefore not reproducible as stated.","major_comments":[{"comment":"The KTH evaluation is not protocol-equivalent to the cited baselines. The text states that the model was tested on a 'random set of samples from full KTH datasets' with only background subtraction, whereas Baccouche et al.'s 3D-ConvNet baseline is evaluated on the KTH1/KTH2 complexity split with a voting scheme, and Ji et al.'s 3DCNN uses ROI sequences extracted by a tracker. The reported 93.42% for 3DPyraNet-F and the 89.40%/90.2% baseline values therefore do not measure the same quantity. Please specify the exact split, class-wise composition, number of runs, standard deviations, and whether results are clip-level or video-level; without these details the KTH rows of Table 4 are not comparable.","section":"Section 5.2, Table 4 (KTH)"},{"comment":"The Maryland state-of-the-art claim is internally inconsistent. The text says 3DPyraNet-F reaches 94.83% and outperforms C3D by 7.17%; that margin is computed from the 87.7% C3D figure in Table 4, but Table 2 reports C3D's Maryland overall accuracy as 78 and 3DPyraNet-F's as 95. The same inconsistency occurs for YUPENN, where Table 3 gives C3D overall 97 and Table 4 gives 98.1. Please reconcile the baseline values and recompute all claimed gains, or state explicitly which C3D configuration each table refers to.","section":"Section 5.3.2 and Tables 2/4 (Maryland)"},{"comment":"The Weizmann protocol is not specified. The paper reports a mean accuracy over five splits but does not define the splits, the number of training/testing sequences per class, or the standard deviation. It also states that the 92.46% 'all-1' result is shown as 3DPyraNet(all-1) in Table 4, but Table 4 contains no such entry. Please provide the split definition, per-split results, and a corrected table or reference.","section":"Section 5.2 and Table 4 (Weizmann)"},{"comment":"The margin statements for YUPENN are mutually inconsistent. The text says the model falls short of state-of-the-art by 5.33% and later says Feichtenhofer et al. (2014) is better by 1.5%, while Table 4 lists 99 and 96.2 for the two Feichtenhofer entries and 93.67 for 3DPyraNet-F, corresponding to differences of 5.33 and 2.53 percentage points. Please correct the text and table so that the claimed margins are reproducible.","section":"Section 5.3.2 (YUPENN comparisons)"}],"minor_comments":[{"comment":"The notation table is referred to as 'Table. ??' in two places, but no notation table appears; please add it or remove the references.","section":"Section 3.2"},{"comment":"The abbreviations 'AR' and 'SR' are used, with 'SR' presumably meaning dynamic scene recognition (DSR); this abbreviation is never defined.","section":"Section 5.1"},{"comment":"The model name appears as '3DPyraNet-FM' in Table 4 and as '3DPyraNet-F M' in the text; please use one consistent name throughout.","section":"Table 4 and Section 4.2"},{"comment":"The sentence 'it overcame reported best result by 3DConvNet model, i.e. 88.26% with an average of 91.07% from ten tests' is ambiguous about which model produced the 91.07% average; please rewrite for clarity.","section":"Section 5.2"},{"comment":"The symbol 'D' is used both for temporal depth and as the upper limit in summations; a consistent table of symbols would help the reader.","section":"Equations 1-2"},{"comment":"There are typographical errors such as 'becuase', 'hihg', and inconsistent capitalization; a careful proofread is needed.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is close to a first submission in format and would benefit from a cycle of major revision. If the authors cannot supply the split definitions and reconcile Tables 2-4, the state-of-the-art claims should be withdrawn. I lean toward major revision rather than rejection because the underlying architecture and derivations are worth evaluating under a properly specified protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the architecture is real and worth a look, but the headline numbers don't survive contact with the paper's own tables. The actual contribution is a 3D pyramidal network with partially shared position-oriented kernels, plus a feature-fusion variant that feeds the top-layer features to a linear SVM. That is honest, non-paradigm-shifting work, and the parameter count (about 0.83M versus C3D's 17.5M) is a genuine selling point for embedded or memory-limited settings. The backpropagation equations look coherent, and the link to earlier pyramidal networks (Cantoni-Petrosino, Phung-Bouzerdoum) is properly acknowledged. Self-citation is not the problem here.\n\nThe problem is the evidence. The stress-test note is correct. The same C3D baseline appears as 78 on Maryland in Table 2 but as 87.7 in Table 4, and as 97 on YUPENN in Table 3 but as 98.1 in Table 4. The paper uses the Table 4 numbers to compute a 7.17% SOTA gain on Maryland, yet Table 2 already shows 3DPyraNet-F at 95 overall. Section 5.3.2 calls 94.83% state-of-the-art while its own per-class table shows 95. Something is wrong in the aggregation, and because the SOTA margins are computed directly from these baseline numbers, the internal contradictions are load-bearing.\n\nThe protocol details are also missing. For KTH, Section 5.2 explicitly says the authors used a random set of samples from the full dataset with only background subtraction, while the cited 3D-ConvNet baseline split KTH into KTH1/KTH2 by complexity and used voting. The paper itself concedes this non-equivalence. For Weizmann, five splits are mentioned but not defined. No class-wise splits, number of runs, or standard deviations are given for any dataset. These are not minor omissions; they make the accuracy tables unverifiable.\n\nWhat the paper does well is present a clear architectural alternative with a principled weighting scheme and a parameter-efficiency story. The per-class results on Maryland are interesting. But as written, the state-of-the-art claim is underdetermined.\n\nMy recommendation: do not desk-reject, but do not accept either. Send it to a referee who knows the KTH and Weizmann protocol literature, with a request for major revision: fix the baseline tables, document the exact splits and evaluation level (clip versus video), and either provide code or enough detail to reproduce the numbers. The architecture deserves referee time; the current empirical write-up does not.","headline":"A legitimate compact 3D pyramidal architecture, but the reported state-of-the-art results are undercut by conflicting baseline numbers and non-equivalent evaluation protocols.","tokens_in":21525,"tokens_out":2886,"would_cite":false,"duration_ms":31548,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A four-layer 3D pyramidal network with position-oriented, partially shared weights is claimed to learn video features that rival or beat far larger 3D convolutional networks, reaching 94.83% on Maryland and 98.99% on Weizmann with only…","keywords":["3D pyramidal neural network","spatio-temporal feature learning","human action recognition","dynamic scene recognition","feature fusion","linear SVM","video classification","parameter-efficient deep learning"],"falsifier":"Run 3DPyraNet-F on KTH and Weizmann under the exact standard protocols used by the cited baselines, using fixed public splits, multiple runs, and per-class standard deviations. If the mean accuracy falls below the reported 93.42% on KTH or 98.99% on Weizmann by more than a couple of points, the state-of-the-art comparison is not robust to split choice.","tokens_in":20564,"feed_emoji":"🎥","tokens_out":8541,"duration_ms":79446,"temperature":0.7,"pith_summary":"This paper tries to establish that a deliberately shallow, biologically inspired pyramid architecture can learn video features that compete with much larger deep 3D convolutional networks. It proposes 3DPyraNet, a 3D correlation network with partially shared, position-oriented weights that process several adjacent frames, and 3DPyraNet-F, a version that fuses the top-layer feature maps and classifies them with a linear SVM. The reported result is that 3DPyraNet-F reaches 98.99% on Weizmann, 93.42% on KTH, 93.67% on YUPENN, and 94.83% on Maryland, using roughly 0.83M parameters where C3D uses 17.5M. If true, the claim matters because it would show that network shape and weight-sharing philosophy, not depth and width alone, can drive accuracy on small labeled video sets, and that a small model can be deployed where memory is scarce.","feed_headline":"A 0.83M-parameter video net beats a 17.5M-parameter C3D","feed_subtitle":"Three-dimensional pyramidal weighting plus feature fusion gives 94.83% on Maryland in a 0.83M-parameter model.","key_machinery":"The load-bearing object is the 3D partially shared weight matrix: for every output neuron, three weight slices and three consecutive input frames are correlated through a receptive field of size RF, with the weight matrix the same size as the input so that the kernel is position-specific and locally shared only to the degree set by an overlap O. This weighting scheme does the work of learning spatial layout and temporal correspondence at once; 3D temporal max pooling then selects maxima across three fields, and the final normalized feature maps are concatenated into one vector that a linear SVM classifies.","core_discovery":"On the paper's own terms, the central discovery is that replacing the sliding convolutional kernel with a 3D, partially shared, position-oriented weight matrix—each output neuron gets a unique local 3D kernel assembled from three consecutive frames—lets a four-layer network learn discriminative spatio-temporal features from raw or lightly preprocessed frames. Training that base model and then extracting and fusing the final normalized feature maps into a single vector for a linear SVM yields 3DPyraNet-F. The paper reports that this fusion step raises accuracy on Weizmann from 90.9% to 98.99%, and on Maryland from 67% to 94.83%, overtaking C3D by 7.17 percentage points there; on KTH and YUPENN it reports results comparable or close to recent deep baselines while keeping parameters below one million.","pith_inferences":["If the protocol issue were resolved, a natural extension would be to train 3DPyraNet-F on a large-scale video corpus; the paper itself suggests this only as future work, and the outcome is not known.","The gap between global fusion (3DPyraNet-F) and local mean fusion (3DPyraNet-F_M) on Weizmann and KTH hints that fusion granularity is a tunable axis; sweeping fusion over intermediate layers rather than only the top layer could lift YUPENN, where the model lags the best baseline by 5.33%.","The position-oriented weight matrix may be interpretable by construction: visualizing the trained weight matrices, as the paper reports for the first layer, could provide a direct test of which spatial positions and temporal offsets the model relies on for each class."],"forward_implications":["A shallow 3D pyramidal model can match or beat much deeper 3D convolutional networks on small action and scene datasets, so depth is not the only route to discriminative video features.","Fusing the highest-layer learned features and classifying with a linear SVM is itself a performance lever: the paper reports Weizmann rising from 90.9% to 98.99% with this addition.","The roughly 0.83M-parameter model, compared with C3D's 17.5M, implies video recognition models can run in embedded or memory-constrained settings if the accuracy carries over.","The strong Maryland result suggests position-oriented temporal correlation is especially suited to camera-induced motion, where the scene content shifts between frames."],"supporting_citations":[{"why":"Supplies the 3D-ConvNet baseline and evaluation setup on Weizmann and KTH whose scores the paper compares against and partly exceeds.","marker":"Baccouche et al. [2011]"},{"why":"Establishes 3D CNNs for action recognition and provides the KTH comparison point the paper seeks to approach with fewer layers.","marker":"Ji et al. [2013]"},{"why":"Original pyramidal neural network structure that 3DPyraNet explicitly extends to three dimensions.","marker":"Cantoni and Petrosino [2002]"},{"why":"Provides the pyramidal weighting scheme and back-propagation update rules that 3DPyraNet adapts.","marker":"Phung and Bouzerdoum [2007]"},{"why":"C3D is the chief deep baseline on YUPENN and Maryland and the reference for the parameter-count comparison (17.5M vs 0.83M).","marker":"Tran et al. [2015]"},{"why":"DPCF is the per-class and overall accuracy baseline for YUPENN and Maryland that the paper claims to approach or overtake on Maryland.","marker":"Feichtenhofer et al. [2016]"},{"why":"Introduces the KTH dataset and a local SVM baseline, giving the action-recognition protocol that the paper partially follows.","marker":"Sch¨uldt et al. [2004]"},{"why":"Supplies the early/global fusion framing for video features and the large-scale pretraining context that the paper contrasts with its small-dataset training.","marker":"Karpathy and Leung [2014]"}],"fun_headline_variants":["Sub-million parameter video net beats C3D on Maryland","3DPyraNet-F: feature fusion for low-cost spatio-temporal learning","3DPyraNet-F: sub-million params, top results on 3 benchmarks","Fusion of 3D pyramidal features lifts video action recognition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline numbers rest on the assumption that the paper's loosely described training and evaluation splits match the baselines it is compared with, so the reported margins are not an artifact of easier or harder test sets.","fun_headline_variants_meta":{"raw":{"variants":["Sub-million parameter video net beats C3D on Maryland","3DPyraNet-F: feature fusion for low-cost spatio-temporal learning","3DPyraNet-F: sub-million params, top results on 3 benchmarks","Fusion of 3D pyramidal features lifts video action recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001434,"raw_usage":{"total_tokens":5829,"prompt_tokens":1042,"completion_tokens":4787,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":4707}},"tokens_in":658,"tokens_out":4787,"duration_ms":34103,"temperature":1.0,"reasoning_tokens":4707,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:03:18.902426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run 3DPyraNet-F on KTH and Weizmann under the exact standard protocols used by the cited baselines, using fixed public splits, multiple runs, and per-class standard deviations. If the mean accuracy falls below the reported 93.42% on KTH or 98.99% on Weizmann by more than a couple of points, the state-of-the-art comparison is not robust to split choice.","supporting_citations":[],"review_version":1}