{"id":"9796bf5b-3d6b-48e9-b81a-792729943c40","arxiv_id":"2509.09151","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.","lead":"This survey frames the history of video understanding as a story of datasets shaping model design, linking motion, temporal span, hierarchy, and multimodality to architecture families. It is a synthesis and roadmap, not a new experiment, and it aims to guide dataset construction and model selection.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset-driven causal claim is undercut by the paper's own benchmark protocol: Table III selects the best model per dataset and contradicts the claim in Section IV.A that two-stream/3D CNNs dominate short-clip datasets.","rationale":"The reader's verdict REJECT is supported. The paper's central thesis is causal: dataset structural properties drive architecture evolution. The evidence provided is a correlation between dataset attributes and best-performing models in Tables III and IV. However, Table III explicitly selects the best-performing variant per dataset, which systematically hides the fact that on short-clip datasets like HMDB51 and UCF101, the highest accuracies are achieved by modern transformers/MAEs, not by the two-stream or 3D CNNs that the text credits with dominance. This internal contradiction is not a peripheral issue; it is the cornerstone of the causal argument. The paper asserts the causal arrow in Section III.A and V.B without any controlled comparisons that would rule out alternative drivers such as compute, model scale, pretraining data, or leaderboard incentives. Given that the survey itself acknowledges protocol fragmentation and architecture incentives driven by benchmarks, the empirical foundation for 'datasets as the principal structural force' collapses. A revised paper could survive as a descriptive survey if it softened the causal claim and acknowledged the contradictory evidence, but as written, the central claim is not supported. This is not an ad hominem or a disagreement with consensus; it is an internal inconsistency between the text and the paper's own tables. My agreement with the reader is strong: we both identify the same load-bearing weakness. I recommend REJECT, or at minimum a major revision with CONDITIONAL acceptance if the authors reframe the contribution as a dataset-centric taxonomy rather than a causal history.","tokens_in":40094,"tokens_out":2709,"duration_ms":37379,"concrete_test":"Run a controlled comparison on HMDB51 and UCF101: take a representative 3D CNN (e.g., RGB-I3D) and a masked-autoencoder transformer (e.g., VideoMAE V2), initialize both from the same Kinetics-400 pretrained weights, use identical input sampling and fine-tuning budget, and report top-1 accuracy. If the transformer matches or exceeds the 3D CNN on these short-clip datasets, then the Section IV.A claim that short-clip, motion-centric datasets favor two-stream/3D CNNs is directly refuted, and the dataset-driven causal pattern loses its clearest supporting example. A secondary check: recompute Table III's 'consistently outperform' statement using all reported model entries for HMDB51/UCF101 (not only the best variant) and count how many of the top-5 results are transformers rather than early 3D CNNs.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim, stated in Section V.B ('Datasets are not passive benchmarks; they are the principal structural force shaping model design'), depends on the evidence in Tables III and IV. But Table III states 'the best-performing model variant is reported' for each dataset. This protocol cannot support the causal pattern claimed in Section IV.A, where 'On HMDB51 and UCF101, early Two-Stream variants and 3D CNNs consistently outperform others.' The same table shows VideoMAE V2 at 88.1/99.6 and InternVideo at 89.3 on HMDB51/UCF101, far above Two-Stream'16 (69.2/93.5) and RGB-I3D (74.8/95.6). Selecting the best variant obscures the fact that modern transformers and masked-autoencoder models achieve the highest numbers on the very datasets the paper uses to argue for 3D-CNN dominance. Thus, the empirical core of the paper's causal narrative is an artifact of cherry-picking, not a controlled comparison: no attempt is made to hold pretraining data, compute, input sampling, or evaluation protocol constant across model families. Without such control, the observed correlation between dataset attributes and architecture success is equally explained by model scale, pretraining sophistication, or leaderboard incentives. The paper's own framing in Section V.A even acknowledges that evaluation fragmentation and leaderboard tuning shape architectural incentives, undercutting the exclusive causal role assigned to dataset structure in Section III.A and V.B.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey argues that the evolution of video understanding architectures is fundamentally shaped by dataset structure. It introduces a 'dataset-bias-architecture' framework in which four dataset properties—motion amplitude, temporal span, compositionality/hierarchy, and multimodal richness—impose inductive biases that drive architectural choices. The paper organizes major datasets into these categories (Table II), reviews milestones from two-stream CNNs and 3D CNNs through transformers and video-language models, and uses curated benchmark tables (Tables III and IV) to claim that dataset properties predict which model families succeed. It concludes with a prescriptive roadmap for aligning model design with dataset structure and for constructing future datasets.","tokens_in":40422,"tokens_out":5226,"duration_ms":64808,"significance":"If the central causal claim—that datasets are 'the principal structural force shaping model design'—were established, the survey would offer a useful organizing perspective for a fragmented literature. The paper has clear strengths: Table II is a broad compendium of datasets with structural annotations, the coverage of modern video-language and egocentric datasets is current, the authors provide code and dynamic visualizations, and the roadmap in Section V is actionable. However, the significance is currently undermined by the fact that the empirical support consists of selectively curated benchmark numbers and hand-assigned dataset ratings, with no controlled comparison. The central claim is plausible as a retrospective narrative but is not demonstrated at the strength asserted.","major_comments":[{"comment":"The text states that 'On HMDB51 and UCF101, early Two-Stream variants and 3D CNNs consistently outperform others,' citing Two-Stream'16 (69.2/93.5) and RGB-I3D (74.8/95.6). The same Table III lists VideoMAE V2 at 88.1 and 99.6 on these datasets, and InternVideo at 89.3 on HMDB51. These are 15–20 points higher, so the sentence is factually contradicted by the paper's own table. The table caption also says 'the best-performing model variant is reported,' meaning rows are not comparable under a fixed protocol: models differ in pretraining data, compute, input sampling, and evaluation settings. This contradiction undermines the inference that short-clip datasets 'strongly favor' two-stream/3D CNNs. A similar issue appears in Section IV.B, where early backbones are said to 'consistently excel' in temporal localization, while Table IV lists InternVideo2 at 72.0 mAP on THUMOS'14 versus I3D+Flow","section":"Section IV.A, Table III"},{"comment":"The central thesis that datasets are 'the principal structural force shaping model design' is asserted on the basis of a correlation between hand-assigned dataset attributes (Table II) and selected architecture successes (Tables III–IV). No attempt is made to hold constant model scale, pretraining data, compute budget, or evaluation protocol, so the observed alignment is equally consistent with compute-driven or pretraining-driven evolution. Moreover, Section V.A itself acknowledges that 'evaluation fragmentation' and leaderboard incentives 'shape architectural incentives'—a non-dataset confound. The causal claim is therefore not supported by the presented evidence. Please either weaken the claim to 'an important and underexamined influence' or provide a more rigorous argument, e.g., a historical timeline showing architecture transitions following dataset releases, or citations to ablati","section":"Sections III.A and V.B"},{"comment":"The structural ratings (Amp/Span/Comp/Agents) are central to the framework, but they are assigned without an explicit rubric, operational definitions, inter-annotator agreement, or sensitivity analysis. For example, Kinetics-400 is rated Amp=H, Span=S, Comp=-, Agents=M, but the criteria for these levels are not given, and a different researcher could plausibly rate the same dataset differently. Because these ratings are used to support the paper's main narrative, their subjectivity is load-bearing. Please provide a coding protocol, report reliability, or explicitly relabel the ratings as informal and reduce their role in the causal argument.","section":"Table II"},{"comment":"The selective reporting in Tables III and IV makes it difficult to interpret 'dashes' as capabilities. For example, VideoMAE V2 has no retrieval or QA entries in Table IV, and InternVideo2 lacks QA entries, yet the text interprets such absences as evidence that certain model families are specialized or limited. A model with no reported number may simply have not been evaluated on that benchmark. The tables should include a completeness statement or a reference to the original papers' evaluation suites, and the text should avoid reading missing entries as negative evidence.","section":"Tables III and IV"}],"minor_comments":[{"comment":"The model 'Two-Stream'16' is cited as [184] (Feichtenhofer et al., 2016), but the original two-stream architecture is [38] (Simonyan and Zisserman, 2014). Please clarify the naming to avoid confusion between the two papers.","section":"Section IV.A and References"},{"comment":"The row labeled 'Swin' refers to Video Swin Transformer [239]; using the unqualified name may be confused with the image Swin Transformer. Please rename to 'Video Swin'.","section":"Table III"},{"comment":"The UCF101-24 dataset is listed with year 2024, but UCF101-24 is a subset of UCF101 with spatio-temporal annotations and is much older. Please correct the year or clarify the provenance.","section":"Table II"},{"comment":"The arXiv abstract uses the title 'Video Understanding by Design: How Datasets Shape Video Models,' while the manuscript header title is '... How Datasets Shape Architectures and Insights.' Please unify the title and abstract wording.","section":"Front matter"},{"comment":"The caption says 'Images adopted from [180]'; for a journal submission please confirm that permission or license for reuse is obtained and that the source is clearly credited.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The paper's thesis is interesting but the current evidence base does not support the strong causal language. The internal contradiction between Section IV.A and Table III is the most serious issue and should be fixed before any further consideration. I also note that several of the references are to the first author's own prior survey-like papers (e.g., [8], [11], [12]); this is not by itself a problem, but the authors should ensure that the present survey does not substantially duplicate that material. Editors may also wish to consider whether the manuscript's scope and presentation match the journal's typical format for surveys, given the heavy emphasis on a single interpretive narrative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nTwo things to know. First, this is a genuinely comprehensive dataset-centric survey of video understanding, and Table II — organizing roughly a hundred datasets by motion amplitude, temporal span, compositionality, and agent density — is a genuinely useful reference. The framing of datasets as inductive-bias generators and architectures as responses is a coherent way to retell the field's history, and the sections on hierarchical/compositional structure and egocentric video are informative. Second, the paper's central causal claim is undercut by its own evidence, and not in a subtle way.\n\nThe thesis is stated bluntly in Section V.B: 'Datasets are not passive benchmarks; they are the principal structural force shaping model design.' The supporting evidence is the curated benchmark tables. But Table III's protocol — reporting the best-performing model variant per dataset — directly contradicts the Section IV.A claim that 'On HMDB51 and UCF101, early Two-Stream variants and 3D CNNs consistently outperform others.' VideoMAE V2 reaches 88.1 on HMDB51 and 99.6 on UCF101; InternVideo reaches 89.3. Two-Stream'16 gets 69.2 and 93.5. So the very datasets used to argue for 3D-CNN dominance are actually topped by masked-autoencoder transformers. Selecting the best variant per dataset hides this, and there is no control for pretraining data, compute, input sampling, or evaluation protocol.\n\nThere is also a circular flavor to the framework: the four dataset attributes are hand-assigned H/M/L ratings in Table II, and the causal story is then read back from the same curated examples. The paper itself acknowledges in Section V.A that leaderboard incentives and evaluation fragmentation shape architectural incentives, which dilutes the exclusive causal role assigned to dataset structure. The load-bearing thesis is asserted, not demonstrated.\n\nTo be fair, the survey does solid work as a compilation and a roadmap. If the authors softened the claim from 'principal structural force' to 'one important factor,' and either re-ran the benchmark analysis with proper controls or simply acknowledged that modern transformers win on those same short-clip datasets, the survey could be genuinely useful. As it stands, the internal contradiction makes the central narrative unreliable.\n\nWho should read it: anyone wanting a broad dataset taxonomy or a survey of video architectures. It deserves serious peer review — heavy revision, not a desk reject — because the survey is substantial and the flaw is fixable. I wouldn't cite it until the causal claim is reworked.","headline":"A useful survey with a serious internal contradiction: its own benchmark table refutes its central claim that early 3D CNNs dominate short-clip datasets.","tokens_in":40916,"tokens_out":2352,"would_cite":false,"duration_ms":26998,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that the structure of video datasets—motion complexity, temporal span, compositionality, and multimodal richness—is the principal force shaping model architecture, making dataset design a strategic lever for the field.","keywords":["video understanding","dataset-centric analysis","inductive bias","architecture evolution","action recognition","video transformers","vision-language models","benchmark analysis"],"falsifier":"Run a controlled study that keeps the architecture family fixed, varies a single dataset attribute (e.g., temporal span while holding motion amplitude constant), and shows no systematic performance ordering; alternatively, find two datasets with identical attribute ratings that produced very different dominant architectures. Either result would break the claimed causal link.","tokens_in":39960,"feed_emoji":"🎥","tokens_out":4307,"duration_ms":44982,"temperature":0.7,"pith_summary":"The paper tries to establish a single organizing claim: video-understanding architectures do not evolve on their own; they are responses to the structural properties of the datasets the field builds and adopts. Four properties—motion amplitude, temporal span, compositional/hierarchical structure, and multimodal richness—act as structural pressures that favor specific inductive biases, from two-stream CNNs and 3D convolutions for short motion clips, to temporal reasoning networks and transformers for long procedural sequences, to vision-language models for text-paired corpora. A sympathetic reader should care because this reframes dataset construction as an active design choice that determines which kinds of video intelligence become possible, and it offers a matching rule: pick the architecture whose inductive bias fits the dataset's structure. The paper supports the claim with a large comparative table of datasets and benchmarks spanning two decades.","feed_headline":"Datasets, not models, steer video AI evolution","feed_subtitle":"Four dataset traits—motion, span, composition, multimodality—explain the shift from 3D CNNs to transformers and video-language models.","key_machinery":"The carrying mechanism is the dataset-bias-architecture framework. It breaks video datasets into four structural properties—motion amplitude, temporal span, compositionality/hierarchy, and multimodal richness (plus agent density)—and treats every architecture as an enforced inductive bias that matches (or mismatches) those properties. Table II operationalizes the framework as a compact rating system (H/M/L for amplitude and agents, S/M/L for span, -/C/H for composition) across a century of datasets; Tables III and IV connect those ratings to measured performance of representative models. The framework does the explanatory work: it turns the history of video understanding into a sequence of d","core_discovery":"The central discovery is that datasets operate as inductive-bias generators. Each dataset imposes invariances its model must internalize: coarse high-amplitude motions reward instantaneous motion capture (optical flow, shallow 3D filters); long-horizon, overlapping activities reward temporal memory and hierarchy; multi-agent scenes reward relational or graph representations; and video-text corpora reward cross-modal alignment. On this reading, the milestone trajectory—two-stream networks, 3D CNNs, temporal segment/relation networks, transformers, masked self-supervised models, and video-language foundation models—is not a random succession of fashions but a systematic accommodation of increa","pith_inferences":["Editorial inference: The causal arrow (datasets shape architectures) could be tested directly by controlled experiments that fix the architecture family and vary one structural attribute at a time; the survey does not run such ablations, so the claim remains an interpretation of correlated historical patterns.","Editorial inference: If the framework holds, it predicts that next-generation long-horizon, multi-agent, multimodal corpora will push the field toward memory-augmented and state-space models plus retrieval-augmented video understanding, since those structures directly target temporal span and compositionality.","Editorial inference: The authors' H/M/L ratings are assigned by hand; a community-validated or automatically computed scoring of dataset attributes would let the framework serve as a reusable diagnostic tool for predicting which architecture family suits any new benchmark.","Editorial inference: The same lens could be applied prospectively during dataset construction: deliberately vary the four attributes to probe whether an architecture's inductive bias is genuinely being challenged, rather than relying on leaderboard rankings alone."],"forward_implications":["Matching architecture to dataset structure pays off: short-clip motion datasets favor two-stream and 3D CNN models, compositional and interaction-heavy datasets favor sequential and transformer models, and text-paired corpora favor video-language pretraining.","Training on coarse, motion-only datasets yields fragile transfer; if robustness in nuanced real-world settings is desired, motion granularity must appear in the data.","Simply scaling class counts or clip counts will not yield general video intelligence; the decisive ingredient is structure—procedural hierarchies, temporal continuity, and precise cross-modal alignment.","Future architectures should integrate temporal precision, hierarchical composition, long-horizon attention, and multimodal grounding; future datasets should be built with sub-second audio-text alignment, multi-agent annotations, and compositional evaluation splits.","Dataset design should be treated as a strategic lever, not a scaling exercise, because datasets generate the invariance pressures that architectures evolve to accommodate."],"fun_headline_variants":["Datasets are the real architects of video AI","Why video models mimic their training data","Video AI: shaped by data, not just design","Your dataset decides your video model","From CNNs to transformers: data decides the path"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the four hand-selected dataset attributes are the dominant cause of architectural change—rather than compute availability, leaderboard incentives, or model-family trends—and that the paper's H/M/L ratings of each dataset are accurate and sufficient.","fun_headline_variants_meta":{"raw":{"variants":["Datasets are the real architects of video AI","Why video models mimic their training data","Video AI: shaped by data, not just design","Your dataset decides your video model","From CNNs to transformers: data decides the path"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000148,"raw_usage":{"total_tokens":1051,"prompt_tokens":791,"completion_tokens":260,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":206}},"tokens_in":535,"tokens_out":260,"duration_ms":3134,"temperature":1.0,"reasoning_tokens":206,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T19:34:08.366758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled study that keeps the architecture family fixed, varies a single dataset attribute (e.g., temporal span while holding motion amplitude constant), and shows no systematic performance ordering; alternatively, find two datasets with identical attribute ratings that produced very different dominant architectures. Either result would break the claimed causal link.","supporting_citations":[],"review_version":1}