{"id":"4ba43b62-facc-4a1f-b698-f7793dba4a60","arxiv_id":"2608.10765","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A synthesis protocol builds a regenerable four-level intention benchmark from flat action data and shows a consistent compositional held-out gap across four baseline model families.","lead":"This paper presents a framework that turns a flat action corpus into a four-level hierarchical benchmark with actions, activities, low-level intentions, and high-level intentions. The framework includes an anti-circularity design and yields a compositional generalization gap that appears across all tested model families.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The structural-property claim rests on a single synthesized benchmark instance; the reported seed variance covers only model initialization, not the generator's own sampling randomness.","rationale":"The reader's weakest assumption concerns ontology and Group Y independence. I agree this matters for external validity, but the paper already disclaims that 'intentions are imposed by construction rather than observed' (Section 8), so an idiosyncratic ontology weakens the semantic interpretation, not the internal claim that the gap is structural. The load-bearing condition is regeneration stability: the benchmark is synthesized from a stochastic transition model and sampler, and the central assertion is about the benchmark ('structural property') rather than about these 15,002 episodes. All reported variance comes from model training seeds, which cannot distinguish a property of the protocol from a property of one generated corpus. I therefore partially agree with the reader: the missing significance and robustness analysis is real, but I locate it in benchmark regeneration rather than ontology/Group Y. A regeneration-seed analysis is the one check that would settle whether the headline gap is a stable property or a single-sample artifact. The concern does not change the verdict category; it strengthens the existing CONDITIONAL verdict until the test is run.","tokens_in":11050,"tokens_out":7591,"duration_ms":87864,"concrete_test":"Regenerate the benchmark from the released generator with at least 10 independent synthesis seeds, holding fixed the ontology, transition-model specification, source features, and split-construction procedure. For each seed, reproduce the standard and compositional HLI macro-F1 for at least the R-GCN and hierarchical transformer under the paper's protocol, and record the compositional gap. If the 95% interval over regeneration seeds overlaps zero or spans more than ±0.03, the 'structural property' claim is not supported; if it remains within 0.13-0.17, the concern is resolved. Report also the per-LLI gaps to check that the aggregate is not driven by one of the four multi-parent LLIs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that a 0.13-0.17 compositional held-out gap is 'a structural property of the benchmark rather than a model artifact' (Abstract; Section 7). For that claim to hold, the gap must be stable under the benchmark's own stochastic generation process, because the contribution is a synthesis protocol, not one dataset. The reported variability (Table 6; bootstrap CI [0.161, 0.171]) is computed over three to five model-training seeds only; it does not vary the transition-model, perturbation, or coverage-sampler randomness that produces the 15,002 episodes. A single instantiation of a stochastic generator cannot establish a structural property: if regenerating the benchmark with different synthesis seeds shifts the gap outside 0.13-0.17, or to zero, the observed gap is a property of that one draw, not of the protocol. This is especially pointed because the compositional split is built on only four multi-parent LLIs, so the gap could be sensitive to which parent associations are withheld and how the sampler assigns subjects. The paper's Section 8 explicitly frames the resource as a regenerable protocol, yet no regeneration-seed analysis is reported. The released generator makes the missing check feasible.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a benchmark-generation framework that synthesizes a four-level hierarchical intention benchmark (actions, activities, low-level intentions, high-level intentions) from a flat single-label action corpus (NTU RGB+D 120), while retaining real pre-extracted action features. Episodes are assembled via a transition model under a subject-consistency constraint and a coverage-aware sampler that reduces the subject-usage Gini from 0.566 to 0.248. To avoid circular supervision, sequence-generation rules (Group X) are kept disjoint from evaluation-time first-order-logic rules (Group Y), and a compositional held-out split withholds LLI-HLI parent associations. The instantiated benchmark contains 15,002 episodes. Four reference baselines (hierarchical transformer, sequential transformer, bag-of-actions MLP, relational graph network) are evaluated, yielding a compositional held-out gap of 0.13–0.17 macro-F1 across all baselines. The paper claims this gap is a structural property of the benchmark and not a model artifact, supported by logic-violation rates and an order-destroying control. The ontology, transition model, and generator are released.","tokens_in":11289,"tokens_out":6879,"duration_ms":72957,"significance":"If the central result holds, the paper makes a useful contribution to compositional and hierarchical action recognition: it provides a regenerable protocol that turns a flat action corpus into a deep intention hierarchy with explicit anti-circularity safeguards, coverage balancing, and compositional splits. The released generator and ontology enable others to regenerate and extend the benchmark. The compositional gap observed across four model families, including a graph-aware model, is an interesting finding that could guide future benchmark design. However, the strength of the claim that the gap is 'structural' is currently limited by the lack of regeneration-seed variation and by uncontrolled split comparisons. The paper is a benchmark-protocol paper, not a recognition-methods paper, and its value depends on the empirical stability of the reported gap.","major_comments":[{"comment":"The claim that the compositional gap of 0.13–0.17 macro-F1 is 'a structural property of the benchmark rather than a model artifact' is under-supported because the reported variability (three to five seeds) covers only model-initialization randomness, not the synthesis generator's own randomness (transition-model perturbations, coverage sampler, episode sampling). The paper frames the contribution as a regenerable protocol (Section 8; Data Availability), so a single instantiation of the stochastic generator cannot establish a structural property. Please report the gap's stability over multiple independent regeneration seeds (e.g., 10 regenerated benchmark instances with the same ontology and hyperparameters), or temper the claim accordingly.","section":"Section 7, Table 6 and Abstract"},{"comment":"The decomposition of the gap into 1,854 novel-context episodes and 321 ordering-pattern episodes is not defined anywhere in the paper: no section explains how these subsets are derived or what 'held out by ordering pattern alone' means. Without this definition, the claim that the gap is 'concentrated on genuine novel compositions' cannot be checked. Furthermore, the standard and compositional splits appear to be different episode sets; a matched test set (identical held-out episodes, differing only in which LLI-HLI associations appear in training) is needed to rule out the possibility that the gap reflects raw distribution shift rather than compositional failure.","section":"Section 7, 'Compositional generalization' paragraph"},{"comment":"The text states that destroying order 'changes macro-F1 by less than one standard deviation at every level and baseline,' but the reported numbers are inconsistent with this claim: for example, the hierarchical transformer's action-level drop of +2.9 with standard deviation 2.0 is 1.45 standard deviations, and the sequential transformer's action-level drop of +5.2(5.5) is near one standard deviation but the text's 'around 0.02 to 0.04' range does not match the table's range of 0.009–0.052. Please correct the text and provide proper significance testing (e.g., a paired bootstrap or permutation test across episodes) before concluding that the order-destroying control shows no effect.","section":"Section 7, Table 8 and surrounding text"},{"comment":"The claim that 'every model exceeds the 2.5% Group Y data floor, confirming genuine error rather than generator recovery' is based on mean differences of 0.3–1.7 percentage points without any significance test or confidence interval. Because the 2.5% floor is itself an estimate from a single generated dataset, please report a test (e.g., bootstrap over episodes or seeds) for whether each model's Group Y violation rate is significantly above the data floor. As written, the evidence is suggestive but not statistically established.","section":"Section 7, Table 7"},{"comment":"The anti-circularity argument hinges on the assertion that the 22 Group Y rules are 'semantic constraints the generator does not use' and 'any annotator would apply independently.' Since both the generator's transition model and the Group Y rules are authored by the same researchers on the basis of the same ontology, this independence is not automatic and is not empirically demonstrated. Please provide a direct test, for example by measuring the correlation between the transition model's soft ordering preferences and the Group Y rules, or by showing that the Group Y rules cannot be derived from the ontology's transition structure. Without such a check, the validity claim is only asserted.","section":"Section 5, 'Disjoint generation and evaluation rules'"}],"minor_comments":[{"comment":"The text says 'five independent annotators judged whether the action sequences of 150 sampled episodes form coherent behavior' but then reports 'majority-vote plausibility rate was 0.86 (43/50)'. The numbers 150 and 50 are inconsistent; please clarify how many episodes were annotated and how the 43/50 was computed.","section":"Section 5, plausibility check"},{"comment":"There is a typo in the subsection heading: 'heterogeneous-graphepisode' should be 'heterogeneous-graph episode'.","section":"Section 4, 'Heterogeneous graph construction'"},{"comment":"The caption says 'bootstrap 95% confidence interval on the hierarchical-transformer gap over five seeds is [0.161,0.171]' but it is not stated explicitly that this interval is over model-training seeds only and not over regeneration seeds; please state this clearly to avoid misunderstanding.","section":"Section 7, Table 6 caption"},{"comment":"Table 8's caption says 'mean±standard deviation in parentheses over three seeds,' while Tables 5 and 6 report HT and R-GCN over five seeds and seq-T over three. Please clarify whether Table 8 used three seeds for all baselines and why, or align the captions.","section":"Section 7, Table 8 caption"},{"comment":"The claim that the framework 'transfers to other flat single-label action corpora' is stated as a property, but no transfer demonstration is provided. Since a second corpus is explicitly left to future work, please soften the claim or add a discussion of the conditions under which transfer might fail.","section":"Section 8, 'Discussion and Conclusion'"}],"recommendation":"major_revision","confidential_remarks":"This is a benchmark-protocol paper. The central idea of synthesizing a deep intention hierarchy with anti-circularity safeguards is sound and topical. The main risk is that the central empirical claim ('structural gap') is asserted from a single generated dataset; a regeneration-seed study is feasible because the generator is released, and it should be required before the claim can be accepted. Additionally, the split-decomposition and order-control inconsistencies should be corrected. The paper is likely to be a useful contribution after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, well-scoped benchmark-synthesis paper with a real empirical finding, but the headline claim that the compositional gap is a structural property of the benchmark goes beyond the evidence. The gap is stable across model seeds, not across regeneration of the benchmark itself, and the paper's own framing as a regenerable protocol makes that distinction load-bearing.\n\nWhat's new: the combination of a four-level intention hierarchy, subject-consistent coverage-aware sampling, disjoint generation and evaluation rules, and compositional held-out splits in one protocol is genuinely not in the earlier literature. Individual pieces exist, but the integration is useful and they've implemented it carefully. The anti-circularity design is the best part: keeping the 22 Group Y semantic rules out of the generator, then showing that logic-free baselines violate them above the 2.5% intrinsic data floor, is a clean way to demonstrate that the benchmark isn't solvable by generator recovery. The order-destroying control is also handled honestly—they treat it as a generator-consistency check, not as evidence about temporal order, and the near-zero effect at the intention level is appropriately interpreted.\n\nThe compositional gap itself is credible as an empirical observation: 0.13–0.17 macro-F1 drop across four model families, concentrated on episodes with truly novel LLI–HLI pairings. The bootstrap CI on the hierarchical transformer is a nice touch, though it only covers model seed variation.\n\nSoft spots, in proportion. The main one is the stress-test issue: the gap is measured on a single instantiation of a stochastic generator. The transition-model perturbations, sampling randomness, and coverage sampler are never varied, so we don't know whether 0.13–0.17 is robust or a property of one draw. For a paper whose contribution is a protocol, that's a missing experiment. Second, the Group Y violation-rate differences between baselines (2.8–4.2%) are small and no significance test is reported; they might be within noise. Third, the ontology is expert-authored and the paper concedes intentions are imposed, not observed. That limitation is stated clearly, but it means the benchmark measures reasoning over a constructed hierarchy, which is fine as long as it's not oversold. Finally, code and artifacts are not yet released, so the regenerability claim is currently unverifiable.\n\nWho this is for: researchers building or evaluating hierarchical action-recognition models, and anyone working on compositional generalization benchmarks. It's a solid contribution that needs revision, not rejection. I'd send it to a serious referee, with the request that the authors either add regeneration-seed analysis or soften the structural-property language.","headline":"A credible compositional-gap result on a thoughtful benchmark-synthesis protocol, but the structural-property claim needs regeneration-seed analysis before it holds.","tokens_in":11801,"tokens_out":1966,"would_cite":true,"duration_ms":19654,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that a benchmark synthesized from flat action clips forces all tested model families to lose 0.13–0.17 macro-F1 on new combinations of known intention components, and that the gap is structural because a graph-aware model…","keywords":["hierarchical action recognition","benchmark synthesis","compositional generalization","neuro-symbolic learning","first-order logic constraints","hierarchical transformer","multimodal action recognition"],"falsifier":"Regenerate the benchmark with an independently authored ontology and rule set produced by a different set of experts; if the compositional held-out gap shrinks to near zero on this second version, the gap is an artifact of the particular semantic standard rather than a structural property of compositional intention recognition.","tokens_in":10887,"feed_emoji":"🧩","tokens_out":9315,"duration_ms":77883,"temperature":0.7,"pith_summary":"This paper proposes a way to build a four-level hierarchical-intention benchmark — actions, activities, low-level intentions, and high-level intentions — from a flat, single-label action corpus, while keeping real pre-extracted features at the action level. The central claim is that this benchmark exposes a structural limitation in current action-recognition models: when new combinations of known intention components are held out from training, every tested model family loses 0.13 to 0.17 macro-F1 at the high-level-intention tier, and a graph-aware model that otherwise recognizes best does not close the gap. Because the drop appears across architectures and is concentrated on genuinely novel compositions rather than distribution shift, the authors conclude it is a property of the compositional task itself. If correct, the work provides a regenerable testbed for controlled evaluation of hierarchical and compositional reasoning about human behavior.","feed_headline":"Models lose 0.13–0.17 macro-F1 on held-out intentions","feed_subtitle":"A synthesized benchmark shows the gap is structural: even the best graph model cannot close it.","key_machinery":"The load-bearing mechanism is the benchmark synthesis protocol itself: a four-level ontology (actions → activities → low-level intentions → high-level intentions) over 85 action classes, a transition model that assembles subject-consistent episodes, a coverage-aware sampler that reduces the subject-usage Gini from 0.566 to 0.248, and a validity design that keeps the sequence-generation rules disjoint from the 22 evaluation-time first-order-logic Group Y rules. Compositional held-out splits withhold one LLI→HLI association per multi-parent LLI, forcing transfer to unseen high-level contexts. An order-destroying control permutes sibling order while fixing labels, performers, and membership, serving as a generator-consistency check. Together these components make hierarchy, coverage, circularity, and compositional generalization measurable in one setting.","core_discovery":"The paper's central discovery is the compositional held-out gap: a consistent drop in high-level-intention macro-F1 of 0.13 to 0.17 when LLI-to-HLI associations are withheld from training, observed across all four reference baselines (hierarchical transformer, sequential transformer, bag-of-actions MLP, and relational graph network) and every seed. The graph-aware R-GCN is the strongest recognizer yet shows the largest gap, which the authors take as evidence that the gap is structural rather than a capacity artifact. A decomposition shows the drop is far larger on episodes containing truly novel LLI→HLI pairings (0.47) than on episodes held out by ordering pattern alone (0.30), confirming genuine composition rather than distribution shift. The paper also finds that logic-free baselines violate the evaluation-only semantic rules above the 2.5% intrinsic violation rate of the ground-truth labels, arguing the benchmark cannot be solved by recovering the generator.","pith_inferences":["The anti-circularity protocol — measuring violation rates on evaluation-only rules and checking that ground-truth labels violate them at a non-zero rate — could be adopted by other synthetic benchmarks as a cheap red flag for generator recovery.","If the ontology were rebuilt from scratch by a different expert panel, the size of the compositional gap on the new version would reveal how much of the difficulty depends on this particular semantic standard rather than on compositional structure in general.","The order-flexibility finding suggests a design axis for benchmark generators: tuning the transition model's temporal strength can modulate how much of the task is about sequencing versus co-occurrence, letting researchers target specific reasoning abilities.","A model that explicitly factorizes LLI and HLI components might close part of the gap; testing this would show whether the structural property is an irreducible interaction between component features."],"forward_implications":["Benchmark scores on standard splits overstate compositional ability: models must also be evaluated on held-out intention combinations to measure genuine generalization.","The framework transfers to any flat single-label action corpus with per-performer metadata and sufficient class coverage, enabling hierarchy benchmarks in other domains without new recording campaigns.","The released ontology, transition model, and generator allow the benchmark to be regenerated and extended, supporting controlled comparisons of neuro-symbolic models that incorporate logical constraints.","Because order destruction leaves intention-level performance nearly unchanged, high-level intention in this benchmark is recoverable from co-occurrence of sub-behaviors; future versions can inject stronger temporal signal to test order-based reasoning."],"supporting_citations":[{"why":"Supplies the flat single-label action corpus with per-performer metadata and pre-extracted skeleton features that the synthesis composes into episodes.","marker":"[11]"},{"why":"Provides the stochastic-grammar grounding for the transition model that assembles ordered episodes from atomic clips.","marker":"[17]"},{"why":"Articulates the circular-supervision risk in synthetic evaluation, motivating the disjoint generation-versus-evaluation rule design.","marker":"[7]"},{"why":"Contributes the compositional held-out split protocol that the benchmark applies to LLI-to-HLI associations.","marker":"[14]"},{"why":"Establishes the compositional-generalization challenge of held-out unseen combinations that the benchmark adapts to intention levels.","marker":"[9]"},{"why":"Demonstrates zero-shot compositional action recognition, a neighboring task whose evaluation this benchmark extends to deeper hierarchy.","marker":"[10]"},{"why":"Frames dataset validity and documentation, which the paper operationalizes into anti-circularity and order-control checks.","marker":"[6]"},{"why":"Shows procedural generation of training data, the synthetic-data lineage this work extends to semantic label-structure synthesis.","marker":"[5]"}],"fun_headline_variants":["Structural gap: all models fail on novel action compositions","Even best graph model can't close 0.13–0.17 compositional gap","Held-out intentions reveal structural benchmark gap, not model limit","New benchmark: models lose 0.13–0.17 F1 on unseen intention combos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that the experts' definition of what counts as a coherent intention, and the 22 rules used at evaluation, are a valid and stable standard independent of the generator; if that standard is idiosyncratic or shaped by the generator, the benchmark tests an arbitrary construct rather than human intentions.","fun_headline_variants_meta":{"raw":{"variants":["Structural gap: all models fail on novel action compositions","Even best graph model can't close 0.13–0.17 compositional gap","Held-out intentions reveal structural benchmark gap, not model limit","New benchmark: models lose 0.13–0.17 F1 on unseen intention combos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000826,"raw_usage":{"total_tokens":3675,"prompt_tokens":1072,"completion_tokens":2603,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":688,"completion_tokens_details":{"reasoning_tokens":2521}},"tokens_in":688,"tokens_out":2603,"duration_ms":19723,"temperature":1.0,"reasoning_tokens":2521,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:50:15.015481+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate the benchmark with an independently authored ontology and rule set produced by a different set of experts; if the compositional held-out gap shrinks to near zero on this second version, the gap is an artifact of the particular semantic standard rather than a structural property of compositional intention recognition.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the flat single-label action corpus with per-performer metadata and pre-extracted skeleton features that the synthesis composes into episodes."},{"cited_title":"Predicting human activities using stochastic grammar, in: Proceedings of the IEEE International Conference on Computer Vision, pp","cited_arxiv_id":null,"evidence_quote":"Provides the stochastic-grammar grounding for the transition model that assembles ordered episodes from atomic clips."},{"cited_title":"1049– 1059","cited_arxiv_id":null,"evidence_quote":"Contributes the compositional held-out split protocol that the benchmark applies to LLI-to-HLI associations."},{"cited_title":"Cogs: A compositional generalization challenge based on semantic interpretation, in: Proceedings of the 2020conferenceonempiricalmethodsinnaturallanguageprocessing (emnlp), pp","cited_arxiv_id":null,"evidence_quote":"Establishes the compositional-generalization challenge of held-out unseen combinations that the benchmark adapts to intention levels."},{"cited_title":"C2c: Component-to-composition learning for zero- shot compositional action recognition, in: European Conference on Computer Vision, Springer","cited_arxiv_id":null,"evidence_quote":"Demonstrates zero-shot compositional action recognition, a neighboring task whose evaluation this benchmark extends to deeper hierarchy."},{"cited_title":"Datasheetsfordatasets","cited_arxiv_id":null,"evidence_quote":"Frames dataset validity and documentation, which the paper operationalizes into anti-circularity and order-control checks."},{"cited_title":"Procedu- ralgenerationofvideostotraindeepactionrecognitionnetworks,in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp","cited_arxiv_id":null,"evidence_quote":"Shows procedural generation of training data, the synthetic-data lineage this work extends to semantic label-structure synthesis."}],"review_version":1}