{"id":"d55aae7a-c508-4b6e-908f-d2ba977b23af","arxiv_id":"2502.02817","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of Action Quality Assessment organizes the past decade of research into 7 trends, 9 dataset domains, and performance comparisons across 195 papers.","lead":"This paper surveys the field of Action Quality Assessment, reviewing 195 papers, 26 datasets, and 7 research trends from the past decade. It aims to serve as a structured entry point for researchers and practitioners working on scoring human actions in sports, healthcare, and industry.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'largest systematic survey' claim rests on a PRISMA search that is not reproducible: the manuscript omits the query, inclusion criteria, and flow details, and says 'over 200' papers in the abstract while reporting 195 included; completeness is therefore unverifiable.","rationale":"The reader identified the right load-bearing assumption: completeness and unbiasedness of the PRISMA search. I agree, and the manuscript text provides further evidence. The survey has independent strengths: it catalogues 26 datasets, organizes the literature into seven trends, and includes performance tables that are useful entry points. But the headline claim is quantitative and comparative, and the reporting does not support it. PRISMA (ref [20]) is invoked, yet the mandatory flow diagram and explicit eligibility criteria are absent; instead, Section 1 gives only a four-stage count summary. The abstract and Section 6 say 'over 200' papers, while Section 1 says 195 included, an internal inconsistency in the central count. The existence of two other 2024 surveys ([216], [217]) means 'largest' cannot be accepted on faith; it must be demonstrated by coverage comparison. The EgoExo4D omission in the dataset table reinforces that the inclusion rule is unclear. None of this requires doubting author integrity; it simply means the paper's central claim is not yet supported by the evidence it reports. If the authors release the search protocol, reconcile counts, and provide a coverage comparison, the survey could support its claim; until then, conditional acceptance is appropriate. This matches the reader's verdict, so no change is needed.","tokens_in":44976,"tokens_out":6885,"duration_ms":59760,"concrete_test":"Take the union of all AQA papers cited in prior surveys [18], [19], [216], [217] and check whether this survey's 195 included papers (or its bibliography) contains all of them. Any relevant AQA paper from that union that is missing without a documented exclusion reason refutes the 'comprehensive' claim; for 'largest', the survey must also contain additional relevant papers beyond that union. This coverage-overlap test is the decisive check of the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 6) is that this is 'the largest and most comprehensive survey to date' on AQA, supported by the PRISMA pipeline in Section 1: 505 papers identified from Web of Science and Google Scholar, 276 after title/abstract screening, 195 after full-text eligibility. The attack is that this process is not auditable, so the completeness claim cannot be checked. The manuscript gives neither the exact search strings nor the per-database date ranges, and it does not state the inclusion/exclusion criteria used at screening or eligibility stages. Without these, 'largest' and 'comprehensive' are assertions rather than demonstrated properties. The internal count mismatch makes this concrete: the abstract and Section 6 claim 'over 200' papers were reviewed, but Section 1 reports 195 included papers. Additionally, the paper cites two other 2024 surveys ([216], [217]) but provides no coverage comparison against them, so the comparative 'largest' claim is unsupported. A related inventory inconsistency is that Section 3.3 introduces EgoExo4D [223] as an AQA-relevant dataset, yet Table 1 lists exactly 26 datasets and omits it. These points do not undermine the survey's useful taxonomy, dataset summaries, or trend analysis, but they do block verification of the headline claim and should be resolved before the paper is used as the authoritative systematic review.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims to present the largest and most comprehensive systematic survey of Action Quality Assessment (AQA) to date, using the PRISMA framework to review papers published up to December 2024. It proposes definitions, metrics, a dataset inventory of 26 datasets in 9 domains, a taxonomy of research methods organized into 7 principal trends since 2014, and a discussion of challenges and future directions. The survey's central value would lie in a complete, auditable map of the AQA literature, but as submitted the completeness claim is not verifiable from the reported methodology.","tokens_in":45268,"tokens_out":5010,"duration_ms":49523,"significance":"If the search process is made auditable and the inventory is corrected, this survey would be a useful structured reference for AQA researchers: it consolidates dataset descriptions, provides comparative performance tables, organizes the methodological literature into seven trends, and makes its data available at a public URL. The PRISMA-based claim is an explicit strength in principle, but the missing search protocol and internal inconsistencies currently prevent the reader from checking the headline completeness claim.","major_comments":[{"comment":"The paper reports inconsistent inclusion counts. The abstract and Section 6 say 'over 200' papers were systematically reviewed, while Section 1 reports that 195 papers met the inclusion criteria. If the 200+ figure refers to identified or screened papers rather than included papers, that should be stated explicitly; as written, the discrepancy undermines the precision of the headline claim and should be reconciled.","section":"Abstract, §1, §6"},{"comment":"The PRISMA process is not reproducible. The manuscript reports only the keyword 'action quality assessment' and the sources Web of Science and Google Scholar, with counts 505, 276, and 195, but omits the exact search strings, per-database date ranges, deduplication method, and the inclusion/exclusion criteria used at the title/abstract and full-text stages. A PRISMA flow diagram and a search protocol (in the paper or in the linked data repository) are needed to support the claim that the survey covers the full AQA literature rather than a convenience sample.","section":"§1 (PRISMA pipeline)"},{"comment":"The dataset inventory is internally inconsistent. Section 3.3 introduces EgoExo4D as an AQA-relevant dataset with 1224 samples and 16 hours of video, but Table 1, which is presented as the complete list of 26 publicly available datasets, does not include EgoExo4D. Either the dataset should be added to Table 1 and the counts and domain summaries updated, or its inclusion in Section 3.3 should be explicitly justified as outside the table's scope.","section":"§3.3 and Table 1"},{"comment":"The 'largest and most comprehensive survey to date' claim is unsupported by comparison with the two 2024 surveys cited as [216] and [217]. The manuscript does not state how many papers, datasets, or time periods those surveys cover, nor does it quantify the overlap or additionally covered items. The claim should be either substantiated with a concrete coverage comparison or softened to a description of scope.","section":"§6 and references [216],[217]"}],"minor_comments":[{"comment":"The text says Pirsiavash et al. 'first pioneered the assessment of action quality in 2014', which conflicts with Section 4.1's statement that AQA research dates back to Gordon in 1995; the sentence should be qualified as referring to modern deep-learning-based AQA or the first AQA dataset.","section":"§2.4"},{"comment":"The cross-reference to 'Tab. 6' for the objective-evaluation methods appears to be incorrect: the objective-evaluation performance table is labeled Table 5, while Table 6 covers asymmetric relationships.","section":"§4.2.1(3)"},{"comment":"'NeurlIPS 2024' should read 'NeurIPS 2024'.","section":"§3.8"},{"comment":"'comrehensively' should be 'comprehensively'.","section":"§5.2"},{"comment":"'withI3D-T ransformer decoder' has a spacing typo and should be 'with I3D-Transformer decoder'.","section":"§4.1"},{"comment":"The year axis appears to skip 2016, showing 14, 15, 17, 18, ...; please check whether this is a rendering issue or a labeling error.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The substantial number of self-citations by one co-author is not, by itself, circular, and several of those works are genuine pillars of the AQA literature. However, the evaluative language around those works and the lack of coverage comparison with the two recent surveys [216,217] should be scrutinized editorially. The main grounds for major revision are the non-reproducible PRISMA protocol, the inclusion-count inconsistency, and the EgoExo4D omission from Table 1; all are fixable within the paper's current scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the most useful AQA survey I have seen to date: it organizes the field into seven research trends and nine dataset domains, compiles performance tables, and covers work through December 2024, including AI-generated video evaluation. The taxonomy is sensible, the dataset summaries are practical, and the methodology timeline gives newcomers a real map of how the field evolved. That is a genuine contribution even though the paper contains no new methods or data.\n\nThe soft spots are real but concentrated in the systematic-review apparatus. The abstract says 'over 200' papers were reviewed; Section 1 reports 195 included papers. That is a concrete inconsistency. More importantly, the PRISMA pipeline is not auditable: no search strings, no per-database date ranges, no screening or eligibility criteria, no flow diagram. Without those, the 'largest and most comprehensive' claim is an assertion, not a demonstrated property. The paper also cites two competing 2024 surveys but never compares coverage against them, which weakens the comparative claim. The EgoExo4D mention in Section 3.3 without a Table 1 entry is a further inventory inconsistency. These are fixable, but they block verification of the headline claim.\n\nThe internal cross-reference errors, like swapped table references, look like copyediting rather than deep problems. The self-citations to Parmar are not themselves a flaw: he is a central contributor in AQA, and the survey's utility does not depend on those being fewer. The real issue is that the search process is undocumented, not who appears in the reference list.\n\nBottom line: the central taxonomy, dataset tables, and trend analysis hold up. This paper deserves a serious referee and likely publication after revision. I would ask the authors to release a PRISMA flow diagram and exact inclusion criteria, reconcile the 195/over-200 counts, fix the dataset inconsistency, and either add coverage comparisons with the other surveys or soften the 'largest' claim to match what is actually documented. If that happens, I would cite this as the standard entry point to AQA.","headline":"A genuinely useful AQA survey whose 'largest systematic review' claim is not yet auditable; worth serious refereeing after PRISMA transparency and consistency fixes.","tokens_in":45771,"tokens_out":1886,"would_cite":true,"duration_ms":21340,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the largest and most systematic survey of action quality assessment to date, reviewing over 200 papers through the PRISMA framework and organizing 26 datasets and 7 research trends.","keywords":["action quality assessment","skills assessment","video understanding","computer vision","systematic review","PRISMA","deep learning","survey"],"falsifier":"Run the same PRISMA search on Web of Science and Google Scholar for 'action quality assessment' with an explicit query, per-database date ranges, and documented inclusion and exclusion criteria; if the reproducible flow yields substantially more than 195 qualifying papers, or if tracing the reported 505 to 276 to 195 numbers proves impossible because no query can reproduce them, the survey's 'largest to date' claim is not settled. A quicker check: the abstract says 'over 200 research papers' while the method section says 195 papers met inclusion criteria, so reconciling that count is a concrete first test.","tokens_in":44785,"feed_emoji":"📊","tokens_out":5608,"duration_ms":48253,"temperature":0.7,"pith_summary":"Action quality assessment (AQA) asks how well a person performs an action, not just what action it is. This paper claims to be the largest and most systematic survey of the field to date, reviewing over 200 research papers under the PRISMA framework through December 2024. It organizes 26 public datasets into 9 application domains, distills a decade of methodology into 7 research trends, and lays out challenges and future directions. A reader would care because AQA underpins physiotherapy, sports training, surgical skill evaluation, and quality control for AI-generated video, and a structured map of the field makes its progress and open problems visible.","feed_headline":"Survey maps a decade of action quality research","feed_subtitle":"PRISMA review of 200+ papers organizes 26 datasets and 7 methodology trends.","key_machinery":"The load-bearing apparatus is the PRISMA systematic-review framework, applied as a four-stage pipeline: identification (keyword search in Web of Science and Google Scholar), screening (title/abstract deduplication and relevance filtering), eligibility (full-text quality and novelty review), and inclusion (195 papers, 26 datasets). The survey's organizing grid is a 2D taxonomy: 9 dataset domains by application scenario and 7 methodology trends by research objective, with performance tables comparing representative models under Spearman's rank correlation, relative $\\ell^2$ distance, or accuracy. This machinery is what converts a literature collection into the claimed systematic map.","core_discovery":"The paper's central claim is that it provides the most complete structured synthesis of AQA research produced so far. Concretely, it reports a four-stage PRISMA selection process that moved from 505 candidate records to 276 screened papers to 195 included papers (with 26 datasets), spans work from the first AQA datasets in 2014 through December 2024, and organizes the included literature into 7 principal trends: fine-grained analysis, multitask and multimodal methods, generalization, continual learning, explainability, comprehensive assessment, and self-supervised representation learning. It also classifies datasets into 9 domains (surgery, rehabilitation, daily activities, music, fitness, industrial manufacturing, dance, AI-generated video, and sports) and identifies persistent challenges in actions, datasets, and methodologies. If this claim is correct, the survey is a reference map that researchers can use to locate methods, benchmarks, and open problems in AQA.","pith_inferences":["Because the survey stops at December 2024 and cites two other December 2024 surveys in the same space, the 'largest to date' status is time-sensitive; an updated or living review would be needed to keep the map authoritative.","The proposed future direction of using text-to-video models to generate large-scale multi-action AQA datasets is testable: one could generate prompt-conditioned videos with known skill levels and measure whether models trained on them transfer to real-world judging benchmarks.","The 7-trend taxonomy could serve as a coding scheme for a follow-up meta-analysis that tracks the share of AQA papers per trend over time, turning the survey's qualitative 'rising trend' statements into quantitative evidence.","If gold-standard action samples become established, reference-based contrastive regression methods could be re-evaluated against that fixed anchor rather than against in-dataset best samples."],"forward_implications":["The field has moved from coarse score regression toward fine-grained, explainable, and multimodal assessment; the survey's 7-trend taxonomy makes that trajectory explicit.","Researchers and practitioners can use the 26-dataset, 9-domain inventory to choose benchmarks for sports, surgery, rehabilitation, fitness, dance, music, manufacturing, daily activities, or AI-generated video.","The survey's challenge analysis points to concrete next steps: larger multi-action real-world datasets, AI-generated multi-action datasets built with text-to-video models, and gold-standard action samples for reference-based scoring.","Methodology gaps identified, especially real-time lightweight models, interpretability, and robustness to missing modalities, define near-term research targets.","Standard metrics (SRC, R-l2, accuracy) are consolidated with formulas and usage guidance, supporting cross-paper comparisons."],"supporting_citations":[{"why":"The PRISMA 2020 statement supplies the systematic-review methodology and reporting structure the survey follows.","marker":"[20]"},{"why":"Introduced the first AQA datasets, the supervised regression formulation, and Spearman's Rank Correlation as the field's primary metric.","marker":"[2]"},{"why":"Established deep spatiotemporal feature extraction with C3D backbones and the UNLV-Dive and UNLV-Vault datasets that many later methods build on.","marker":"[3]"},{"why":"Provided JIGSAWS, the first surgical skill dataset, and anchors the surgery domain.","marker":"[4]"},{"why":"Introduced MTL-AQA with fine-grained action breakdowns, text descriptions, and the multi-task learning paradigm.","marker":"[21]"},{"why":"Provided Fitness-AQA, the largest AQA dataset, and self-supervised representations for workout form assessment.","marker":"[8]"},{"why":"Provided GAIA, the first AQA dataset for AI-generated videos, used to argue AQA can evaluate text-to-video generation quality.","marker":"[13]"},{"why":"Provided FineDiving, the first fine-grained sports video dataset with procedure segmentation, underpinning the fine-grained trend.","marker":"[38]"}],"fun_headline_variants":["Decade of action quality: 200+ papers, 26 datasets, 7 trends","Largest AQA survey yet: PRISMA review of 200+ papers over 10 years","Action quality mapped: 7 trends, 9 domains, and lingering challenges","From surgery to sports: 200+ AQA papers organized into 7 trends","AQA's decade: 200+ papers, 26 datasets, 7 key research trends"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the PRISMA literature search was complete and unbiased; the paper reports only the headline numbers (505 to 276 to 195 papers) without the search strings, databases' date ranges, or inclusion and exclusion criteria, so a reader cannot verify that the survey covers the full AQA literature rather than a convenience sample.","fun_headline_variants_meta":{"raw":{"variants":["Decade of action quality: 200+ papers, 26 datasets, 7 trends","Largest AQA survey yet: PRISMA review of 200+ papers over 10 years","Action quality mapped: 7 trends, 9 domains, and lingering challenges","From surgery to sports: 200+ AQA papers organized into 7 trends","AQA's decade: 200+ papers, 26 datasets, 7 key research trends"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001198,"raw_usage":{"total_tokens":4933,"prompt_tokens":933,"completion_tokens":4000,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":3886}},"tokens_in":549,"tokens_out":4000,"duration_ms":28985,"temperature":1.0,"reasoning_tokens":3886,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T10:58:56.415690+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same PRISMA search on Web of Science and Google Scholar for 'action quality assessment' with an explicit query, per-database date ranges, and documented inclusion and exclusion criteria; if the reproducible flow yields substantially more than 195 qualifying papers, or if tracing the reported 505 to 276 to 195 numbers proves impossible because no query can reproduce them, the survey's 'largest to date' claim is not settled. A quicker check: the abstract says 'over 200 research papers' while the method section says 195 papers met inclusion criteria, so reconciling that count is a concrete first test.","supporting_citations":[],"review_version":1}