Pith. sign in

Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods are mainly concerned with what and where the action is, which is unable to meet the requirements. Meanwhile, most of the existing datasets lack the labels indicating the degree of action standardization, and the action quality assessment datasets lack explainability and detailed feedback. Therefore, we define a new Human Action Form Assessment (AFA) task, and introduce a new diverse dataset CoT-AFA, which contains a large scale of fitness and martial arts videos with multi-level annotations for comprehensive video analysis. We enrich the CoT-AFA dataset with a novel Chain-of-Thought explanation paradigm. Instead of offering isolated feedback, our explanations provide a complete reasoning process--from identifying an action step to analyzing its outcome and proposing a concrete solution. Furthermore, we propose a framework named Explainable Fitness Assessor, which can not only judge an action but also explain why and provide a solution. This framework employs two parallel processing streams and a dynamic gating mechanism to fuse visual and semantic information, thereby boosting its analytical capabilities. The experimental results demonstrate that our method has achieved improvements in explanation generation (e.g., +16.0% in CIDEr), action classification (+2.7% in accuracy) and quality assessment (+2.1% in accuracy), revealing great potential of CoT-AFA for future studies. Our dataset and source code is available at https://github.com/MICLAB-BUPT/EFA.

fields

cs.CV 3

years

2026 3

verdicts

UNVERDICTED 3

representative citing papers

A DVDrive Approach for doScenes Instructed Driving Challenge

cs.CV · 2026-06-19 · unverdicted · novelty 3.0

The submission adapts OmniDrive with a DVPE-style divided-view perception module to enhance instruction-conditioned ego trajectory prediction on nuScenes scenes for the doScenes challenge.

Leveraging Metric Depth for Relative Depth Prediction

cs.CV · 2026-06-09 · unverdicted · novelty 2.0

Competition solution applies zero-shot pretrained models for metric depth to achieve relative depth prediction in football scenes with limited data, scoring 2.68e-3.

citing papers explorer

Showing 3 of 3 citing papers.

  • A DVDrive Approach for doScenes Instructed Driving Challenge cs.CV · 2026-06-19 · unverdicted · none · ref 16 · internal anchor

    The submission adapts OmniDrive with a DVPE-style divided-view perception module to enhance instruction-conditioned ego trajectory prediction on nuScenes scenes for the doScenes challenge.

  • Leveraging Metric Depth for Relative Depth Prediction cs.CV · 2026-06-09 · unverdicted · none · ref 13 · internal anchor

    Competition solution applies zero-shot pretrained models for metric depth to achieve relative depth prediction in football scenes with limited data, scoring 2.68e-3.

  • A VideoMAE-v2 Approach to Zero-Shot Traffic Accident Anticipation cs.CV · 2026-06-08 · unverdicted · none · ref 15 · internal anchor

    VideoMAE-v2 backbone with per-frame head achieves 2nd place in 2026 CVPR zero-shot traffic accident anticipation competition by training solely on public binary-labeled data.