Pith. sign in

REVIEW 1 cited by

Towards Good Practices for Missing Modality Robust Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.13916 v2 pith:D7ZBWMLB submitted 2022-11-25 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords missingmodalitymodalitiesmodelmulti-modalactiongoodinference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Standard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances degrade drastically if any modality is missing in the inference stage. We ask: how can we train a model that is robust to missing modalities? This paper seeks a set of good practices for multi-modal action recognition, with a particular interest in circumstances where some modalities are not available at an inference time. First, we study how to effectively regularize the model during training (e.g., data augmentation). Second, we investigate on fusion methods for robustness to missing modalities: we find that transformer-based fusion shows better robustness for missing modality than summation or concatenation. Third, we propose a simple modular network, ActionMAE, which learns missing modality predictive coding by randomly dropping modality features and tries to reconstruct them with the remaining modality features. Coupling these good practices, we build a model that is not only effective in multi-modal action recognition but also robust to modality missing. Our model achieves the state-of-the-arts on multiple benchmarks and maintains competitive performances even in missing modality scenarios. Codes are available at https://github.com/sangminwoo/ActionMAE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformer-based Multisensor Data Fusion of Ultrasonic Guided Wave and FBG-based Strain Measurements for Multitask Aerospace Structural Health Monitoring

    eess.SP 2026-06 conditional novelty 6.0 of 10

    Transformer fusion of asynchronous PZT guided-wave and FBG strain data yields HI MAE/RMSE <0.1 and localization MAE/RMSE <0.0465/0.1571, beating single-sensor and SOTA DNN baselines by ~60% on ReMAP composite fatigue panels.

Pith tools