Pith. sign in

REVIEW 2 cited by

Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.13369 v2 pith:26VAQ7I4 submitted 2019-07-31 cs.CV

Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video Recognition

classification cs.CV
keywords framesamplingrecognitionnetworkuntrimmedvideoclassificationgreat
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Video Recognition has drawn great research interest and great progress has been made. A suitable frame sampling strategy can improve the accuracy and efficiency of recognition. However, mainstream solutions generally adopt hand-crafted frame sampling strategies for recognition. It could degrade the performance, especially in untrimmed videos, due to the variation of frame-level saliency. To this end, we concentrate on improving untrimmed video classification via developing a learning-based frame sampling strategy. We intuitively formulate the frame sampling procedure as multiple parallel Markov decision processes, each of which aims at picking out a frame/clip by gradually adjusting an initial sampling. Then we propose to solve the problems with multi-agent reinforcement learning (MARL). Our MARL framework is composed of a novel RNN-based context-aware observation network which jointly models context information among nearby agents and historical states of a specific agent, a policy network which generates the probability distribution over a predefined action space at each step and a classification network for reward calculation as well as final recognition. Extensive experimental results show that our MARL-based scheme remarkably outperforms hand-crafted strategies with various 2D and 3D baseline methods. Our single RGB model achieves a comparable performance of ActivityNet v1.3 champion submission with multi-modal multi-model fusion and new state-of-the-art results on YouTube Birds and YouTube Cars.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A training-free frame-selection method that weights frame embeddings by query relevance and maximizes the selected subspace's volume improves keyframe recall and VQA accuracy across eight MLLMs on Video-MME and LongVi...

  2. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A query-conditioned, training-free frame selector that unifies relevance and diversity into a single volume-maximization objective improves keyframe recall and long-video question-answering accuracy.