An end-to-end transformer couples detection with a temporal matching penalty and feedback queries, boosting temporal consistency of scene-graph predictions on Action Genome, OpenPVSG, and MEVA.
BLoad: Enhancing Neural Network Training with Efficient Sequential Data Handling
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The increasing complexity of modern deep neural network models and the expanding sizes of datasets necessitate the development of optimized and scalable training methods. In this white paper, we addressed the challenge of efficiently training neural network models using sequences of varying sizes. To address this challenge, we propose a novel training scheme that enables efficient distributed data-parallel training on sequences of different sizes with minimal overhead. By using this scheme we were able to reduce the padding amount by more than 100$x$ while not deleting a single frame, resulting in an overall increased performance on both training time and Recall in our experiments.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Temporally Consistent Dynamic Scene Graphs: An End-to-End Approach for Action Tracklet Generation
An end-to-end transformer couples detection with a temporal matching penalty and feedback queries, boosting temporal consistency of scene-graph predictions on Action Genome, OpenPVSG, and MEVA.