No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training Techniques

Alexander Schwing; Derek Hoiem; Tanmay Gupta

arxiv: 1811.05967 · v2 · pith:DXZDTXG2new · submitted 2018-11-14 · 💻 cs.CV

No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training Techniques

Tanmay Gupta , Alexander Schwing , Derek Hoiem This is my paper

classification 💻 cs.CV

keywords trainingdetectionlayouttechniquesappearanceapproachesencodingsfactors

0 comments

read the original abstract

We show that for human-object interaction detection a relatively simple factorized model with appearance and layout encodings constructed from pre-trained object detectors outperforms more sophisticated approaches. Our model includes factors for detection scores, human and object appearance, and coarse (box-pair configuration) and optionally fine-grained layout (human pose). We also develop training techniques that improve learning efficiency by: (1) eliminating a train-inference mismatch; (2) rejecting easy negatives during mini-batch training; and (3) using a ratio of negatives to positives that is two orders of magnitude larger than existing approaches. We conduct a thorough ablation study to understand the importance of different factors and training techniques using the challenging HICO-Det dataset.

This paper has not been read by Pith yet.

No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training Techniques

discussion (0)