Pith. sign in

REVIEW 2 cited by

Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.10634 v2 pith:PC7P4OC5 submitted 2021-06-20 cs.CV

Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding

classification cs.CV
keywords augmentedd-tanproposeapproachboundingfirstgroundinghuman-centric
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We propose an effective two-stage approach to tackle the problem of language-based Human-centric Spatio-Temporal Video Grounding (HC-STVG) task. In the first stage, we propose an Augmented 2D Temporal Adjacent Network (Augmented 2D-TAN) to temporally ground the target moment corresponding to the given description. Primarily, we improve the original 2D-TAN from two aspects: First, a temporal context-aware Bi-LSTM Aggregation Module is developed to aggregate clip-level representations, replacing the original max-pooling. Second, we propose to employ Random Concatenation Augmentation (RCA) mechanism during the training phase. In the second stage, we use pretrained MDETR model to generate per-frame bounding boxes via language query, and design a set of hand-crafted rules to select the best matching bounding box outputted by MDETR for each frame within the grounded moment.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

    cs.CV 2026-07 conditional novelty 6.0

    A coarse-to-fine video-grounding framework improves temporal boundary accuracy by densely re-examining frames around coarse start/end predictions, achieving SOTA on HC-STVGv1/v2 and VidSTG.

  2. Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization

    cs.CV 2026-04 unverdicted novelty 5.0

    ST-GD adapts Grounding DINO with about 10 million trainable parameters via adapters and a temporal decoder to achieve competitive performance on limited-data spatio-temporal video grounding benchmarks.