Pith. sign in

REVIEW 1 cited by

ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.00383 v2 pith:GCIJPJ2U submitted 2022-07-01 cs.CV cs.IR

classification cs.CVcs.IR
keywords videochallengelanguagequeriessubmissionclipego4dnatural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this report, we present the ReLER@ZJU-Alibaba submission to the Ego4D Natural Language Queries (NLQ) Challenge in CVPR 2022. Given a video clip and a text query, the goal of this challenge is to locate a temporal moment of the video clip where the answer to the query can be obtained. To tackle this task, we propose a multi-scale cross-modal transformer and a video frame-level contrastive loss to fully uncover the correlation between language queries and video clips. Besides, we propose two data augmentation strategies to increase the diversity of training samples. The experimental results demonstrate the effectiveness of our method. The final submission ranked first on the leaderboard.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GazeNLQ @ Ego4D Natural Language Queries Challenge 2025

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GazeNLQ adds contrastively pretrained gaze embeddings to a GroundNLQ-style grounding model, reporting 27.82 R1@0.3 on the Ego4D NLQ test split only when ensembled with GroundVQA.

Pith tools