Pith. sign in

REVIEW 1 cited by

VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17074 v7 pith:7N3HYAGA submitted 2023-11-27 cs.CV

classification cs.CV
keywords villsfeaturelearningre-identificationspatialconsistentdesignsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and imperfect feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross-Spectral Body Recognition with Side Information Embedding: Benchmarks on LLCM and Analyzing Range-Induced Occlusions on IJB-MDF

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Encoding camera identity, rather than spectral domain, in a side embedding reaches state-of-the-art visible-infrared person matching on LLCM, while cross-range matching on IJB-MDF reveals a large scale-induced perform...

Pith tools