Pith. sign in

REVIEW 2 cited by

CoLo-CAM: Class Activation Mapping for Object Co-Localization in Weakly-Labeled Unconstrained Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.09044 v5 pith:QE6FHIQY submitted 2023-03-16 cs.CV

classification cs.CV
keywords co-localizationframesobjectwsvolactivationcolorinformationlocalization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Leveraging spatiotemporal information in videos is critical for weakly supervised video object localization (WSVOL) tasks. However, state-of-the-art methods only rely on visual and motion cues, while discarding discriminative information, making them susceptible to inaccurate localizations. Recently, discriminative models have been explored for WSVOL tasks using a temporal class activation mapping (CAM) method. Although their results are promising, objects are assumed to have limited movement from frame to frame, leading to degradation in performance for relatively long-term dependencies. This paper proposes a novel CAM method for WSVOL that exploits spatiotemporal information in activation maps during training without constraining an object's position. Its training relies on Co-Localization, hence, the name CoLo-CAM. Given a sequence of frames, localization is jointly learned based on color cues extracted across the corresponding maps, by assuming that an object has similar color in consecutive frames. CAM activations are constrained to respond similarly over pixels with similar colors, achieving co-localization. This improves localization performance because the joint learning creates direct communication among pixels across all image locations and over all frames, allowing for transfer, aggregation, and correction of localizations. Co-localization is integrated into training by minimizing the color term of a conditional random field (CRF) loss over a sequence of frames/CAMs. Extensive experiments on two challenging YouTube-Objects datasets of unconstrained videos show the merits of our method, and its robustness to long-term dependencies, leading to new state-of-the-art performance for WSVOL task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal-consistent CAMs for Weakly Supervised Video Segmentation in Waste Sorting

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Training a before/after classifier with a temporal reconstruction loss on motion-compensated saliency maps improves weakly supervised video segmentation of waste on a conveyor belt.

  2. Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A multi-task CLIP model with LoRA, multi-level CLS fusion, and spherical feature interpolation reports 95% accuracy and strong retrieval scores on the ENTRep endoscopy benchmark.

Pith tools