Pith. sign in

REVIEW 2 cited by

XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07115 v2 pith:6QBODRVA submitted 2022-07-14 cs.CV

classification cs.CV
keywords memoryfeaturelong-termmodelxmematkinson-shiffrinobjectsegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present XMem, a video object segmentation architecture for long videos with unified feature memory stores inspired by the Atkinson-Shiffrin memory model. Prior work on video object segmentation typically only uses one type of feature memory. For videos longer than a minute, a single feature memory model tightly links memory consumption and accuracy. In contrast, following the Atkinson-Shiffrin model, we develop an architecture that incorporates multiple independent yet deeply-connected feature memory stores: a rapidly updated sensory memory, a high-resolution working memory, and a compact thus sustained long-term memory. Crucially, we develop a memory potentiation algorithm that routinely consolidates actively used working memory elements into the long-term memory, which avoids memory explosion and minimizes performance decay for long-term prediction. Combined with a new memory reading mechanism, XMem greatly exceeds state-of-the-art performance on long-video datasets while being on par with state-of-the-art methods (that do not work on long videos) on short-video datasets. Code is available at https://hkchengrex.github.io/XMem

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Time-Contrastive Pretraining for In-Context Image and Video Segmentation

    cs.CV 2025-06 reject novelty 5.0 of 10

    Time-contrastive prompt retrieval plus SAM2 video object segmentation yields strong FLARE 2022 Dice scores, but the reported gains are confounded by unequal fine-tuning and missing error bars.

  2. Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems

    cs.CV 2025-06 reject novelty 3.0 of 10

    A survey that organizes vision-language segmentation methods for intelligent transportation, but its synthesis is undermined by fabricated references and unverifiable benchmarks.

Pith tools