Pith. sign in

REVIEW 4 cited by

Putting the Object Back into Video Object Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12982 v2 pith:NLZ3O65F submitted 2023-10-19 cs.CV

classification cs.CV
keywords objectcutiememorysegmentationreadingvideobackbottom-up
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Cutie, a video object segmentation (VOS) network with object-level memory reading, which puts the object representation from memory back into the video object segmentation result. Recent works on VOS employ bottom-up pixel-level memory reading which struggles due to matching noise, especially in the presence of distractors, resulting in lower performance in more challenging data. In contrast, Cutie performs top-down object-level memory reading by adapting a small set of object queries. Via those, it interacts with the bottom-up pixel features iteratively with a query-based object transformer (qt, hence Cutie). The object queries act as a high-level summary of the target object, while high-resolution feature maps are retained for accurate segmentation. Together with foreground-background masked attention, Cutie cleanly separates the semantics of the foreground object from the background. On the challenging MOSE dataset, Cutie improves by 8.7 J&F over XMem with a similar running time and improves by 4.2 J&F over DeAOT while being three times faster. Code is available at: https://hkchengrex.github.io/Cutie

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Enabling Extensible Embodied Capabilities with Tools

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    Introduces Embodied Tool Protocol and tool externalization to improve embodied AI performance on perception and cognition tasks, with measured gains but limits on execution capabilities.

  2. SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    SigLoMa enables dynamic loco-manipulation on quadrupeds from ego-centric 5 Hz vision alone by using Sigma Points for scalable exteroception, an ego-centric Kalman Filter for high-rate state estimation, and an active s...

  3. ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning

    cs.RO 2026-07 conditional novelty 5.0 of 10

    ACE decouples LLM workflow reasoning from a mask-conditioned pick-and-place policy, reaching 50–70% success on zero-shot complex tabletop tasks where end-to-end baselines score 0%.

  4. 4D Vessel Reconstruction for Benchtop Thrombectomy Analysis

    eess.IV 2026-04 conditional novelty 5.0 of 10

    A nine-camera multi-view workflow with 4D Gaussian Splatting reconstructs dynamic vessel surfaces in thrombectomy phantoms to enable standardized comparative displacement and stress-proxy tracking.

Pith tools