Pith. sign in

Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.CV 4

years

2026 3 2025 1

verdicts

UNVERDICTED 4

representative citing papers

Adapting MLLMs for Nuanced Video Retrieval

cs.CV · 2025-12-15 · unverdicted · novelty 7.0

Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.

DiLA: Disentangled Latent Action World Models

cs.CV · 2026-05-15 · unverdicted · novelty 6.0

DiLA uses content-structure disentanglement driven by predictive bottlenecks to create semantically structured latent actions for high-fidelity video world models.

citing papers explorer

Showing 4 of 4 citing papers.

  • VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection cs.CV · 2026-04-23 · unverdicted · none · ref 38 · 2 links

    VFM4SDG is a dual-prior framework that distills cross-domain stable relations from VFMs into DETR encoders and injects semantic-contextual priors into decoder queries to reduce missed detections in single-domain generalized object detection.

  • Adapting MLLMs for Nuanced Video Retrieval cs.CV · 2025-12-15 · unverdicted · none · ref 66

    Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.

  • Continual Visual and Verbal Learning Through a Child's Egocentric Input cs.CV · 2026-06-03 · unverdicted · none · ref 81

    BabyCL learns word-referent mappings from egocentric video in a single chronological pass via streaming visual learning, dual replay, and three contrastive losses, outperforming streaming baselines on the SAYCam 4AFC benchmark.

  • DiLA: Disentangled Latent Action World Models cs.CV · 2026-05-15 · unverdicted · none · ref 25

    DiLA uses content-structure disentanglement driven by predictive bottlenecks to create semantically structured latent actions for high-fidelity video world models.