Reflect-R1 introduces the first evidence-driven self-correction framework for long video understanding using a three-stage pipeline, stage-decoupled RL via SD-GRPO, and a 120K dataset to achieve SOTA on VideoMME and LongVideoBench.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
AnchorPrune prunes visual tokens by first selecting a protected query-relevance anchor and then greedily adding important, non-redundant context, preserving up to 97.6% of full-token accuracy with only 160 of 2,880 tokens.
citing papers explorer
-
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Reflect-R1 introduces the first evidence-driven self-correction framework for long video understanding using a three-stage pipeline, stage-decoupled RL via SD-GRPO, and a 120K dataset to achieve SOTA on VideoMME and LongVideoBench.
-
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
AnchorPrune prunes visual tokens by first selecting a protected query-relevance anchor and then greedily adding important, non-redundant context, preserving up to 97.6% of full-token accuracy with only 160 of 2,880 tokens.
- Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising