Pith. sign in

REVIEW 1 cited by

RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08576 v1 pith:OFBHCHEM submitted 2025-03-11 cs.CV

classification cs.CV
keywords videosamplingbenchmarksrag-adapterunderstandinglongmllmstesting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal Large Language Models (MLLMs) capable of video understanding are advancing rapidly. To effectively assess their video comprehension capabilities, long video understanding benchmarks, such as Video-MME and MLVU, are proposed. However, these benchmarks directly use uniform frame sampling for testing, which results in significant information loss and affects the accuracy of the evaluations in reflecting the true abilities of MLLMs. To address this, we propose RAG-Adapter, a plug-and-play framework that reduces information loss during testing by sampling frames most relevant to the given question. Additionally, we introduce a Grouped-supervised Contrastive Learning (GCL) method to further enhance sampling effectiveness of RAG-Adapter through fine-tuning on our constructed MMAT dataset. Finally, we test numerous baseline MLLMs on various video understanding benchmarks, finding that RAG-Adapter sampling consistently outperforms uniform sampling (e.g., Accuracy of GPT-4o increases by 9.3 percent on Video-MME), providing a more accurate testing method for long video benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Event-Grounded Question Answering over Long Audio via Structured Retrieval

    eess.AS 2026-02 reject novelty 5.0 of 10

    LA-RAG stores timestamped audio events from a grounding model in SQL and uses intent-aware retrieval plus an LLM to answer long-audio questions, claiming 76.88% accuracy on synthetic home audio.

Pith tools