MAVIS introduces a multi-agent framework that parses videos into a structured semantic library and uses logic-aware debate among agents to retrieve relevant videos competitively without task-specific fine-tuning.
arXiv preprint arXiv:2407.12508 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
UNVERDICTED 2representative citing papers
CDGLT achieves SOTA on MET-Meme for multimodal metaphor identification by using SLERP-based concept drift and prompt-adapted LayerNorm tuning with reduced compute.
citing papers explorer
-
MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding
MAVIS introduces a multi-agent framework that parses videos into a structured semantic library and uses logic-aware debate among agents to retrieve relevant videos competitively without task-specific fine-tuning.
-
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
CDGLT achieves SOTA on MET-Meme for multimodal metaphor identification by using SLERP-based concept drift and prompt-adapted LayerNorm tuning with reduced compute.