R²ScP recovers missing audio-visual data in question answering by retrieving semantically consistent examples and purifying noise, outperforming generative imputation in incomplete scenarios.
InProceed- ings of the AAAI Conference on Artificial Intelligence, volume 39, pages 10483–10491
4 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 4years
2026 4representative citing papers
MAVIS introduces a multi-agent framework that parses videos into a structured semantic library and uses logic-aware debate among agents to retrieve relevant videos competitively without task-specific fine-tuning.
RSICCLLM introduces a post-training framework with RSICI dataset, difference-aware supervised fine-tuning, and dual-negative preference optimization that claims to outperform much larger models on remote sensing image change captioning.
citing papers explorer
-
Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
R²ScP recovers missing audio-visual data in question answering by retrieving semantically consistent examples and purifying noise, outperforming generative imputation in incomplete scenarios.
-
MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding
MAVIS introduces a multi-agent framework that parses videos into a structured semantic library and uses logic-aware debate among agents to retrieve relevant videos competitively without task-specific fine-tuning.
-
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
RSICCLLM introduces a post-training framework with RSICI dataset, difference-aware supervised fine-tuning, and dual-negative preference optimization that claims to outperform much larger models on remote sensing image change captioning.
- AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition