In multimodal KB-VQA, gold evidence at the first prompt slot beats gold at the last by 16–26 points, flipping the classic U-shaped lost-in-the-middle pattern into primacy bias.
arXiv preprint arXiv:2505.24073 (2025),https://arxiv.org/abs/2505.24073
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
background 1representative citing papers
Geometry-indexed depth–pose retrieval plus schema SFT and GRPO planning improves faithfulness, consistency, and controllability of plot-to-short-drama video generation over multi-agent and text-only baselines.
A lightweight hierarchical multimodal graph RAG that fuses entity-grounded visual objects with text nodes and propagates relevance via multi-granularity PPR, delivering SOTA multimodal task performance at far lower construction cost.
citing papers explorer
-
Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
In multimodal KB-VQA, gold evidence at the first prompt slot beats gold at the last by 16–26 points, flipping the classic U-shaped lost-in-the-middle pattern into primacy bias.
-
DramaDirector: Geometry-Guided Short Drama Generation
Geometry-indexed depth–pose retrieval plus schema SFT and GRPO planning improves faithfulness, consistency, and controllability of plot-to-short-drama video generation over multi-agent and text-only baselines.
-
MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation
A lightweight hierarchical multimodal graph RAG that fuses entity-grounded visual objects with text nodes and propagates relevance via multi-granularity PPR, delivering SOTA multimodal task performance at far lower construction cost.