Using a single Llama-3.2-11B-Vision-Instruct model per task with RAG, reranking, multi-task fine-tuning, and refusal-data augmentation, the solution ranked 1st on Task3 and 3rd on Tasks 1 and 2 in the CRAG-MM challenge.
Gonzalez, Ion Stoica, Sohier Dane, Maggie Demkin, and Nate Keating
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Solution for Meta KDD Cup'25: A Comprehensive Three-Step Framework for Vision Question Answering
Using a single Llama-3.2-11B-Vision-Instruct model per task with RAG, reranking, multi-task fine-tuning, and refusal-data augmentation, the solution ranked 1st on Task3 and 3rd on Tasks 1 and 2 in the CRAG-MM challenge.