RRG trains MLLMs via a reinforced multimodal reference game with contrastive rewards on hard positives and negatives to produce accurate, discriminative concept descriptions, achieving SOTA on personalization benchmarks.
Personalized large vision-language models
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3years
2026 3roles
background 1polarities
background 1representative citing papers
ICPT converts a few reference images of a personalized concept into an adaptive-length visual prompt plus a label embedding, letting a frozen LVLM add and reason about multiple concepts on the fly.
Introduces Personal VCL formalization and benchmark revealing LMM context gaps, plus an Agentic Context Bank baseline that boosts personalized visual reasoning.
citing papers explorer
-
Personalizing MLLMs via Reinforced Multimodal Reference Game
RRG trains MLLMs via a reinforced multimodal reference game with contrastive rewards on hard positives and negatives to produce accurate, discriminative concept descriptions, achieving SOTA on personalization benchmarks.
-
Personalize Your Large Vision-language Models With In-context Prompt Tuning
ICPT converts a few reference images of a personalized concept into an adaptive-length visual prompt plus a label embedding, letting a frozen LVLM add and reason about multiple concepts on the fly.
-
Personal Visual Context Learning in Large Multimodal Models
Introduces Personal VCL formalization and benchmark revealing LMM context gaps, plus an Agentic Context Bank baseline that boosts personalized visual reasoning.