← back to paper
arxiv: 2509.08777 · 2 revisions
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles