Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:10:00.626842Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2509.08777.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:10:00.626842Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:52:45.867032Z
A source-named dated measurement, never combined with another source.
Source: cited_works
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3608cf9-115f-4bf5-b4e5-07ebb093b7b4 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Spice: Semantic propositional image cap- tion evaluation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e6c8d3f1-6575-4cfb-8b47-53e70e1ee3fe · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Lawrence Zitnick, and Devi Parikh
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6bd06c55-e959-4e66-a6be-e108851d6205 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els.arXiv, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f656ce2-34cf-4ab0-88f8-96953b6a04cf · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 296632bc-bb85-455f-925a-bb7ee389111d · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles The control of the false discovery rate in multiple testing under dependency
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c3889dd4-8ff4-4d03-ae37-b04576e9e4d5 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improving image genera- tion with better captions.arXiv, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72389f87-8d8c-4384-844c-f9a1c4df1373 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Visit-bench: A benchmark for vision- language instruction following inspired by real-world use
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f9e2392-17a6-4326-a0ae-d838dfc86f8d · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f40a452-7d07-4743-8df3-d967181035f0 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6bce0c68-48d2-4d24-9aa6-6c113d35c710 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles X-iqe: explainable image quality evaluation for text-to-image generation with visual large language models.arXiv, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd27c001-1f68-470a-a1f1-bbc3fbfcdd66 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154966d0-46b3-4f1c-b0c8-985b62991e5a · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.ICCV, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c98be1fb-0aaa-4adc-b28f-389233eeb660 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles A coefficient of agreement for nominal scales
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db31190b-3c9a-4b8d-83aa-5880698cb02a · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Ex- ploring GPT-4 vision for text-to-image synthesis evaluation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4713f7d9-ba43-42c1-9b38-6cb3fab6fe3f · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Instructblip: Towards general- purpose vision-language models with instruction tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e9f98e67-98d7-424b-9c43-f51cbbc06d97 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improving selective visual question answering by learning from your peers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2d64d8c7-47ba-4d50-916b-1859b72ac210 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Moura, Devi Parikh, and Dhruv Batra
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2fc28f98-7f8e-4b43-be9b-78b3ffc050c4 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles The relationship be- tween precision-recall and ROC curves
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c1a58c9a-e552-4a14-8230-cdfd66fb20ca · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Mllm-bench, evaluating multi- modal llms using gpt-4v.arXiv, 2023
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf889458-d5e7-4a37-812d-ada0289494a3 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Selective classification for deep neural networks.NeurIPS, 30, 2017
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ac3fd95-a6bd-4817-a1f5-a02601c56ebf · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Practical variational inference for neural net- works.Advances in neural information processing systems, 24, 2011
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 50ed7c0d-0e39-4e1c-89b7-341d31b53deb · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles A Survey on LLM-as-a-Judge
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c043b7d-f93c-4c06-86aa-a88c76a3c2f3 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Weinberger
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1968aae-65ba-443b-abce-e9e7212bf6af · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Springer, New York, NY , USA, 2001
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 961b2e05-ad1e-4d19-babe-1ef2bd9d6add · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Glass, and Yulia Tsvetkov
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 50c20587-0fed-4899-9c3b-829f1045575b · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Machine learning with a re- ject option: A survey.Machine Learning, pages 1–38, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1535f5b7-42ce-4705-b41b-377cce914983 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4d1d5d0c-7ed4-4c99-a083-b21f4a11deb8 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Promptboosting: Black-box text classification with ten forward passes
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ee14fb5c-be1b-4b8e-969a-4c5db627c4b7 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Are llm-based evaluators con- fusing nlg quality criteria?arXiv, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d2fd8af-c33a-46a9-ba26-84f3af4004ca · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles GPT-4o System Card
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630922ed-d7f5-43b2-a4ff-fc218f9e3f37 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Calibrating language models via augmented prompt ensembles.ICML Workshop on Deployable Generative AI, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 63b39703-7552-43eb-a441-c6312f9fc3f2 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles How can we know what language models know?Trans- actions of the Association for Computational Linguistics, 8: 423–438, 2020
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4fcf8bfc-96e5-40bc-972c-0348b95f539c · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Billion- scale similarity search with GPUs.IEEE Transactions on Big Data, 2019
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b565dc-d1de-4c51-9258-28f5acd50047 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc03e82d-2585-4d28-8ac6-dfc2660a5cd1 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Benchmarking cog- nitive biases in large language models as evaluators.arXiv,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6457c77b-21a7-4ee5-9e07-6bd84e102aa9 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improved precision and recall met- ric for assessing generative models.Advances in neural in- formation processing systems, 32, 2019
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6730b91e-1eb6-47f2-9a98-f5c6c36a263e · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Seed-bench: Bench- marking multimodal large language models.CVPR, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aacab920-6bf6-4e5b-856f-6e58fd94d9bb · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles From generation to judgment: Op- portunities and challenges of llm-as-a-judge.arXiv preprint arXiv: 2411.16594, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1813db-8738-4375-b8de-e443211a01e9 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6dc6f29-878c-4124-9224-42e5c027a030 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles ROUGE: A package for automatic evaluation of summaries.Text Summarization Branches Out, 2004
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eedef4f5-10bf-4e1f-8883-18834e9609e5 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Microsoft coco: Common objects in context.ECCV,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 34c4dde9-8ae8-4694-b073-83a8c2f1213a · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Mitigating hallucination in large multi-modal models via robust instruction tuning.ICLR,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f9f3cc8-6c8c-4959-9956-418a5551f759 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Visual instruction tuning.NeurIPS, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 08920eaf-a063-4fbd-b6dd-bdec8601d243 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Llms as narcissistic evaluators: When ego inflates evaluation scores
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79738d18-7d1f-4dda-b52b-0a6dc496f15a · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Predict re- sponsibly: improving fairness and accuracy by learning to defer.NeurIPS, 31, 2018
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ff6be8c5-c650-4e3d-8f34-aa2d59ad8fbb · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gpt-4v(ision) technical work and authors.https: //openai.com/contributions/gpt-4v/, 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18587033-0c7e-4bae-b88d-3b52140a17ad · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Openai o1 system card.https://openai
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 81130e22-1f21-4b9b-9d3b-0b945da0266e · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Bowman, and Shi Feng
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 29d9ffa6-db7a-46ab-a660-9fa37b651757 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Bleu: a method for automatic evaluation of machine translation.ACL, 2002
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6302416-04e4-41e4-94bd-df1e4bafa69a · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Kosmos-2: Grounding multimodal large language models to the world.arXiv, 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9fe06e04-fb45-4444-b18f-cc46b54d7cbb · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Sdxl: improving latent diffusion models for high-resolution image synthesis
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6642fe0-938a-4e65-b5af-5f2aa834ca84 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd110cf1-eb76-4eb6-ac18-88681c72771d · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Learning transferable visual models from natural language supervi- sion
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acda0efd-f74e-452c-b6a4-2b9bf2eef5b0 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles High-resolution image synthesis with latent diffusion models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c673ea0-015c-43c3-9ff1-55243a793312 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd90b93e-f9d4-49d2-9b90-d0c73c5af6e5 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Verbosity Bias in Preference Labeling by Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f718aff-0726-4723-83f1-39f0762985ac · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improved techniques for training gans.Advances in neural information processing systems, 29, 2016
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a7356c-e99d-486f-9520-91e1f63a4e2f · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gemini: A Family of Highly Capable Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dac2553-0751-4480-9372-0f948155de8d · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Bayesian prompt ensembles: Model uncer- tainty estimation for black-box large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05aab997-1bb7-4874-83e8-5beefa44aae4 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Large language models are not fair evaluators.arXiv, 2023
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2312aa86-2571-4756-85a9-3ac0906d2d99 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46695519-4fe1-4fdf-bb09-bbff6705d6a6 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Reliable visual question answering: Abstain rather than answer incorrectly
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2790cf08-e40c-488b-ad53-ccdbcdb1149c · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Strength in numbers: Estimating confidence of large language models by prompt agreement
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bac3ba20-eb94-4ada-8c22-672fa9814ea6 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gpt- 4v(ision) is a human-aligned evaluator for text-to-3d genera- tion.CVPR, 2024
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 040f3b76-5409-42df-a13f-ef97dae5b022 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5d8eb4-7abe-4a9a-a27a-5c59e2c77edf · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18d541b-1b04-44b3-b11c-a8c401a50f3b · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling.arXiv, 2024
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86a15652-b7a6-47e4-abb7-6685669a3438 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b50544-9ab8-47d9-84b5-a70d76dc94b0 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark.NeurIPS, Datasets and Benchmarks, 2023
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 890b33c7-a2e6-4119-a72a-ecfe12e73217 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.TACL, 2014
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 361b64b5-0a25-4096-baf4-c940b3ce35c1 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Llama-adapter: Efficient fine-tuning of language models with zero-init attention.ICLR, 2024
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf8ea24a-ae78-4a0f-bbff-491f68b5e0ff · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gpt-4v (ision) as a generalist eval- uator for vision-language tasks.arXiv, 2023
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4fff451b-3ad9-4207-ad5a-611a9e740573 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Xing, Haotong Zhang, Joseph Gon- zalez, and Ion Stoica
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3119628-c4d5-446c-bd6b-21d8f795f5f7 · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Minigpt-4: Enhancing vision-language understanding with advanced large language models.ICLR,
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9d90768-44cc-4730-ac0b-4e6b611cee8a · outbound
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles aug- mented
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9873766-246d-4eda-90e3-1c18aa61fda0 · inbound
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.