Pith. sign in

Paper Citation Record · LEDGER

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

As of 18 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 1 inbound Pith citation observation for arXiv:2509.08777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08777 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:10:00.626842Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:52:45.867032Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3608cf9-115f-4bf5-b4e5-07ebb093b7b4 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Spice: Semantic propositional image cap- tion evaluation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.693800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.156285Z digest=sha256:32bf7d43ae475bff35231436fcb989f8c46c67d47807e516f5cf3a697a663333

Observation e6c8d3f1-6575-4cfb-8b47-53e70e1ee3fe · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.670583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.166414Z digest=sha256:3545c8f1ea8b32999a133ccba2fb8d62b78def0360bf0568d93f455551928958

Observation 6bd06c55-e959-4e66-a6be-e108851d6205 · outbound

This paper cites Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els.arXiv, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.650461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.173348Z digest=sha256:f9dd4c6387eec121267cf4057945f922cdb85f93c740b3de594b65df0f168529

Observation 0f656ce2-34cf-4ab0-88f8-96953b6a04cf · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv,.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.629248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.180401Z digest=sha256:c2c478b0d5c2be8b589a253de9f56120950e40cba477b40bcfeeca42081553dc

Observation 296632bc-bb85-455f-925a-bb7ee389111d · outbound

This paper cites The control of the false discovery rate in multiple testing under dependency.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles The control of the false discovery rate in multiple testing under dependency

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.607339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.187866Z digest=sha256:87764af3162de7671e0eabd150349323faaacf5eba21e2f468553d7a3c7c5a23

Observation c3889dd4-8ff4-4d03-ae37-b04576e9e4d5 · outbound

This paper cites Improving image genera- tion with better captions.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improving image genera- tion with better captions.arXiv, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.583340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.193866Z digest=sha256:e3e84a266317c1dee721f5048e4ac32eed801badef34d2ff177beb439ea79c68

Observation 72389f87-8d8c-4384-844c-f9a1c4df1373 · outbound

This paper cites Visit-bench: A benchmark for vision- language instruction following inspired by real-world use.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Visit-bench: A benchmark for vision- language instruction following inspired by real-world use

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.558709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.203766Z digest=sha256:78c346a581fb51ea45657d043a35b0157d9afb451cad32b609cab96d0bcb80c0

Observation 8f9e2392-17a6-4326-a0ae-d838dfc86f8d · outbound

This paper cites an unresolved cited work.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:10:02.538394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.214998Z digest=sha256:5b474fc6a4b0763e521c14a4dd8d1346a3fbfcdcf096eea7a88c157a2adf0112

Observation 8f40a452-7d07-4743-8df3-d967181035f0 · outbound

This paper cites an unresolved cited work.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:10:02.517223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.223854Z digest=sha256:e2b9f1481d631eac23c7fc52214e7e4db5f90bb090c15780f0924ac4d87be3d0

Observation 6bce0c68-48d2-4d24-9aa6-6c113d35c710 · outbound

This paper cites X-iqe: explainable image quality evaluation for text-to-image generation with visual large language models.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles X-iqe: explainable image quality evaluation for text-to-image generation with visual large language models.arXiv, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.499482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.236650Z digest=sha256:ebefd908ae4ffb7179d1a6a70835e457592bf06f0271c70b8cfcfd83cf0612e8

Observation cd27c001-1f68-470a-a1f1-bbc3fbfcdd66 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.244946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.244946Z digest=sha256:6cf78483672dc3b4680ca5f53d2eae6e4ca5cb2a93d890b255f5e41c9c377b06

Observation 154966d0-46b3-4f1c-b0c8-985b62991e5a · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.ICCV, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.ICCV, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.481840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.250740Z digest=sha256:8effaeb48776395fce96c52c4d961cff2b86150299513d3e45799bdd0e63c4ee

Observation c98be1fb-0aaa-4adc-b28f-389233eeb660 · outbound

This paper cites A coefficient of agreement for nominal scales.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles A coefficient of agreement for nominal scales

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.257632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.257632Z digest=sha256:b5ba6e05155aa0ca70e34de39b0e897bdcd3181c88ab657967c0bfa43d9b68b9

Observation db31190b-3c9a-4b8d-83aa-5880698cb02a · outbound

This paper cites Ex- ploring GPT-4 vision for text-to-image synthesis evaluation.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Ex- ploring GPT-4 vision for text-to-image synthesis evaluation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.452085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.264208Z digest=sha256:c1a6b3e4a0b9ffe080cd9ed10431524fa61e2d2c5794cb43c2e4d61474728b17

Observation 4713f7d9-ba43-42c1-9b38-6cb3fab6fe3f · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.432589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.269383Z digest=sha256:4385389c8aa3f3f9073dda37b98c2edad72c9c905f9c5717b2967ec4d2248d43

Observation e9f98e67-98d7-424b-9c43-f51cbbc06d97 · outbound

This paper cites Improving selective visual question answering by learning from your peers.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improving selective visual question answering by learning from your peers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.395471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.275429Z digest=sha256:c5a887116298fde0b1d68931c48ab1e2ac83a203fea5046e72e4e0ffebf380a0

Observation 2d64d8c7-47ba-4d50-916b-1859b72ac210 · outbound

This paper cites Moura, Devi Parikh, and Dhruv Batra.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Moura, Devi Parikh, and Dhruv Batra

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.366788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.280966Z digest=sha256:cece854e49a4144175962baa502711f6cca820e52a9de56d2181c157501ca871

Observation 2fc28f98-7f8e-4b43-be9b-78b3ffc050c4 · outbound

This paper cites The relationship be- tween precision-recall and ROC curves.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles The relationship be- tween precision-recall and ROC curves

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.339782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.286873Z digest=sha256:5fe6e82e1f811a728f55b51cd612e45c7937ee0ce88ba6cd4491a4b0b1757c62

Observation c1a58c9a-e552-4a14-8230-cdfd66fb20ca · outbound

This paper cites Mllm-bench, evaluating multi- modal llms using gpt-4v.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Mllm-bench, evaluating multi- modal llms using gpt-4v.arXiv, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.308668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.292656Z digest=sha256:35f0d8463515cd731673520830ba6c69307685dc6a2e11235750448b0ee2430a

Observation cf889458-d5e7-4a37-812d-ada0289494a3 · outbound

This paper cites Selective classification for deep neural networks.NeurIPS, 30, 2017.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Selective classification for deep neural networks.NeurIPS, 30, 2017

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.279222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.297373Z digest=sha256:0e004bd2be9000ddcc2e0ea280f439fee3039da16bd018d48ce8ed6b2f0bd4f3

Observation 3ac3fd95-a6bd-4817-a1f5-a02601c56ebf · outbound

This paper cites Practical variational inference for neural net- works.Advances in neural information processing systems, 24, 2011.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Practical variational inference for neural net- works.Advances in neural information processing systems, 24, 2011

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.243263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.303656Z digest=sha256:8c55db171b05adfb6c404f509eddd27d151ff210212d331477b64f1ad8487ca0

Observation 50ed7c0d-0e39-4e1c-89b7-341d31b53deb · outbound

This paper cites A Survey on LLM-as-a-Judge.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles A Survey on LLM-as-a-Judge

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.308515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.308515Z digest=sha256:6fbcb745c690b2e22c5883cabc0696034265ffab946aaf88863ec9ded61de2cb

Observation 0c043b7d-f93c-4c06-86aa-a88c76a3c2f3 · outbound

This paper cites Weinberger.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Weinberger

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.207848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.315283Z digest=sha256:e8bec5cddd6d1c3a69b9edf4d353c507092b0f9bc6400e2d53681b5418dfc02e

Observation d1968aae-65ba-443b-abce-e9e7212bf6af · outbound

This paper cites Springer, New York, NY , USA, 2001.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Springer, New York, NY , USA, 2001

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.185588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.322231Z digest=sha256:339d203e3b7def1fdad12b313e6ba948dbcb00ba444efe696e7b5b51be37644f

Observation 961b2e05-ad1e-4d19-babe-1ef2bd9d6add · outbound

This paper cites Glass, and Yulia Tsvetkov.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Glass, and Yulia Tsvetkov

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.153021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.327430Z digest=sha256:a22839781285cee5950481e0757ebb4789c45b20d490e6f30becc72b6aa60797

Observation 50c20587-0fed-4899-9c3b-829f1045575b · outbound

This paper cites Machine learning with a re- ject option: A survey.Machine Learning, pages 1–38, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Machine learning with a re- ject option: A survey.Machine Learning, pages 1–38, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.122870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.333238Z digest=sha256:fa4df745618abdb764f52460c11d0712db29dd3203f65545581f2bf7a49e3bf1

Observation 1535f5b7-42ce-4705-b41b-377cce914983 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:02.072632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.338036Z digest=sha256:004e0532abba9385b30ac360ac7c2060086c6f66baa4d4f82b27cc337f636b31

Observation 4d1d5d0c-7ed4-4c99-a083-b21f4a11deb8 · outbound

This paper cites Promptboosting: Black-box text classification with ten forward passes.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Promptboosting: Black-box text classification with ten forward passes

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.822777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.342781Z digest=sha256:ce5b471402993bebc61b44e125f46129563ae74e672eb84d6e65b585dce07304

Observation ee14fb5c-be1b-4b8e-969a-4c5db627c4b7 · outbound

This paper cites Are llm-based evaluators con- fusing nlg quality criteria?arXiv, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Are llm-based evaluators con- fusing nlg quality criteria?arXiv, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.791347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.347986Z digest=sha256:f58164e8c5d155c23247bbcb971c2ad92e400d314804003d1110e902ea011425

Observation 1d2fd8af-c33a-46a9-ba26-84f3af4004ca · outbound

This paper cites GPT-4o System Card.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.353769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.353769Z digest=sha256:3ac6482cc8ced8d2e6e476960deb081e22720b8a088229356a1228d4fb27c0d2

Observation 630922ed-d7f5-43b2-a4ff-fc218f9e3f37 · outbound

This paper cites Calibrating language models via augmented prompt ensembles.ICML Workshop on Deployable Generative AI, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Calibrating language models via augmented prompt ensembles.ICML Workshop on Deployable Generative AI, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.772494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.358862Z digest=sha256:c98ffbde3e329550b96b6e2d80a60697c324c989d581c142809f48c50475ae60

Observation 63b39703-7552-43eb-a441-c6312f9fc3f2 · outbound

This paper cites How can we know what language models know?Trans- actions of the Association for Computational Linguistics, 8: 423–438, 2020.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles How can we know what language models know?Trans- actions of the Association for Computational Linguistics, 8: 423–438, 2020

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.753863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.364453Z digest=sha256:4b3cda36d5c4ea5870060908d0d61ec6096f1235ba16556cf3683a73459f80ea

Observation 4fcf8bfc-96e5-40bc-972c-0348b95f539c · outbound

This paper cites Billion- scale similarity search with GPUs.IEEE Transactions on Big Data, 2019.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Billion- scale similarity search with GPUs.IEEE Transactions on Big Data, 2019

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.369750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.369750Z digest=sha256:08d43d775fffc163f4d228e4d392c2e1f3ae17786a6911c9c235ba1e0113a274

Observation b5b565dc-d1de-4c51-9258-28f5acd50047 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.375084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.375084Z digest=sha256:ff694698f530cd6daf92eb805919a0b52b19d7c8d46e2411bad3d5caa1a68141

Observation dc03e82d-2585-4d28-8ac6-dfc2660a5cd1 · outbound

This paper cites Benchmarking cog- nitive biases in large language models as evaluators.arXiv,.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Benchmarking cog- nitive biases in large language models as evaluators.arXiv,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.722060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.381149Z digest=sha256:7dbfda2324c6a237c6f1a6b50585a323ad524336e82b261870d9e08ee7f13eb7

Observation 6457c77b-21a7-4ee5-9e07-6bd84e102aa9 · outbound

This paper cites Improved precision and recall met- ric for assessing generative models.Advances in neural in- formation processing systems, 32, 2019.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improved precision and recall met- ric for assessing generative models.Advances in neural in- formation processing systems, 32, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.698675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.386597Z digest=sha256:e991be57730c030c7e1a69d9c6c77c4a912a3fb25b1b1162aa55cdb5b3b3d687

Observation 6730b91e-1eb6-47f2-9a98-f5c6c36a263e · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.CVPR, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Seed-bench: Bench- marking multimodal large language models.CVPR, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.680752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.392057Z digest=sha256:e32c6c365d6c0a63ef090ab3acd493137dafa415b9386090ad4da6e5ae8b8d68

Observation aacab920-6bf6-4e5b-856f-6e58fd94d9bb · outbound

This paper cites From generation to judgment: Op- portunities and challenges of llm-as-a-judge.arXiv preprint arXiv: 2411.16594, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles From generation to judgment: Op- portunities and challenges of llm-as-a-judge.arXiv preprint arXiv: 2411.16594, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.398173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.398173Z digest=sha256:25407bdda0c96595ee610463b1ef596914b4b4acfbfd55146ee86203ef812a72

Observation ed1813db-8738-4375-b8de-e443211a01e9 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.403033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.403033Z digest=sha256:5bfa6adb7e86f6d949a1f7004701bc05c8bb6fdbe97ce5eb02961213c5f51c1f

Observation e6dc6f29-878c-4124-9224-42e5c027a030 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.Text Summarization Branches Out, 2004.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles ROUGE: A package for automatic evaluation of summaries.Text Summarization Branches Out, 2004

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.663637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.408036Z digest=sha256:15cee9178474141b02f3c49cc7faaab3cb00b81dc9a7354ade2befe45e88813f

Observation eedef4f5-10bf-4e1f-8883-18834e9609e5 · outbound

This paper cites Microsoft coco: Common objects in context.ECCV,.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Microsoft coco: Common objects in context.ECCV,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.649144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.412807Z digest=sha256:8ef01915b4a354446cbfb96ba4592119a7d890bdce62c69c838f60103aca1a2a

Observation 34c4dde9-8ae8-4694-b073-83a8c2f1213a · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.ICLR,.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Mitigating hallucination in large multi-modal models via robust instruction tuning.ICLR,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.634648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.418416Z digest=sha256:3c0aae3afb627436dde2277cb0e5a7ad7efbdd259a19c957fb0ff858f3ce758a

Observation 1f9f3cc8-6c8c-4959-9956-418a5551f759 · outbound

This paper cites Visual instruction tuning.NeurIPS, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Visual instruction tuning.NeurIPS, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.616525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.424190Z digest=sha256:414445411e165ce73762ba965c29478b8c5c3642e171687cf2ebcd2774edff33

Observation 08920eaf-a063-4fbd-b6dd-bdec8601d243 · outbound

This paper cites Llms as narcissistic evaluators: When ego inflates evaluation scores.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Llms as narcissistic evaluators: When ego inflates evaluation scores

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.599673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.429547Z digest=sha256:684c9c1efcb0f294a2bf2ce5c8e68a8c8749e52d5c60217c0ca79b7af6fe6995

Observation 79738d18-7d1f-4dda-b52b-0a6dc496f15a · outbound

This paper cites Predict re- sponsibly: improving fairness and accuracy by learning to defer.NeurIPS, 31, 2018.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Predict re- sponsibly: improving fairness and accuracy by learning to defer.NeurIPS, 31, 2018

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.583978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.437345Z digest=sha256:aa97183de1a8700a704f6f1d1358b55ebe78364eef8e7f60501f2e6c1a1fc948

Observation ff6be8c5-c650-4e3d-8f34-aa2d59ad8fbb · outbound

This paper cites Gpt-4v(ision) technical work and authors.https: //openai.com/contributions/gpt-4v/, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gpt-4v(ision) technical work and authors.https: //openai.com/contributions/gpt-4v/, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.567076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.447853Z digest=sha256:8d436ce9e58935328d6b68c547b755982e76debeb568d96af1acaa67974f2eaf

Observation 18587033-0c7e-4bae-b88d-3b52140a17ad · outbound

This paper cites Openai o1 system card.https://openai.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Openai o1 system card.https://openai

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.548783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.454060Z digest=sha256:3fcde223254ef89d033272973aa67c5bf91ea8b2f5689601b40aa25dae5241b0

Observation 81130e22-1f21-4b9b-9d3b-0b945da0266e · outbound

This paper cites Bowman, and Shi Feng.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Bowman, and Shi Feng

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.529864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.459708Z digest=sha256:66c6b8950124cb341b600ae401a20291b2f1714919fcd4c85cadc95d78273e95

Observation 29d9ffa6-db7a-46ab-a660-9fa37b651757 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.ACL, 2002.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Bleu: a method for automatic evaluation of machine translation.ACL, 2002

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.513432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.465601Z digest=sha256:0060f6acb8581cf9960165807b573a3fe7bfbb6abc13e1e40b8fd9853f1ca414

Observation a6302416-04e4-41e4-94bd-df1e4bafa69a · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Kosmos-2: Grounding multimodal large language models to the world.arXiv, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.496197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.470734Z digest=sha256:c8bb2612a4a1d95a66480f014dd174a00230dcf3dda9247ffde78eeec8ae1e51

Observation 9fe06e04-fb45-4444-b18f-cc46b54d7cbb · outbound

This paper cites Sdxl: improving latent diffusion models for high-resolution image synthesis.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Sdxl: improving latent diffusion models for high-resolution image synthesis

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.479329Z digest=sha256:c3930ec5136fd9f09615a2c230f0ac365b6893a1845079e6803d8f769c9ffa59

Observation a6642fe0-938a-4e65-b5af-5f2aa834ca84 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.465835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.485703Z digest=sha256:76822fa01de548e99642378e3a6faacac2542477461851adc7f8257d280f798c

Observation dd110cf1-eb76-4eb6-ac18-88681c72771d · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.491246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.491246Z digest=sha256:791c84c15bb074d32d11504be49816070562d1ceca746fe48c3e6bf1ac94b17f

Observation acda0efd-f74e-452c-b6a4-2b9bf2eef5b0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles High-resolution image synthesis with latent diffusion models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.497977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.497977Z digest=sha256:5aba2dec1867b595a47d3264866af5ecd7b0eda5b1d5cd9a6ca07bd92192c030

Observation 8c673ea0-015c-43c3-9ff1-55243a793312 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.504942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.504942Z digest=sha256:55316fc015c7606fef2aeeab1cc86aee09aa5d0b72e7cb9d0a5901dc58ef0834

Observation bd90b93e-f9d4-49d2-9b90-d0c73c5af6e5 · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Verbosity Bias in Preference Labeling by Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.510105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.510105Z digest=sha256:5a437e353c03eabed41f062a6197e6a9e13ef2282cf413d6d2d331140ce178ba

Observation 6f718aff-0726-4723-83f1-39f0762985ac · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.516193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.516193Z digest=sha256:94aad88a93217c57383195ed6d03b721fdfe9964e039dccafb70a6110b09ec55

Observation 74a7356c-e99d-486f-9520-91e1f63a4e2f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.522704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.522704Z digest=sha256:8499bb401912b1b0db1ae35f3f21187848dacac10afa01b5918671e9e8e761eb

Observation 2dac2553-0751-4480-9372-0f948155de8d · outbound

This paper cites Bayesian prompt ensembles: Model uncer- tainty estimation for black-box large language models.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Bayesian prompt ensembles: Model uncer- tainty estimation for black-box large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.411155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.528704Z digest=sha256:0084c6f485de166c17cb4b1bfeae3de5f2b7f87361884c8da65bd2e6b7bbb7d2

Observation 05aab997-1bb7-4874-83e8-5beefa44aae4 · outbound

This paper cites Large language models are not fair evaluators.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Large language models are not fair evaluators.arXiv, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.364011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.535789Z digest=sha256:482f076057158b1dc715705ec7b9bf2da5b07c5f94af1f5cb53fd88887411afe

Observation 2312aa86-2571-4756-85a9-3ac0906d2d99 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.541896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.541896Z digest=sha256:df1c13e028f0964e67f3dbaf5fc4c09238f389ac21d0cf8b29e292a34946accf

Observation 46695519-4fe1-4fdf-bb09-bbff6705d6a6 · outbound

This paper cites Reliable visual question answering: Abstain rather than answer incorrectly.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Reliable visual question answering: Abstain rather than answer incorrectly

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.347422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.548729Z digest=sha256:35c6306b04a34363214b17380567f2696b106f06fb54ef7dddb8a49f8df5f012

Observation 2790cf08-e40c-488b-ad53-ccdbcdb1149c · outbound

This paper cites Strength in numbers: Estimating confidence of large language models by prompt agreement.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Strength in numbers: Estimating confidence of large language models by prompt agreement

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.329704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.553882Z digest=sha256:c747a31160802fbba3620606a54e9ae7597876a7d900b7ed9f1b7e44fb59d2a8

Observation bac3ba20-eb94-4ada-8c22-672fa9814ea6 · outbound

This paper cites Gpt- 4v(ision) is a human-aligned evaluator for text-to-3d genera- tion.CVPR, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gpt- 4v(ision) is a human-aligned evaluator for text-to-3d genera- tion.CVPR, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.310636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.560479Z digest=sha256:7df7c59cabc2a584f8b0c9f28d5a09f965714a85be666668069e8422aaca34bd

Observation 040f3b76-5409-42df-a13f-ef97dae5b022 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.568962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.568962Z digest=sha256:b231819bd1ebad80faec6e08ac9b1a31e9a654eb8e1e99b52919dfeeffc63259

Observation ed5d8eb4-7abe-4a9a-a27a-5c59e2c77edf · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.575267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.575267Z digest=sha256:66a23265016b03a8d6bcb4d9ab19df59803b12d992302013eac3112bc7dbc246

Observation d18d541b-1b04-44b3-b11c-a8c401a50f3b · outbound

This paper cites Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling.arXiv, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling.arXiv, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.293121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.580346Z digest=sha256:f1a27f27e1c82d3b6ec74f29384967d99dcf8cff958ad167ce79d4c3a9120f39

Observation 86a15652-b7a6-47e4-abb7-6685669a3438 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T16:10:00.585677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:10:00.585677Z digest=sha256:25d9240896fa2fb390505b60f029e57884fd72973fd770c170ddbb699a593fcf

Observation 69b50544-9ab8-47d9-84b5-a70d76dc94b0 · outbound

This paper cites Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark.NeurIPS, Datasets and Benchmarks, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Lamm: Language-assisted multi- modal instruction-tuning dataset, framework, and bench- mark.NeurIPS, Datasets and Benchmarks, 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.215242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.591692Z digest=sha256:3ec4660f98c60e112d70b9ceb5fd32dcd88478e7850ddfe02ce5cf113e7450be

Observation 890b33c7-a2e6-4119-a72a-ecfe12e73217 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.TACL, 2014.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.TACL, 2014

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.170686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.596836Z digest=sha256:b8461f80e5685edaaf94359a9b5ebbb7ecf6613dc26c5d5f7a8e532e1e2b6b3f

Observation 361b64b5-0a25-4096-baf4-c940b3ce35c1 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention.ICLR, 2024.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Llama-adapter: Efficient fine-tuning of language models with zero-init attention.ICLR, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.151904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.601866Z digest=sha256:cbfcfc947f0e4cec4fd62ac2c6deb799c567a9543a9e9f6831b10b77ca3a7fdd

Observation cf8ea24a-ae78-4a0f-bbff-491f68b5e0ff · outbound

This paper cites Gpt-4v (ision) as a generalist eval- uator for vision-language tasks.arXiv, 2023.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Gpt-4v (ision) as a generalist eval- uator for vision-language tasks.arXiv, 2023

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.135983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.608245Z digest=sha256:9c97aea7abfdbcf7ca0b349ee430cd32c8bb7361e5e21fa3189659a6b5186e01

Observation 4fff451b-3ad9-4207-ad5a-611a9e740573 · outbound

This paper cites Xing, Haotong Zhang, Joseph Gon- zalez, and Ion Stoica.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Xing, Haotong Zhang, Joseph Gon- zalez, and Ion Stoica

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.114231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.614297Z digest=sha256:227dd062ed096a01f6cc27fd9658b1f16925d92d3f2df422a35ee4f4db9dcb35

Observation a3119628-c4d5-446c-bd6b-21d8f795f5f7 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.ICLR,.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles Minigpt-4: Enhancing vision-language understanding with advanced large language models.ICLR,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:10:01.094516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.621266Z digest=sha256:71f195255959f383d63b8eaf513534dce3342e1af774b64ae421a539b6f1ae9b

Observation a9d90768-44cc-4730-ac0b-4e6b611cee8a · outbound

This paper cites aug- mented.

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles aug- mented

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T16:10:01.077119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:10:00.626842Z digest=sha256:bfce735de2d023c11fc70ef5caf6c7a2c6e3a9abd81f95b503145d5674ada054

Pith citing papers

Observation a9873766-246d-4eda-90e3-1c18aa61fda0 · inbound

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics cites this paper.

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T11:52:45.867032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:52:45.867032Z digest=sha256:fd792c02dd6a21971742db88d87dc1b97a3a81921fa662c4e5bf8f3e629e3ec1