Pith. sign in

Paper Citation Record · LEDGER

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2505.15240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15240 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:10.894434Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:53:02.187519Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T18:53:02.683433Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact4
  • verified fuzzy21
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eadde58f-0925-492b-a928-73b06b038893 · outbound

This paper cites GPT-4 Technical Report.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:07.772078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:07.772078Z digest=sha256:33370a264676d403cb2e6c2aee0873dff81822711e8b97eb0e4fffb963ecabca

Observation 467dbc23-a0d5-4c5f-9049-756a0e40b7c0 · outbound

This paper cites Categorical data analysis.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Categorical data analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.953900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:07.872856Z digest=sha256:4f0f474963fc2872edeadd1657a6b5a9603c87746ef68c8a3c1346d157754ec1

Observation db7db87a-aea4-42ba-bf0e-65311de07201 · outbound

This paper cites A computationally intensive ranking system for paired comparison data.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge A computationally intensive ranking system for paired comparison data

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.838672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:07.925438Z digest=sha256:146f0955d862d42751ce728ba4d48940a893e5d264512cf001b8ebb8ee679dc3

Observation 96127373-87f6-45f5-a8eb-cbeca7724a9f · outbound

This paper cites Rank analysis of incomplete block designs: I.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:07.965543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:07.965543Z digest=sha256:72251733afd56149ffdd1e54a52db8933ef8b0b4adb4cd6fba87fabb91501ee6

Observation 35d1a2ed-7104-4683-83e3-433242852d34 · outbound

This paper cites Language models are few-shot learners.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.029766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.029766Z digest=sha256:bd54903b2b3848c32622fad9c78b1adc68bb59ac246ee9bcbf8f4eb958a77b6e

Observation e5936fee-505a-43d0-9255-cb5dda048f1c · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.130142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.130142Z digest=sha256:a218e514794417a2f5d709d78db24c6808cb73749714f36c7573c6366c6d567f

Observation 5d8571d4-de8d-4bcf-a2d5-ef4679f4ee63 · outbound

This paper cites Learning to rank: from pairwise approach to listwise approach.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank: from pairwise approach to listwise approach

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.179818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.179818Z digest=sha256:fb120933ab257618b096a7d37cd020291c4e237dbcc347aada06d113424b99a1

Observation 356dde46-a9d5-46f6-888c-3f5bdc7e283d · outbound

This paper cites Efficient bayesian inference for generalized bradley--terry models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient bayesian inference for generalized bradley--terry models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.252459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.252459Z digest=sha256:a1811fcfbd909d5cd42e761ed218a350603eef188ccdd039a5ac66effccdd831

Observation 45c945fa-246a-4147-8f9f-64fecc9a50e3 · outbound

This paper cites Models for paired comparison data: A review with emphasis on dependent data.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Models for paired comparison data: A review with emphasis on dependent data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.688879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.346278Z digest=sha256:1cb74e63d8d6fe9f87cca46f039c88f0be08a120e102d02885af43fc67d49d27

Observation 1e7ab943-58d5-4e45-8989-d6ed586af534 · outbound

This paper cites Humans or llms as the judge? a study on judgement biases, 2024 a.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Humans or llms as the judge? a study on judgement biases, 2024 a

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.587348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.409355Z digest=sha256:a9c336919e85ffb5df0b9d94fb849424f30ac48ffa456d4d0bd4659f4fd24470

Observation 0d7ef445-02df-496d-bb81-bf337597a535 · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Premise Order Matters in Reasoning with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.450887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.450887Z digest=sha256:5cfb3860dda1eed38e0cf9a6c221fb6d10d00ebb26cf6600b8d7e9f24d70d078

Observation d0015be6-7a2c-4c43-a796-579b9d683950 · outbound

This paper cites Of human criteria and automatic metrics: A benchmark of the evaluation of story generation.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Of human criteria and automatic metrics: A benchmark of the evaluation of story generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.488265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.502773Z digest=sha256:266df2ecd0b6eb1374806bf1dd5e18229847d0646dea4b4bd1ed72660c96b575

Observation f9799cb4-970b-4e35-ba8e-c3918c478614 · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.594167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.594167Z digest=sha256:1e11a1da3965389b7ad9afb5d68508159ff6a78e8659f4bb6f26bb4c1788385e

Observation 2d6691a0-6d51-4edd-9f5e-8926df2563c6 · outbound

This paper cites Scaling instruction-finetuned language models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Scaling instruction-finetuned language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.397921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.640171Z digest=sha256:0e48309b3a732f65f605a13410736ec5f146cb1ccfd1359df106133145d39ce9

Observation 1a6aa0b7-0ce2-484c-9fec-a6d176015723 · outbound

This paper cites Ranking by pairwise comparisons for swiss-system tournaments.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Ranking by pairwise comparisons for swiss-system tournaments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.315039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.679816Z digest=sha256:15bde5b553903658f68b55951c32a6bb2254a9a1e16f1a25a0e43b50b152308b

Observation 6069fd9a-d81c-47ed-9d81-93c7aaf9f6cc · outbound

This paper cites The method of paired comparisons, volume 12.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The method of paired comparisons, volume 12

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.212792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.741755Z digest=sha256:022fd5ea14c29fdbfe1a04e615ef91a8181642f6b0e1d8cf7ba3dbbca6b4d7d5

Observation 446665cf-8dd5-4de3-8efc-e0ff5235425c · outbound

This paper cites The Llama 3 Herd of Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.826293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.826293Z digest=sha256:cbe525e44ce5b6486bcc0f47b0745889dc6e85b4e7509521d29c563984d14448

Observation f2db2e4d-2a6d-443b-90f4-870843a00bba · outbound

This paper cites Rank aggregation methods for the web.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank aggregation methods for the web

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.106964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.860405Z digest=sha256:4b433092a0dc118beefc58ff30d5134d43402665a90456a8d0285ba3abe26583

Observation da8ea7c2-855e-4f5a-bf8a-017ff2f09038 · outbound

This paper cites Summeval: Re-evaluating summarization evaluation.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Summeval: Re-evaluating summarization evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.961343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.964922Z digest=sha256:904c44c637f3593bb05d34df3f52b77e4c07946f18d85a2a8825b365aa0b1a38

Observation 8042578b-947a-4264-b9c7-a3660f7d2c4f · outbound

This paper cites GPTScore: Evaluate as You Desire.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPTScore: Evaluate as You Desire

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.027206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.027206Z digest=sha256:25090b8990576844ed1f4d4f1a634418cc7bcff602e6920b96602f22e91841bd

Observation 8b01b668-9df0-496a-961b-60ca7706f7f0 · outbound

This paper cites Trueskill™: a bayesian skill rating system.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill™: a bayesian skill rating system

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.835903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.064810Z digest=sha256:31bb3e1febeb3f5a93dbff7cc1e305e8f7100d778dead3c183044ee967687abe

Observation b56986d4-cd59-4ed1-8c13-8dfcba911fe0 · outbound

This paper cites an unresolved cited work.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:13.709267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.174549Z digest=sha256:6e0bc0e13e6f5faf25b47935ff7e75893692f25b5a7beebe9e9f4180250cac23

Observation e6ff5acb-ce24-4ec7-8b84-622957033dce · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.226377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.226377Z digest=sha256:2b0d137413a4980b3a88a5e118f9e2baf7f5eca42c3d243bd3dcd88f264793c9

Observation 339a8c9d-66c9-4d53-acbf-8cf41c0a0d9b · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.268446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.268446Z digest=sha256:b252de8902924c24d500c86e88d90d947ddc27fecd5014bb2a976a52d79b0558

Observation 73b61eed-2a0f-4020-a8a6-cbf2e2bca907 · outbound

This paper cites Learning to rank for information retrieval.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank for information retrieval

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.310157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.310157Z digest=sha256:8243fe3f90bfd7ac9a281c6311a855ac8b1b8e144f5c660370884d1eaf9dfb29

Observation fffee44e-23f9-4fd5-a59a-e7f9814163e9 · outbound

This paper cites G -eval: NLG evaluation using gpt-4 with better human alignment.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge G -eval: NLG evaluation using gpt-4 with better human alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.389780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.389780Z digest=sha256:d988f1fee3dafef6669092e4a7f0405d15f910b901484ecd6056198c0fb9e71b

Observation 6340000d-5367-4002-8518-90d0d56a1d97 · outbound

This paper cites Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.546621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.452271Z digest=sha256:72ee9e573bb9cc374d6b2544077c924758040904471d70fcae23dd97bcd4ce2d

Observation 15275cc5-d69f-428f-9750-5ce9be79c1f0 · outbound

This paper cites Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.788655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.490840Z digest=sha256:fecc25f02beb03ebc5063ad81568fa7f5417bfac31fd51c08427ca8e309d6091

Observation 4dc69aba-b45f-4250-bafe-7e67660805b5 · outbound

This paper cites LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.241221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.555067Z digest=sha256:a6a97bb8b887b1f59ef71bbb702a1fc677ded813c1ce0f157e5d80911419a32a

Observation 2da9566a-ddd2-4c3b-b617-7e6400494f5a · outbound

This paper cites Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.676956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.676956Z digest=sha256:391c59942dc403b0a1e9d1166c0c61ee1d7bdf879aaafe6510ce8b02425334e7

Observation 8d705c09-3f70-4972-8cf6-05be56ecf021 · outbound

This paper cites Stated choice methods: analysis and applications.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stated choice methods: analysis and applications

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.992603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.717192Z digest=sha256:645b6db0b064e3ac625149727974d70ab618c33ff6c8abbaee466ad2515d6d6e

Observation 604b0b71-aa6f-440e-9a13-713aa2d9803b · outbound

This paper cites The structure of random utility models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The structure of random utility models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.905404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.761167Z digest=sha256:1d81633da83e8d715fd10c37756f46bba8672f3ef7b6ee0de45d91f85e21d809

Observation e5bcca16-bf0f-4c06-8e03-87463ceed9af · outbound

This paper cites Trueskill 2: An improved bayesian skill rating system.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill 2: An improved bayesian skill rating system

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.795135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.855334Z digest=sha256:40338577a787335ec199221a54aa3308929df6208dc524a0b8cf222e5b3b521c

Observation 1ed99d0f-dab3-457f-8066-f050f4cb1d0b · outbound

This paper cites Efficient computation of rankings from pairwise comparisons.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient computation of rankings from pairwise comparisons

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.683651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.912909Z digest=sha256:8d409c4cf6f58f4e0261b4d39a2a79d8827718eb9aed1dc07a772b647039476a

Observation f3c346b9-2726-4ab0-a5ca-f87e319bb147 · outbound

This paper cites Training language models to follow instructions with human feedback.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Training language models to follow instructions with human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.957908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.957908Z digest=sha256:ad2b5634412e83b300a7a88443b47145e97a652ddad814bd51c9e0dc7820b7c8

Observation a898b7d7-b1e6-4c88-990c-ea67df9c3fb8 · outbound

This paper cites PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.613917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.001331Z digest=sha256:87476f6ee1cfb41862ccded92fd90733cdf89903627038512d6517b0de02e546

Observation e1e43ffc-2cc6-4775-9367-77bafad316bf · outbound

This paper cites Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.098915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.098915Z digest=sha256:4a1d3600b00d384be38912dbbf2a710615ddb366630b8a82e3ca16c37d28b1ea

Observation a45b21d5-04e9-4df4-9199-7dcac28bfe06 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Qwen2.5: A party of foundation models, September 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.160968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.160968Z digest=sha256:8398d0f71d3e8b2247e7eba9f6ec0b1caf43015b59241f5509b0553692aaaf05

Observation f1abfbcf-c988-4af7-b6c5-960383d42a60 · outbound

This paper cites Finetuning LLMs for Comparative Assessment Tasks.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Finetuning LLMs for Comparative Assessment Tasks

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.376856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.202374Z digest=sha256:dcd33af27dda33206fa2d1a61304c023dfdd10b9b94cee718e6373e4ee767eec

Observation 1b2d98d3-156a-4685-a01b-f136671e54a9 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stanford alpaca: An instruction-following llama model, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.260157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.260157Z digest=sha256:d5034d71b6ef9d3cedb6573a11e79a45936b9078c36c9b23248434d7f6afbf80

Observation 9ef6c73d-62fc-4021-b73b-af03aada0b01 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.352664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.352664Z digest=sha256:51b8145e11598c5c7945b58ee320ff1ea642a949e716dd3b2a8c051a0a2f3111

Observation e62c44d6-1bf2-4314-b9dd-834bdb2ca0b6 · outbound

This paper cites Large language models are not fair evaluators, 2023 b.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large language models are not fair evaluators, 2023 b

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.506759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.387688Z digest=sha256:526f57561bbb04b8e614f70b4f6bc05084d38a8ce67de7f0590305fa9e90676f

Observation d010cadd-bc90-4b67-98f0-71747f35ac9e · outbound

This paper cites Primacy effect of C hat GPT.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Primacy effect of C hat GPT

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.453435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.453435Z digest=sha256:98d5cf1a95517bde113079baac0bb07d71dd0e849098eea41e86e4ca2f176558

Observation ab765c02-662e-40be-9d11-00c87efb2024 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.548443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.548443Z digest=sha256:c2ac645bd4203e42f579a0f6bde3f7c27991e355522d27351b977c3ebf85a422

Observation e96b7956-0621-4486-b72c-160692212958 · outbound

This paper cites an unresolved cited work.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work

Reference 45

Resolution
verified exact
doi, observed 2026-08-07T15:28:11.117158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.610236Z digest=sha256:bb0eb1b1ce8f44e1ae8635e610aa8a33afcc99031b0ac056489d23a42c11bc64

Observation 240c239b-42aa-4f59-a97f-d2e88c222feb · outbound

This paper cites Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.399371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.723001Z digest=sha256:4cbbc7f18a4636593de60971974946ccbe0f79839092eb96aca4b6432925686d

Observation af9ba3e4-36ba-4be3-85fa-bb043de809d4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.779957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.779957Z digest=sha256:8090fe4787917d289938db16ee77bc568b447b43d07ad1eda6cb92a2587ef448

Observation 3954a661-deb0-4836-96b3-456b129590b5 · outbound

This paper cites Lima: Less is more for alignment.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Lima: Less is more for alignment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.275766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.828110Z digest=sha256:c21a34fb384654f123040969804f998e89d30b43d8502259c79710cde7887b1a

Observation 9b426243-0836-4127-a569-ebc5eb8d8e14 · outbound

This paper cites Judgelm: Fine-tuned large language models are scalable judges.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judgelm: Fine-tuned large language models are scalable judges

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.145963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.894434Z digest=sha256:61558af2b9676acbb55a5c1480c3fa1b77b5b0a3926a0042bbb37a4a636a036a

Pith citing papers

Observation 27443c4a-b1ef-430f-89dd-8f762f430d72 · inbound

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth cites this paper.

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:49:20.019956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:49:20.019956Z digest=sha256:7ec6f631c959658d9a5640bd184e823a8564471f77225b1a080f80864fb95272

Observation 8fec4e1d-53fa-47cd-94f0-2b042fbe1a96 · inbound

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation cites this paper.

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T18:53:02.689460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-08T18:53:02.187519Z digest=sha256:43d4e22da6f045bb30fc04eb6e91c07e8af7423f95d69b5495827949873e713b