Pith. sign in

Paper Citation Record · LEDGER

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2505.15240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15240 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:10.894434Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:49:20.019956Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact4
  • verified fuzzy21
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eadde58f-0925-492b-a928-73b06b038893 · outbound

This paper cites GPT-4 Technical Report.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:07.772078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:07.772078Z digest=sha256:29585054c9d83b01ed51366a807d9ec294aad95ff13bddf80b1a704a05a3a8b1

Observation 467dbc23-a0d5-4c5f-9049-756a0e40b7c0 · outbound

This paper cites Categorical data analysis.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Categorical data analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.953900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:07.872856Z digest=sha256:2be631a64d3b4c738fc68bdbf457edea7886a910adb2bae2fd02477335cfab14

Observation db7db87a-aea4-42ba-bf0e-65311de07201 · outbound

This paper cites A computationally intensive ranking system for paired comparison data.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge A computationally intensive ranking system for paired comparison data

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.838672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:07.925438Z digest=sha256:4a410147a5b0a7db9509e026c85ab6d58f63db0771dfae057a87d4ab9eb47626

Observation 96127373-87f6-45f5-a8eb-cbeca7724a9f · outbound

This paper cites Rank analysis of incomplete block designs: I.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:07.965543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:07.965543Z digest=sha256:4572c2b3a8d74edad97b7b5450df82d4e3aee1f99cbcbc9356521acdcb2cf1c4

Observation 35d1a2ed-7104-4683-83e3-433242852d34 · outbound

This paper cites Language models are few-shot learners.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.029766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.029766Z digest=sha256:fed69cb6545e98469b80e77aa603a136ec120011eac82899ad4e4b93486cb0fa

Observation e5936fee-505a-43d0-9255-cb5dda048f1c · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.130142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.130142Z digest=sha256:03eb14940cd7872ab62adaf220268748153e537d5d481761f0fc68e6ed5a48f5

Observation 5d8571d4-de8d-4bcf-a2d5-ef4679f4ee63 · outbound

This paper cites Learning to rank: from pairwise approach to listwise approach.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank: from pairwise approach to listwise approach

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.179818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.179818Z digest=sha256:4f8146e098e65b4175c0fb003a32bf150f2975a54f27bfe1973a8f656e8e7b8e

Observation 356dde46-a9d5-46f6-888c-3f5bdc7e283d · outbound

This paper cites Efficient bayesian inference for generalized bradley--terry models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient bayesian inference for generalized bradley--terry models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.252459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.252459Z digest=sha256:4c30567b164b0639941746fda05ea88fe01d190c058153f02a240ee6d922b584

Observation 45c945fa-246a-4147-8f9f-64fecc9a50e3 · outbound

This paper cites Models for paired comparison data: A review with emphasis on dependent data.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Models for paired comparison data: A review with emphasis on dependent data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.688879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.346278Z digest=sha256:5978779e4737f57f046511db6b37bd5c02e6c6b1c3831547a74794abbfedcae1

Observation 1e7ab943-58d5-4e45-8989-d6ed586af534 · outbound

This paper cites Humans or llms as the judge? a study on judgement biases, 2024 a.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Humans or llms as the judge? a study on judgement biases, 2024 a

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.587348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.409355Z digest=sha256:98defe9b986859002e9213f77d986c6eeebbd55da648c263cad53c3f87a49322

Observation 0d7ef445-02df-496d-bb81-bf337597a535 · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Premise Order Matters in Reasoning with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.450887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.450887Z digest=sha256:323c38e7d42e3e2dcbc5041985278ba495f1e8e0af3932121ec4c6b9b288c095

Observation d0015be6-7a2c-4c43-a796-579b9d683950 · outbound

This paper cites Of human criteria and automatic metrics: A benchmark of the evaluation of story generation.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Of human criteria and automatic metrics: A benchmark of the evaluation of story generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.488265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.502773Z digest=sha256:b47982a12e7f067c64852acb953decf301613aaa5e7c8bd435ff9b4386c54597

Observation f9799cb4-970b-4e35-ba8e-c3918c478614 · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.594167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.594167Z digest=sha256:2ba527727e2712a37c54970720b39549ff574cd5bbab17e60c3a9762cfff9665

Observation 2d6691a0-6d51-4edd-9f5e-8926df2563c6 · outbound

This paper cites Scaling instruction-finetuned language models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Scaling instruction-finetuned language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.397921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.640171Z digest=sha256:ebd4ec887f34e8bb5f4ca5eee91bf81a0e7442f78adbce335bd1295dc17bd3e8

Observation 1a6aa0b7-0ce2-484c-9fec-a6d176015723 · outbound

This paper cites Ranking by pairwise comparisons for swiss-system tournaments.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Ranking by pairwise comparisons for swiss-system tournaments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.315039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.679816Z digest=sha256:7531be1dc34199796318c1173bdeb9444612b0b6ef392b973d83fa243e71f7f4

Observation 6069fd9a-d81c-47ed-9d81-93c7aaf9f6cc · outbound

This paper cites The method of paired comparisons, volume 12.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The method of paired comparisons, volume 12

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.212792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.741755Z digest=sha256:b739dc77a21cee9a61f1d665356d505154002ccb38e5749efa78bb9f7840db7c

Observation 446665cf-8dd5-4de3-8efc-e0ff5235425c · outbound

This paper cites The Llama 3 Herd of Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.826293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.826293Z digest=sha256:94b06aa875112ccf6248bbd746bd3495e9655c0669ccdb69ceb0c016beafbbe7

Observation f2db2e4d-2a6d-443b-90f4-870843a00bba · outbound

This paper cites Rank aggregation methods for the web.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank aggregation methods for the web

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.106964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.860405Z digest=sha256:f91c667c82b44852d89d28152547545f0a79631e4d056b5d25ddf3e81d34cf4d

Observation da8ea7c2-855e-4f5a-bf8a-017ff2f09038 · outbound

This paper cites Summeval: Re-evaluating summarization evaluation.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Summeval: Re-evaluating summarization evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.961343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.964922Z digest=sha256:c199322e8872fe4e75674cfed4196f1a59bd1e2176985b2d69ac786470562e21

Observation 8042578b-947a-4264-b9c7-a3660f7d2c4f · outbound

This paper cites GPTScore: Evaluate as You Desire.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPTScore: Evaluate as You Desire

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.027206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.027206Z digest=sha256:6bde7cb54dd318e5d35313fd2a9600b4006ed4130a99518f00cde27f1d7599d6

Observation 8b01b668-9df0-496a-961b-60ca7706f7f0 · outbound

This paper cites Trueskill™: a bayesian skill rating system.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill™: a bayesian skill rating system

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.835903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.064810Z digest=sha256:a7baed52ee9d92a2ca2bce70baa0512edfb88f5de1d31bb80aa32e01a49ac4b1

Observation b56986d4-cd59-4ed1-8c13-8dfcba911fe0 · outbound

This paper cites an unresolved cited work.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:13.709267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.174549Z digest=sha256:d74a4559e0ee7680178ac8e8aa18f207ce22a192afcfb7c2274fc1a920f053e0

Observation e6ff5acb-ce24-4ec7-8b84-622957033dce · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.226377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.226377Z digest=sha256:560bdcf4f81673936e59d784f87fedd1b72e47edc53aa7dfa31eef5b35f99584

Observation 339a8c9d-66c9-4d53-acbf-8cf41c0a0d9b · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.268446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.268446Z digest=sha256:106275aaea5f0d56136f42fb7b2a386e68d7c8b8e2f96a65f35fb2f7ea0da42b

Observation 73b61eed-2a0f-4020-a8a6-cbf2e2bca907 · outbound

This paper cites Learning to rank for information retrieval.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank for information retrieval

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.310157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.310157Z digest=sha256:68ecca85d9fd29846ff880b01c598051e88029ba684b29ecccbc89134b0d456a

Observation fffee44e-23f9-4fd5-a59a-e7f9814163e9 · outbound

This paper cites G -eval: NLG evaluation using gpt-4 with better human alignment.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge G -eval: NLG evaluation using gpt-4 with better human alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.389780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.389780Z digest=sha256:50fc04f92ffb99e4973064ad37d70dcf937db00f385e04c5a6184ddeb01871a0

Observation 6340000d-5367-4002-8518-90d0d56a1d97 · outbound

This paper cites Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.546621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.452271Z digest=sha256:38b491c4ae735d535b81cd0f9238536e106e4438b6e068e26dfea7402aab8c8a

Observation 15275cc5-d69f-428f-9750-5ce9be79c1f0 · outbound

This paper cites Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.788655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.490840Z digest=sha256:d521667fd4ca989cfbd4e9dfa92ae09a48bd2a5d7e7baefeb8cc2a4975e76caa

Observation 4dc69aba-b45f-4250-bafe-7e67660805b5 · outbound

This paper cites LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.241221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.555067Z digest=sha256:2311b939d72e8ca06ba33f4975045c4981ce94bb70347cd3e737b98844018f7b

Observation 2da9566a-ddd2-4c3b-b617-7e6400494f5a · outbound

This paper cites Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.676956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.676956Z digest=sha256:8f017e9a8979e34aab629ea96a59357556807689ea613e2309fe645cb654dcd0

Observation 8d705c09-3f70-4972-8cf6-05be56ecf021 · outbound

This paper cites Stated choice methods: analysis and applications.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stated choice methods: analysis and applications

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.992603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.717192Z digest=sha256:42d4ba349b228974feb4b77ed430addbdf37dc4a08f50b216bc42b29bf0452a2

Observation 604b0b71-aa6f-440e-9a13-713aa2d9803b · outbound

This paper cites The structure of random utility models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The structure of random utility models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.905404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.761167Z digest=sha256:62a102ea52641e131909cdb5c076af4b3bf8101c052702181b85bc9d65a32fcb

Observation e5bcca16-bf0f-4c06-8e03-87463ceed9af · outbound

This paper cites Trueskill 2: An improved bayesian skill rating system.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill 2: An improved bayesian skill rating system

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.795135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.855334Z digest=sha256:b17222ce88a1ddad503b15e379ad0ddb96a9dceaa84022480ebbcde51f406817

Observation 1ed99d0f-dab3-457f-8066-f050f4cb1d0b · outbound

This paper cites Efficient computation of rankings from pairwise comparisons.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient computation of rankings from pairwise comparisons

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.683651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.912909Z digest=sha256:a1772932a0f89e3d7fe859e76a95145d6c31cef4678a9bb8019428c6263684ad

Observation f3c346b9-2726-4ab0-a5ca-f87e319bb147 · outbound

This paper cites Training language models to follow instructions with human feedback.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Training language models to follow instructions with human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.957908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.957908Z digest=sha256:eb4c6b355d76bd5392cc55145989282fb22b3dc1de299113c7a76d06ebaabf05

Observation a898b7d7-b1e6-4c88-990c-ea67df9c3fb8 · outbound

This paper cites PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.613917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.001331Z digest=sha256:a748adf552aa080f7625996ff0b3951a06ab7ab58ce71183c941458ddcbb1953

Observation e1e43ffc-2cc6-4775-9367-77bafad316bf · outbound

This paper cites Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.098915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.098915Z digest=sha256:3243d621d017a283cac4237c6c2dbc601cd3a3fec4a17bf16e73bac2df2a0af7

Observation a45b21d5-04e9-4df4-9199-7dcac28bfe06 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Qwen2.5: A party of foundation models, September 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.160968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.160968Z digest=sha256:671b16f5db44a4542afc9f6f4976fa06c9cc5d657d28e390aa4a3f65034cb8a3

Observation f1abfbcf-c988-4af7-b6c5-960383d42a60 · outbound

This paper cites Finetuning LLMs for Comparative Assessment Tasks.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Finetuning LLMs for Comparative Assessment Tasks

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.376856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.202374Z digest=sha256:2e38e95f1ec77645b24e87b1b27ef3b044997a9f76f3dc6bd7329581b8d8e10c

Observation 1b2d98d3-156a-4685-a01b-f136671e54a9 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stanford alpaca: An instruction-following llama model, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.260157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.260157Z digest=sha256:5615bbedd9bab6a5efd327d04563cc5b4c30fffd9214da598aca7b395a0740b6

Observation 9ef6c73d-62fc-4021-b73b-af03aada0b01 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.352664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.352664Z digest=sha256:1a1748874cdbb8109db601f1093f9a8215a0413ef330d86add7fb3092685b522

Observation e62c44d6-1bf2-4314-b9dd-834bdb2ca0b6 · outbound

This paper cites Large language models are not fair evaluators, 2023 b.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large language models are not fair evaluators, 2023 b

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.506759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.387688Z digest=sha256:7b973bb9eb979d2a68402739d0048b680aeb5441adaaff99e4ccdad53d9f7608

Observation d010cadd-bc90-4b67-98f0-71747f35ac9e · outbound

This paper cites Primacy effect of C hat GPT.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Primacy effect of C hat GPT

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.453435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.453435Z digest=sha256:dc44dcdb9e28a9d5498224e2a3e95fe33cbe66b0d70cfcf02c08e7a464656476

Observation ab765c02-662e-40be-9d11-00c87efb2024 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.548443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.548443Z digest=sha256:9ae826d37e1efa0db54beff70e3e54dbdfec996e6a6bbd6750bd8886167a0c99

Observation e96b7956-0621-4486-b72c-160692212958 · outbound

This paper cites an unresolved cited work.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work

Reference 45

Resolution
verified exact
doi, observed 2026-08-07T15:28:11.117158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.610236Z digest=sha256:c402e10743f742f3b9891f240869a89722d2620142b37c9b4def1a67cbbfe226

Observation 240c239b-42aa-4f59-a97f-d2e88c222feb · outbound

This paper cites Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.399371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.723001Z digest=sha256:693af3967546f94fa68bb4428d87ccfef695079f6ac51812c18e5b6e440c312b

Observation af9ba3e4-36ba-4be3-85fa-bb043de809d4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.779957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.779957Z digest=sha256:fa2a9c0aff071f8aac76a9693ad2b0d9d2c0a2182c90f09fc24b6565882c0db1

Observation 3954a661-deb0-4836-96b3-456b129590b5 · outbound

This paper cites Lima: Less is more for alignment.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Lima: Less is more for alignment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.275766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.828110Z digest=sha256:6016fc89886761d8141a8f1bfe8488cbe16913d68998f841861897bed0936073

Observation 9b426243-0836-4127-a569-ebc5eb8d8e14 · outbound

This paper cites Judgelm: Fine-tuned large language models are scalable judges.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judgelm: Fine-tuned large language models are scalable judges

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.145963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.894434Z digest=sha256:e2b1c88f2adfe437ab5a97cc6f6691ceec44d0048353cd117c5f3df99b3ddc28

Pith citing papers

Observation 27443c4a-b1ef-430f-89dd-8f762f430d72 · inbound

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth cites this paper.

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:49:20.019956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:49:20.019956Z digest=sha256:6490f8b8299685d12b095721a39116ffb4724c41a44aca522d38024e24b60ebb