Pith. sign in

Paper Citation Record · LEDGER

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process

As of 6 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 2 inbound Pith citation observations for arXiv:2512.23213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.23213 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T19:57:03.999154Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T00:05:31.780655Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T00:12:50.170528Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact31
  • verified fuzzy21
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 55f87f30-114e-44fc-9a70-f0693f5dde68 · outbound

This paper cites write newline.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process write newline

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.542219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:fc3e453bcc21322adc9bfd761abd6abf591ee6340299b12576efb08502a014c0

Observation 1ba0712f-48ed-4533-b3de-3890f8afb36d · outbound

This paper cites GPT-4 Technical Report.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process GPT-4 Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:58:22.793264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:71035ed596f03b215b9a29abc90b90043289e929c6a4aa76010a6a5ece11464a

Observation e679e594-83fb-45be-bcd5-3f36a2caafd3 · outbound

This paper cites an unresolved cited work.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:01:14.539114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:5632e4ed9d8ecafe732099b9830a4184abe4ef7aa17e009b97cb0240736b60d8

Observation f312f3fe-0aca-4a9a-86ab-9b0882f57ccf · outbound

This paper cites Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.790072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:bd52d0a4e5660cda3e9fbf9cd6da1036e2a10d5b6f2773393cf673806961c5b1

Observation 4e51aa58-0d1f-4845-a47a-8b37440cca4c · outbound

This paper cites An automatic and cost-efficient peer-review framework for language generation evaluation.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process An automatic and cost-efficient peer-review framework for language generation evaluation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.783975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4152dce91e268ac55bfa4abffb3679645bdf6965dd63f262c379ac2c88189014

Observation 2e17bd6c-4686-4883-acb4-4a0cc4e71911 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.670667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:69d89d79b209460a08fc56142eb72af8a9244b0a19e4ab71b3da95268cf248e3

Observation cf20bed3-26db-467b-bd96-f3c06a0a6de4 · outbound

This paper cites Adversarial learning from crowds.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Adversarial learning from crowds

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.503775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:0aa3760b29ec25d954bea59fea71590bffbf992cc714f1ec42f6675d6568fa4d

Observation 7be0574a-f48d-448b-befd-80dfcf698294 · outbound

This paper cites Structured probabilistic end-to-end learning from crowds.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Structured probabilistic end-to-end learning from crowds

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:03:23.623798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:f2077b68967034839af9fd8dcc610d0143d93f12df038d68e60246c0c6be3395

Observation fe00d935-a89a-4111-8545-e782cfdeecf4 · outbound

This paper cites Neural-hidden-crf: A robust weakly-supervised sequence labeler.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Neural-hidden-crf: A robust weakly-supervised sequence labeler

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.511592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:df2cbb6aa8eb084761d4af200509894d84363cfe59038048389bb24a035d1ffc

Observation 6026d2f4-0d2b-44b7-877c-2be0cc6bab06 · outbound

This paper cites Harnessing Multiple Large Language Models: A Survey on LLM Ensemble.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.764827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:1f3aff1a3b4eadebfad34a56e40d0afd99061e21449f14a349f11beb3cd8af7f

Observation 7285815a-7cfc-49d1-8dba-63238100db8b · outbound

This paper cites E., et al.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process E., et al

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.506272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4f95efd05f7a395c0b5525f94ae74983a034c5073ad3805ce942770961a8a47b

Observation fcaf17fc-5d76-4cff-9683-91ae2556ed16 · outbound

This paper cites PRE: A Peer Review Based Large Language Model Evaluator.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process PRE: A Peer Review Based Large Language Model Evaluator

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.774949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:fb47dc84fa69481b984f3e6a0ac8f9e7e4b73d7456c41e6afd8f7c7aff05a7d6

Observation 83c10816-cc1f-40f9-b4ef-337d5c9f9655 · outbound

This paper cites W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:41:26.155989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:b8e788daedc00486b98a14724b84ab7d71b58dcb79b8b5cd178f46381b6549cf

Observation 548d295f-3c3d-4a82-8639-c6da529a96d1 · outbound

This paper cites an unresolved cited work.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:01:14.508791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:76cf7d3317f93a0708233faef170ae831d1a4db9d26ba46a05c8633e94fe1877

Observation 287c74e5-7450-4305-bced-e362b1dcc415 · outbound

This paper cites Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.778007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:88829df03df20e34e4b820f5b6a3ae70c77e76b775a05451a90b42e1be172914

Observation 4c3853e6-de2c-4daa-9e8f-bef746d4fe2c · outbound

This paper cites Maximum likelihood from incomplete data via the em algorithm.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Maximum likelihood from incomplete data via the em algorithm

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:03:23.621156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:e64e31617901c23e5e45c5d67580172b75d9dcb2b8a8635de4c2ee16883ca9e4

Observation 67c704b9-ccaf-4ab2-9730-69ec09908fc0 · outbound

This paper cites A survey on ensemble learning.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process A survey on ensemble learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.518249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:a20ee54d98cdfb666652958d59155cf2a98b4d706665a921029bcf85d0606865

Observation c1ef1cf4-96be-47de-ad2a-bba50997bb9a · outbound

This paper cites X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.515068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:6b5bdd9db84542ef3bd234c69a3567391fe83dccb0e9c526e2b98e11f65e2a5f

Observation 606e6660-3212-4f75-bc41-99a7f4d12071 · outbound

This paper cites An introduction to latent variable models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process An introduction to latent variable models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:41:26.153022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4a7a8135bf3f0fc73dbe4c0a29d506fbe72dbbe38e5e1bc5aae8b7f581cab542

Observation 9d2598b1-2d0c-4087-80b3-eb351084f506 · outbound

This paper cites GPTScore: Evaluate as You Desire.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process GPTScore: Evaluate as You Desire

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T17:11:54.379140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:56fd0842410b9d05740e472bd31141d9e6fbb8ff9b57f1a1bfa81c1b6a428213

Observation 3a89d433-5782-4ac5-acb7-3618d5d720bb · outbound

This paper cites Deep learning.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Deep learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:03:23.619107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:8a467bbb2111acc47ae32731f918ad91e89c7d6b5b1545eb5fa7db64548073bb

Observation b5ea98ae-841e-408c-9fc0-da4585180075 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process A Survey on LLM-as-a-Judge

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.771860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:ef0852dd79e3ba59dd31f49d05352ded3d35b491fb8e1871690550178e5deda5

Observation 59898454-34ea-4df1-ac40-9183f50218a3 · outbound

This paper cites Smoothie: Label free language model routing.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Smoothie: Label free language model routing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.561310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:46f06ca82a6c24d385dc0acf3387b7f587644cdef8b0694423bdb51dc7df624e

Observation 459e3f86-be96-4fdf-a295-e402dad61d8c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Measuring Mathematical Problem Solving With the MATH Dataset

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.691508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:9f3ac8a28a119a34a9444018aa77ccc6e1ef592314c33a24467f0d7e5f80fe90

Observation 54063440-a498-4b80-a97a-4c1d3acd80ca · outbound

This paper cites Language model preference evaluation with multiple weak evaluators.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Language model preference evaluation with multiple weak evaluators

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.780983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:8f256c633f708d09fd5afa3bdfc7edde21154c2a81028788d61e8def72f0f631

Observation 7e0b36de-b544-4ed8-9f4d-a51681672a1b · outbound

This paper cites Ensemble learning for heterogeneous large language models with deep parallel collaboration.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Ensemble learning for heterogeneous large language models with deep parallel collaboration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.558397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4144d6fe1d183cb0e2d949016f3919cddafd8e4b281201faf7648f7e275e907a

Observation 5529df80-0829-4460-bc9d-24cbda953884 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.743646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:286e6f35ed41587f038259b41421cc41f2b6402f08c6d2bb6ee01774384bc901

Observation 7b22e705-3058-4af8-9e7b-38f53e63a992 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:58:22.751121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:5002c91b26954abd9687e52008fdce576468bbdbf83c0cb55fc2e9dbcda19d16

Observation d43b8ff0-58d4-4439-988b-4f5774d5b637 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.721897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:2153ac156be504645d21ceecc8ba750bcd1f8bab0472f3145d6dcb58a679acc1

Observation 85b71a9c-0c22-455b-b72e-5f1f10478d90 · outbound

This paper cites Little Giants: Exploring the Potential of Small LLMs as Evaluation Metrics in Summarization in the Eval4NLP 2023 Shared Task.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Little Giants: Exploring the Potential of Small LLMs as Evaluation Metrics in Summarization in the Eval4NLP 2023 Shared Task

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.726394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:c5bb4ba1d087954b1583286484e8bd0c569fe0b5083a8f66f33c304c5344fff7

Observation 604caec7-42f9-47fa-9e96-8e1266405966 · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process From generation to judg- ment: Opportunities and challenges of llm-as-a-judge

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.667123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:5b5efc02843bb893c245e8dff02d6ca6a4827b38f95628aa2f3e780028337ef5

Observation d17ddcf3-9d6f-4481-b86a-9b1e06925d67 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.675035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:3a601fca6f17b5d7073b6da2f49c9d4a79a4d29f4e00385ca55d39344519a44d

Observation f6d20cde-7c1f-4e28-90a7-f03fc4a9e0e8 · outbound

This paper cites More Agents Is All You Need.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process More Agents Is All You Need

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.704184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:a3809977614b1c18756415fb64715a56e3aec997828e6fa81fc9d7c9454dbe01

Observation 168e4c3c-b897-400c-bf24-c3ca69b788c9 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.663206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:1af45206591d69cae31c8c022cea4ddf1f13cd011f4a69d21edb460a0f57e45a

Observation 0e89f12e-e1da-4938-995a-5c759c8e5f21 · outbound

This paper cites Cool-Fusion: Fuse Large Language Models without Training.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Cool-Fusion: Fuse Large Language Models without Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.678914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:6ec9e522caaedc4cbc2e554d0812b68d6721e8e7be24327a0431e825806cc370

Observation 45645ed7-e53e-4f05-a461-cdc540d49cdc · outbound

This paper cites Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.699710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:eeed2bcaab47cb6d3c70123a6d185d7fe1319595662124014a5ac0d79de12515

Observation 09e1e27d-d6b9-46a9-94e4-d2d1b5705446 · outbound

This paper cites Urg: A unified ranking and generation method for ensembling language models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Urg: A unified ranking and generation method for ensembling language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.554892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:5b1645776b156ab35e06546dc314604ec9260c48eaf6aa14fca66f2a38ee6f45

Observation 2772a185-390b-4434-ba2c-9db63c96ba37 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process RouteLLM: Learning to Route LLMs with Preference Data

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.682905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:c50b57c9a7ae44bbc4029168d94d639908a1ee8c8a3b65ff7a9ac58cfdd73ddc

Observation c0b38780-71d8-4aab-958d-29599baf0023 · outbound

This paper cites A Multi-LLM Debiasing Framework.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process A Multi-LLM Debiasing Framework

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.708723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:d0ff008c8b81d9645787f728e630732ec4c671d69c5b0abea137b41472fdc43a

Observation a30c07b4-60c7-4ddf-bf77-59ca8f5e8d95 · outbound

This paper cites Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.734835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:60a77615cc181540379f12c953f53a119dbee25b267424f0236b29017ff1a55f

Observation b1ac5cf7-af69-42a9-9bc6-3c9c1e730a8f · outbound

This paper cites Large Language Model Routing with Benchmark Datasets.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Large Language Model Routing with Benchmark Datasets

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.758141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:37daec3d5644f160f3779507cc38f88888150b6205b91a9c03e4e37d64d93951

Observation 7cf5c060-c938-4dc1-a8aa-4cf6d7c7d82c · outbound

This paper cites Getting more out of mixture of language model reasoning experts.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Getting more out of mixture of language model reasoning experts

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.552140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:25d8c76ccb09fbcd07eb12e46f09fd763025d50a5ef1da650bcf2d0c76f4505f

Observation 5c201e8c-667c-4e71-bc76-ef31ad20ad3e · outbound

This paper cites Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.654514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:cb34882d65678d5ee2e0bc7b4c53189f1d7a73a3354018c195818a50af6c1015

Observation 81d3f557-0f66-4553-a33a-5c199cf3726e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Gemini: A Family of Highly Capable Multimodal Models

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:58:22.717647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:243b43e8b40c44a6ac6c0221309e23a432009e59d831f29d4a1ca46b4713f18c

Observation d7272572-8b78-4577-b097-f6d69785a3cf · outbound

This paper cites Llm-topla: Efficient llm ensemble by maximising diversity.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Llm-topla: Efficient llm ensemble by maximising diversity

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.548975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4e97d35fcf2ccba65b07ebc1022ea78a20fdd7704b5a48933da51c28e1fa1bc7

Observation c942e07e-75bd-4c01-9bf3-d4bc21431884 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.739122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:5c0c6aa6c66144c03150167789bdc92a7fb59606cb43be15052a3a7dac0d808a

Observation 20fe7a45-b6c6-4b2f-92e9-a1c674980534 · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.747282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:9b7f6f6c0eae14fca8dfe2952cb84e54eac2ec08b0dfb10ae98cd15a4ec3f68d

Observation 5755aa23-64ce-41c9-bd4a-74d829f25bf9 · outbound

This paper cites Koala: An Index for Quantifying Overlaps with Pre-training Corpora.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Koala: An Index for Quantifying Overlaps with Pre-training Corpora

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.761678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:3ef7d6b3d02a44cd8983aa12d3b872107a4772d9db56030dca2aff56522b1963

Observation b6881a9e-d489-486b-a526-a7ddfbc891ec · outbound

This paper cites Large Language Models are not Fair Evaluators.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Large Language Models are not Fair Evaluators

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:10:42.624770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:0f4e3be04b92b017a28446d7b45627464cd5ec2c8806c386737d45f410103437

Observation 62dcc8aa-f649-4537-ba7d-a582191aa3e6 · outbound

This paper cites J., and Choi, E.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process J., and Choi, E

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.695340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:54408046adfa1cdec7f3cc675a43f7c6c2133c27e273220a42c434008dc376ce

Observation 18214f72-66ed-49b4-8c62-da18d9e622bd · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:58:22.754673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4410a842462c280f0f2e0b1f7dc04ae4a3aa25df46234cc58bdd8b4559e3042d

Observation 388d9b6e-b63c-45f9-90f4-387e4da831d4 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.687864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:bdc95f3d94aaeade35157cba54cf930e4bb6c8c795056d6ca842b105a3165694

Observation 14fa69d5-1c59-4bde-965b-0ff64564be3d · outbound

This paper cites Bridging the gap between different vocabularies for llm ensemble.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Bridging the gap between different vocabularies for llm ensemble

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.545851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:1f5ffcd59f50b9801ccdb69dabfcc0582d8f78e247bcfca5a78c6dc8ea02fa87

Observation 21509a30-18a5-4b4e-b38e-c577ca7d6648 · outbound

This paper cites Hit the sweet spot! span-level ensemble for large language models.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Hit the sweet spot! span-level ensemble for large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:03:23.617144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:19a764c3ac877f9348ab9be80f01e6189da106f82cb2e7662af7813c9e65ec75

Observation ed024f32-12ae-4d87-aa21-0cf0c8fbd8ba · outbound

This paper cites Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.713390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:05ae84037a9825b7de51a957d8348b0c23b81272898637a2f1e3aeb829842f46

Observation f9e5e6b7-5abd-48cc-b993-8c43459aa85a · outbound

This paper cites WRENCH: A Comprehensive Benchmark for Weak Supervision.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process WRENCH: A Comprehensive Benchmark for Weak Supervision

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.786991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:4e3df559396f6b0b9690a4b5e50e395c3c0dfaffb43579f23fae473e711b8492

Observation 229c22b3-5779-47c8-ae9c-cc972a5893cf · outbound

This paper cites Wider and Deeper LLM Networks are Fairer LLM Evaluators.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Wider and Deeper LLM Networks are Fairer LLM Evaluators

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:58:22.730880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:32c8da40b5fe5985ebb3fa3a26e00e78e1bcfe3eadc62ed0406372c8cfe886fc

Observation 61c09b83-5ba2-4c72-9f50-17fe12a14d2d · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:03:23.614587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:85a8fcd757d94f7ddf8899a0904ee5d73f40ae0a53a2c9a2508342d17b5f5d65

Observation 77b3aeec-5c9b-4203-83b3-d751d4d064dd · outbound

This paper cites Truth inference in crowdsourcing: Is the problem solved? Proceedings of the VLDB Endowment, 10 0 (5): 0 541--552.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Truth inference in crowdsourcing: Is the problem solved? Proceedings of the VLDB Endowment, 10 0 (5): 0 541--552

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.524467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:bffeb42a449e69a7c6817a0686a711aa29a44eaed736ebd73d6b188d5a0a60e2

Observation 73b9649b-b3ae-464b-969b-ffdc0a8a2834 · outbound

This paper cites Lima: Less is more for alignment.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Lima: Less is more for alignment

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:01:14.521685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:581246822d86c7bd6dfa0bac4fe680de8718bc4627778ebf3a0d468eeb33e6e5

Pith citing papers

Observation 209c492a-0f9e-475c-958b-93dd37620545 · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:48:22.803116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:d4e4af993bbb08f73a989f7968e0726132c0960f5197fadf870750881ce91289

Observation c0395aa3-96be-483e-8353-ae36d6fd7d10 · inbound

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems cites this paper.

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.171775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T00:05:31.780655Z digest=sha256:15b7c7d34adb0a8c001d6f4f22b703804d574b94f4a6e0d52bc50a0af8da22a9