Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T19:57:03.999154Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 2 inbound Pith citation observations for arXiv:2512.23213.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T19:57:03.999154Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T00:05:31.780655Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-29T00:12:50.170528Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 55f87f30-114e-44fc-9a70-f0693f5dde68 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process write newline
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ba0712f-48ed-4533-b3de-3890f8afb36d · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e679e594-83fb-45be-bcd5-3f36a2caafd3 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f312f3fe-0aca-4a9a-86ab-9b0882f57ccf · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e51aa58-0d1f-4845-a47a-8b37440cca4c · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process An automatic and cost-efficient peer-review framework for language generation evaluation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e17bd6c-4686-4883-acb4-4a0cc4e71911 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf20bed3-26db-467b-bd96-f3c06a0a6de4 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Adversarial learning from crowds
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7be0574a-f48d-448b-befd-80dfcf698294 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Structured probabilistic end-to-end learning from crowds
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe00d935-a89a-4111-8545-e782cfdeecf4 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Neural-hidden-crf: A robust weakly-supervised sequence labeler
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6026d2f4-0d2b-44b7-877c-2be0cc6bab06 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7285815a-7cfc-49d1-8dba-63238100db8b · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fcaf17fc-5d76-4cff-9683-91ae2556ed16 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process PRE: A Peer Review Based Large Language Model Evaluator
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 83c10816-cc1f-40f9-b4ef-337d5c9f9655 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 548d295f-3c3d-4a82-8639-c6da529a96d1 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 287c74e5-7450-4305-bced-e362b1dcc415 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c3853e6-de2c-4daa-9e8f-bef746d4fe2c · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Maximum likelihood from incomplete data via the em algorithm
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67c704b9-ccaf-4ab2-9730-69ec09908fc0 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process A survey on ensemble learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1ef1cf4-96be-47de-ad2a-bba50997bb9a · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 606e6660-3212-4f75-bc41-99a7f4d12071 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process An introduction to latent variable models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d2598b1-2d0c-4087-80b3-eb351084f506 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process GPTScore: Evaluate as You Desire
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a89d433-5782-4ac5-acb7-3618d5d720bb · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Deep learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5ea98ae-841e-408c-9fc0-da4585180075 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process A Survey on LLM-as-a-Judge
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59898454-34ea-4df1-ac40-9183f50218a3 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Smoothie: Label free language model routing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 459e3f86-be96-4fdf-a295-e402dad61d8c · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Measuring Mathematical Problem Solving With the MATH Dataset
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54063440-a498-4b80-a97a-4c1d3acd80ca · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Language model preference evaluation with multiple weak evaluators
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e0b36de-b544-4ed8-9f4d-a51681672a1b · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Ensemble learning for heterogeneous large language models with deep parallel collaboration
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5529df80-0829-4460-bc9d-24cbda953884 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b22e705-3058-4af8-9e7b-38f53e63a992 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d43b8ff0-58d4-4439-988b-4f5774d5b637 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85b71a9c-0c22-455b-b72e-5f1f10478d90 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Little Giants: Exploring the Potential of Small LLMs as Evaluation Metrics in Summarization in the Eval4NLP 2023 Shared Task
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 604caec7-42f9-47fa-9e96-8e1266405966 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process From generation to judg- ment: Opportunities and challenges of llm-as-a-judge
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d17ddcf3-9d6f-4481-b86a-9b1e06925d67 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6d20cde-7c1f-4e28-90a7-f03fc4a9e0e8 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process More Agents Is All You Need
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 168e4c3c-b897-400c-bf24-c3ca69b788c9 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e89f12e-e1da-4938-995a-5c759c8e5f21 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Cool-Fusion: Fuse Large Language Models without Training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45645ed7-e53e-4f05-a461-cdc540d49cdc · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09e1e27d-d6b9-46a9-94e4-d2d1b5705446 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Urg: A unified ranking and generation method for ensembling language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2772a185-390b-4434-ba2c-9db63c96ba37 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process RouteLLM: Learning to Route LLMs with Preference Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0b38780-71d8-4aab-958d-29599baf0023 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process A Multi-LLM Debiasing Framework
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a30c07b4-60c7-4ddf-bf77-59ca8f5e8d95 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1ac5cf7-af69-42a9-9bc6-3c9c1e730a8f · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Large Language Model Routing with Benchmark Datasets
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7cf5c060-c938-4dc1-a8aa-4cf6d7c7d82c · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Getting more out of mixture of language model reasoning experts
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c201e8c-667c-4e71-bc76-ef31ad20ad3e · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81d3f557-0f66-4553-a33a-5c199cf3726e · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Gemini: A Family of Highly Capable Multimodal Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7272572-8b78-4577-b097-f6d69785a3cf · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Llm-topla: Efficient llm ensemble by maximising diversity
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c942e07e-75bd-4c01-9bf3-d4bc21431884 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 20fe7a45-b6c6-4b2f-92e9-a1c674980534 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5755aa23-64ce-41c9-bd4a-74d829f25bf9 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Koala: An Index for Quantifying Overlaps with Pre-training Corpora
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6881a9e-d489-486b-a526-a7ddfbc891ec · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Large Language Models are not Fair Evaluators
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62dcc8aa-f649-4537-ba7d-a582191aa3e6 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process J., and Choi, E
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18214f72-66ed-49b4-8c62-da18d9e622bd · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 388d9b6e-b63c-45f9-90f4-387e4da831d4 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14fa69d5-1c59-4bde-965b-0ff64564be3d · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Bridging the gap between different vocabularies for llm ensemble
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21509a30-18a5-4b4e-b38e-c577ca7d6648 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Hit the sweet spot! span-level ensemble for large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed024f32-12ae-4d87-aa21-0cf0c8fbd8ba · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9e5e6b7-5abd-48cc-b993-8c43459aa85a · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process WRENCH: A Comprehensive Benchmark for Weak Supervision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 229c22b3-5779-47c8-ae9c-cc972a5893cf · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Wider and Deeper LLM Networks are Fairer LLM Evaluators
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61c09b83-5ba2-4c72-9f50-17fe12a14d2d · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77b3aeec-5c9b-4203-83b3-d751d4d064dd · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Truth inference in crowdsourcing: Is the problem solved? Proceedings of the VLDB Endowment, 10 0 (5): 0 541--552
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73b9649b-b3ae-464b-969b-ffdc0a8a2834 · outbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Lima: Less is more for alignment
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 209c492a-0f9e-475c-958b-93dd37620545 · inbound
Policy Improvement Reinforcement Learning Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0395aa3-96be-483e-8353-ae36d6fd7d10 · inbound
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.