Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:02:43.344256Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.17178.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:02:43.344256Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 65079401-af76-451a-9e28-8fc5522a6d18 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183ee70d-39bc-4dff-b816-464fa854bc3a · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897fe0da-d751-42e4-ba76-d326a6433776 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e01cf3-5634-4125-9be2-74bf851d986f · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3d30b0-21f4-4af7-a3c6-c27c336aea3a · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263757ef-a977-46e0-b7e0-e215def5fd5b · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost From general LLM to translation: How we dramatically improve translation quality using human evaluation data for LLM finetuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e31defa-5c81-4c03-9801-c0ad9a2e5c8f · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1d22ed4d-dcb4-4b55-8273-7c8f6bfd21f7 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474d9235-48cf-4b3e-bf98-cccbc29909cf · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f62d0ab9-a5fa-44a5-94af-b064f3d14fb8 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Does Prompt Formatting Have Any Impact on LLM Performance?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddaebaf-ba4a-4199-a5ec-917255b161ed · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4678c2c8-d8f3-4f49-be03-d3b2c4acf58d · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Bag of baselines for multi-objective joint neural architecture search and hyperparameter optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a91289d-172c-4664-ac42-2b1f64d76a98 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Almost optimal exploration in multi-armed bandits
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 913c0f86-f6c8-4513-b647-2dad7fa3c3c5 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aabda245-64de-4670-94e0-ffa245e30dc4 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5107405a-2842-45b6-bf1a-7cca8549f606 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4fcab8b6-7dad-45fd-9808-56dd4d009178 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost E., Stoica, I., Mooney, P., Dane, S., Howard, A., and Keating, N
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3eb3c3e8-cabe-495c-a774-8c1a07889219 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Aligning with human judgement: The role of pairwise preference in large language model evaluators
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e34c35c1-5b09-4ace-8d71-ca3ab9e5db8d · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ff803e47-8d9f-4f93-840e-4dd68681c9d5 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1a9170-4b93-41c0-9fd0-fa6ecc3be8dd · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost LLM Evaluators Recognize and Favor Their Own Generations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a314ee06-a6a7-460e-b319-dded7a73726d · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost A multi-objective perspective on jointly tuning hardware and hyperparameters
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 134b22ef-4814-468d-8518-f211b8e9749b · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Multi-objective Asynchronous Successive Halving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89c57b62-b27b-49b1-b300-cbfcf43901f7 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Efficient prompt optimization through the lens of best arm identification
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8af58ca6-d7ef-42f4-ab0b-478c1f0e391e · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Fine-tuning and prompt optimization: Two great steps that work better together
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35c2ea36-9ad8-41a0-bd50-685b1ab775fd · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Panda LM : An automatic evaluation benchmark for LLM instruction tuning optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc38d7dc-6883-4f06-af7d-415f3f877f7b · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e378df08-18d0-4050-9482-ca4aa00243ba · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3439735-3525-4b8c-946d-5614fe46b33c · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1032083e-4b32-49cc-894e-fc61a3d3e160 · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d31f7b3-89de-47ca-9839-e8f9055ad91a · outbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.