Pith. sign in

Paper Citation Record · LEDGER

Tuning LLM Judge Design Decisions for 1/1000 of the Cost

As of 11 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.17178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17178 v4

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:02:43.344256Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65079401-af76-451a-9e28-8fc5522a6d18 · outbound

This paper cites write newline.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.126190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.126190Z digest=sha256:64657de1694d266aaf8c55a8909ff2d13ea554485bd3eb14cd24e62905362804

Observation 183ee70d-39bc-4dff-b816-464fa854bc3a · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.135167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.135167Z digest=sha256:2249802fa291ebaf92d550678020a8985d0cac92675d3c392feb0a08182f63a0

Observation 897fe0da-d751-42e4-ba76-d326a6433776 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.141457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.141457Z digest=sha256:feb0b93efbd788fe09842b27f197ada063c48109efd30dc19c90e799437b8869

Observation a0e01cf3-5634-4125-9be2-74bf851d986f · outbound

This paper cites Finding Blind Spots in Evaluator LLMs with Interpretable Checklists.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.147194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.147194Z digest=sha256:a32eaea6f7c06a1e61c9cff45eef4538dc99cf2aa53596bfaa3d1ca82f829a55

Observation 0b3d30b0-21f4-4af7-a3c6-c27c336aea3a · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.152437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.152437Z digest=sha256:4d0cb3ddb9f3ed10ca4debf949569f34fab267b759b00584a6e2016982e57282

Observation 263757ef-a977-46e0-b7e0-e215def5fd5b · outbound

This paper cites From general LLM to translation: How we dramatically improve translation quality using human evaluation data for LLM finetuning.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost From general LLM to translation: How we dramatically improve translation quality using human evaluation data for LLM finetuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.157860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.157860Z digest=sha256:5f65a887cfea0d29de9aedfa76f2ea72f883ec96ebfbd1410c883e76e560e19f

Observation 7e31defa-5c81-4c03-9801-c0ad9a2e5c8f · outbound

This paper cites an unresolved cited work.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:02:43.955862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.163347Z digest=sha256:424e70689f1af69c969841effbccd047eb8f5a81a8712d9e9b0f7b7d828b32bf

Observation 1d22ed4d-dcb4-4b55-8273-7c8f6bfd21f7 · outbound

This paper cites Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.177549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.177549Z digest=sha256:452b34d51fdfb20dd178d787e15a99f40b3da9b15aa6589a48c11fd2beec0705

Observation 474d9235-48cf-4b3e-bf98-cccbc29909cf · outbound

This paper cites The Llama 3 Herd of Models.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.182676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.182676Z digest=sha256:f0295e5ff3800ce854e16bc5a7999f8a14cc5339f26e56d1f7a23841bec230d4

Observation f62d0ab9-a5fa-44a5-94af-b064f3d14fb8 · outbound

This paper cites Does Prompt Formatting Have Any Impact on LLM Performance?.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Does Prompt Formatting Have Any Impact on LLM Performance?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.191164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.191164Z digest=sha256:9764d6b92aada417852c8061a3513f733001f3d82d91b7bceea9ecb1916660a2

Observation 6ddaebaf-ba4a-4199-a5ec-917255b161ed · outbound

This paper cites An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.196827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.196827Z digest=sha256:61e6bbdd36c4d82b4e6cee7835723f78014c5dd13da93cc07c14a2f4859f9a7d

Observation 4678c2c8-d8f3-4f49-be03-d3b2c4acf58d · outbound

This paper cites Bag of baselines for multi-objective joint neural architecture search and hyperparameter optimization.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Bag of baselines for multi-objective joint neural architecture search and hyperparameter optimization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.940346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.202521Z digest=sha256:406124c9920d5fba7a6920b9b42fd5409386171944102d9ddfa95d156eb98c4d

Observation 5a91289d-172c-4664-ac42-2b1f64d76a98 · outbound

This paper cites Almost optimal exploration in multi-armed bandits.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Almost optimal exploration in multi-armed bandits

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.923053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.209724Z digest=sha256:cbb3953355c449a28f0704fb428dcab490b8b6d612d8c8602d85cd42207f9c11

Observation 913c0f86-f6c8-4513-b647-2dad7fa3c3c5 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.216786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.216786Z digest=sha256:51725797a73e4d6fc9eafae8181e51b251f620a08916e6f016b7a594431c8393

Observation aabda245-64de-4670-94e0-ffa245e30dc4 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.230563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.230563Z digest=sha256:8053349bf352b1de10406bf9268ba84a8d97615861b39bdd8f015e46367cdcc5

Observation 5107405a-2842-45b6-bf1a-7cca8549f606 · outbound

This paper cites an unresolved cited work.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:02:43.907606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.238897Z digest=sha256:4f5b291a3c2b30e01989ff552c3c185c98096e5f217b7c8ca91d3f354e16fb24

Observation 4fcab8b6-7dad-45fd-9808-56dd4d009178 · outbound

This paper cites E., Stoica, I., Mooney, P., Dane, S., Howard, A., and Keating, N.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost E., Stoica, I., Mooney, P., Dane, S., Howard, A., and Keating, N

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.890384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.247605Z digest=sha256:cf01b6e7749b6ce09b4d2a5b1b00117207c8e3c091e7bea5873eaa2506b3f5b1

Observation 3eb3c3e8-cabe-495c-a774-8c1a07889219 · outbound

This paper cites Aligning with human judgement: The role of pairwise preference in large language model evaluators.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Aligning with human judgement: The role of pairwise preference in large language model evaluators

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.869102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.258842Z digest=sha256:d9b6b18b41b8d0522eef57ba81fef5569ac575982e3f66e58b4b17c2316d1059

Observation e34c35c1-5b09-4ace-8d71-ca3ab9e5db8d · outbound

This paper cites an unresolved cited work.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:02:43.849295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.263443Z digest=sha256:5973fcbd66f4c45b26d84aee95f0ee4342857e6458fc56f86f1a0cd83d7bf699

Observation ff803e47-8d9f-4f93-840e-4dd68681c9d5 · outbound

This paper cites MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.268780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.268780Z digest=sha256:6130fcc65e367c09a37d814e3207a302dcfd6b83d8fc2eab950094829ce38ef7

Observation 0f1a9170-4b93-41c0-9fd0-fa6ecc3be8dd · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost LLM Evaluators Recognize and Favor Their Own Generations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.274806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.274806Z digest=sha256:5615c7392fe643a6787dba1e665539cded2ea868f10f97215ca29de203541416

Observation a314ee06-a6a7-460e-b319-dded7a73726d · outbound

This paper cites A multi-objective perspective on jointly tuning hardware and hyperparameters.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost A multi-objective perspective on jointly tuning hardware and hyperparameters

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:02:43.491145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.284693Z digest=sha256:9ecde996626962324f4a53a690ec902c9ecdb68294e5d6f351d18505c1f19b29

Observation 134b22ef-4814-468d-8518-f211b8e9749b · outbound

This paper cites Multi-objective Asynchronous Successive Halving.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Multi-objective Asynchronous Successive Halving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.291444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.291444Z digest=sha256:47e109c436f3f5aceb046f218b56a9d987a65adb527a72133311938d7efc0a2f

Observation 89c57b62-b27b-49b1-b300-cbfcf43901f7 · outbound

This paper cites Efficient prompt optimization through the lens of best arm identification.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Efficient prompt optimization through the lens of best arm identification

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.826367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.297699Z digest=sha256:eb5eba25eefc493603c472ed79105011f60d0b164fd99acaad04f57c719b2bbb

Observation 8af58ca6-d7ef-42f4-ab0b-478c1f0e391e · outbound

This paper cites Fine-tuning and prompt optimization: Two great steps that work better together.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Fine-tuning and prompt optimization: Two great steps that work better together

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.801370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.305459Z digest=sha256:858e59334e8033d13c66dbdb1e617bb967134a45e13c3227e6a03e61552d7482

Observation 35c2ea36-9ad8-41a0-bd50-685b1ab775fd · outbound

This paper cites Panda LM : An automatic evaluation benchmark for LLM instruction tuning optimization.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Panda LM : An automatic evaluation benchmark for LLM instruction tuning optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:02:43.770207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:02:43.310841Z digest=sha256:eaa4f16fb3248f56f7e7b1036fd567878eec4c1293c92b5b0b581095e78768c3

Observation cc38d7dc-6883-4f06-af7d-415f3f877f7b · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.317107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.317107Z digest=sha256:ec74219cedc8b2a22de3eecf80d3088452e95230426f78bd71607b08b054744b

Observation e378df08-18d0-4050-9482-ca4aa00243ba · outbound

This paper cites an unresolved cited work.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.323365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.323365Z digest=sha256:9a4c6d40e94b02346a2a47fdd22e0f61e8286866224c1973561944ed6715aba7

Observation d3439735-3525-4b8c-946d-5614fe46b33c · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.332839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.332839Z digest=sha256:9f809cc62e943411144b069f1c72d6434433de06259da46d0e331ba377e0c0ac

Observation 1032083e-4b32-49cc-894e-fc61a3d3e160 · outbound

This paper cites Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.338988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.338988Z digest=sha256:e6010085879ad8ea31e8b2b71580f65f302d72308bbb80867ec06ed939e53ca6

Observation 4d31f7b3-89de-47ca-9839-e8f9055ad91a · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.344256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.344256Z digest=sha256:d329c70fb858a0dd0f2705ddfe9ccd4611686ee61b277e6b06c0feedfce4aeb1

Pith citing papers

No inbound Pith citation observations are available.