Pith. sign in

Paper Citation Record · LEDGER

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.10403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10403 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:13.658964Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T15:14:05.529138Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ef315790-9b74-4f37-80c1-6a150431efaa · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.584381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.584381Z digest=sha256:d031b738a0c44d35893a8635a15300c1608ab2dd4b4abc25c6eb9a65ee02cab7

Observation 6acd8fde-133d-4d07-9f32-eedf388fa975 · outbound

This paper cites N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.; 5 Gonzalez, J.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.; 5 Gonzalez, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.882859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:08.700690Z digest=sha256:165737243e97d9c1244d5cc4840c2897f851e8f2fe27298823c8d393cd4f2ca0

Observation 113822d6-e3d9-4b00-a9c9-05185fb377f3 · outbound

This paper cites Advances in Neural Information Processing Systems 2023, 36, 46595–46623.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Advances in Neural Information Processing Systems 2023, 36, 46595–46623

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.658458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:08.819739Z digest=sha256:f05f5c1358c31d71ba7b9a7bd55ed68ac8f13322222270bfc07e7ae8882aff66

Observation a029ec1d-50a9-4bef-9736-cd09ca5b1c94 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Agent-as-a-Judge: Evaluate Agents with Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.958962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.958962Z digest=sha256:1ebd078f842144b1c86e2a1bd7f8035cdf55be7fd4ca8ae9307c31a2a06d3d7d

Observation 290c994f-6f21-4700-9db7-cf29e9336f1d · outbound

This paper cites QuRating: Selecting High-Quality Data for Training Language Models.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation QuRating: Selecting High-Quality Data for Training Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.115749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.115749Z digest=sha256:8a5f4fe4613edc7cb9ce07aea9c2bc357fecd8de2350387e9ac8b4429b853b52

Observation a2b0b47b-a7e8-4a48-85b5-8d71ab66010a · outbound

This paper cites F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.485017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:09.307806Z digest=sha256:a7e2e1cebbe2d9b42e170cce5886395845afe63cda4769e65ea13c362882492a

Observation a74c5cd5-db0d-4001-9307-d494038856ad · outbound

This paper cites an unresolved cited work.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:18.243036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:09.434801Z digest=sha256:37735d574de8966c9036faf64c9da000adcfdef6f33350b47c8334824d99843a

Observation f980c253-bbe6-46ce-83bc-1e9c8ad4a062 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.627591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.627591Z digest=sha256:bde61b4e5a3bee93944733b194cc1d5198aa14a48157913b43903c2db20e286e

Observation 5d5285b6-e5ed-4824-bb79-f08bef70f2f6 · outbound

This paper cites xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.757567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.757567Z digest=sha256:1c30a92ee689dd15bb2dbc56906f2a8059bb24601eadec5c86e1c9ad12bd7dde

Observation 19dac899-d440-48c8-8b57-7f5e6991fc71 · outbound

This paper cites Discovering Bias in Latent Space: An Unsupervised Debiasing Approach.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Discovering Bias in Latent Space: An Unsupervised Debiasing Approach

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.900260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:09.900260Z digest=sha256:6fc0b9bda3dbafc3931e5eec319de8bd5c93352a87b9b51f7a95dbecb5b00126

Observation d5d4530c-185d-4291-8809-faa66fc8ed31 · outbound

This paper cites H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.013754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:10.023607Z digest=sha256:fd5efbd5e6bd63d67ef557506f3638748e84c381608dd6147ad3c43610f8937c

Observation 8f607aac-142f-4e00-b349-0551a68d1b97 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.196137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:10.196137Z digest=sha256:dbc240160fa7b02167d7bee77a47e830a7d43aac334ff77ddc7a4e2f656fddec

Observation c07bf243-e9b2-49b9-9f36-fdb29e17b4f4 · outbound

This paper cites J.; De Sa, C.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation J.; De Sa, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.805382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:10.307887Z digest=sha256:69f9d4160afb50a6e5e5bb43164594414abaa2d323ddcd02e9736634ef833684

Observation e8c7f540-ff3c-4302-b88c-0909217293e3 · outbound

This paper cites H.; Ehrenberg, H.; Fries, J.; Wu, S.; Ré, C.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation H.; Ehrenberg, H.; Fries, J.; Wu, S.; Ré, C

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.526067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:10.487327Z digest=sha256:57005874a86b6b5fa5411d6f315aafef4703ece03a1630ef12eae2828f188a31

Observation ca798256-8249-4e82-8760-09a04738b7a4 · outbound

This paper cites Training complex models with multi-task weak supervision.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Training complex models with multi-task weak supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.359182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:10.653766Z digest=sha256:da50fc3aba7e9861441729bca700436bedb9bed9114eb0c855390eca60ba8135

Observation affb8cf4-53ab-49ed-bdea-b30ec53d5f5a · outbound

This paper cites Fast and three-rious: Speeding up weak supervision with triplet methods.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Fast and three-rious: Speeding up weak supervision with triplet methods

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.089083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:10.819345Z digest=sha256:1fc6908e8db06d2c6c7b53927cf5193fe2c831c11843247f8764b5d0bd6d19db

Observation d07bcf6c-d883-4dbd-83f3-8cbd8345e7f8 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation RewardBench: Evaluating Reward Models for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.968389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:10.968389Z digest=sha256:5ce1da97ad5ef7e75bd75a52f40af91ba8469d304bd55f7213e90bfb3b303513

Observation 0e7e9cc8-1810-4351-af7f-ae67533c8332 · outbound

This paper cites The Twelfth International Conference on Learning Representations.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation The Twelfth International Conference on Learning Representations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.956616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:11.147019Z digest=sha256:7a5e0e095b9ab397ee6b35257ce5bfd1ca2f79d8d0bc35144ecb12ec5d9cdace

Observation 08063b97-5a06-43dc-b46d-d612811d8fcc · outbound

This paper cites GPT-4o System Card.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.322747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:11.322747Z digest=sha256:c04840ee350e15d2b6f0f3ef975b8ef89a0704e0614f402ffce18625cebb72c9

Observation 7849f027-4f10-4708-9357-5ee5d7c19757 · outbound

This paper cites an unresolved cited work.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:16.873686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:11.471540Z digest=sha256:888bd1ab9b1834a5fce34fbe47236e56919cdb820b1afd0c3ec1a437922856e4

Observation 292117a4-594a-45ad-826b-945007f1d7f5 · outbound

This paper cites P.; Fishburne Jr, R.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation P.; Fishburne Jr, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.587401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:11.632589Z digest=sha256:c9a6f6e85b24b617b2abe34f291b1d3b77b8e2a7feb3cc96973ee940cffbfb51

Observation 89546b02-b98c-41dd-a339-c03f10bba31b · outbound

This paper cites FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.299829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:11.812496Z digest=sha256:b276d08e27c225db22677f397f1c0f14590b51cd68d5d6a8162cd849efeea39e

Observation 8562dfc4-bc17-4ef9-bcd6-9bcc0311d637 · outbound

This paper cites Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.018586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:11.939726Z digest=sha256:ef87062a22a303eeb8b3a6c87cb41a7aff7e98179dd1007712f678d4cac14a49

Observation 6bdd328f-b03a-4208-8ac1-24a3eca19819 · outbound

This paper cites Universalizing weak supervision.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Universalizing weak supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.672364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:12.071052Z digest=sha256:7b86975092018e4492473a345f29e2cc69997d96cba0e423ceed6ced00610578

Observation 5711fb54-8843-4ca0-8982-a659fa51253a · outbound

This paper cites Weak-to-Strong Generalization Through the Data-Centric Lens.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Weak-to-Strong Generalization Through the Data-Centric Lens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.353380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:12.237890Z digest=sha256:fd61dc43f72f4e20d6c8b92119d84ab52e3c34d44a45904bea8f3a342b313001

Observation f6833448-268e-4a5a-8a56-f5b92c2d2589 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.060847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:12.346391Z digest=sha256:6da13ae8a80f9950ab2203e3cdc8d41cd37c475fe245f8b6c0a58a521f9f1a99

Observation bb382b29-30e9-41dc-8491-3df9c32420b7 · outbound

This paper cites GPT-4 Technical Report.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.492714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.492714Z digest=sha256:be5eb49666d86f8c7d83b15cbcf9045d5fbabd26ac685d462a29ae9b723bec0b

Observation 1a7646b5-cff3-46c1-b0d7-07c9a51d592e · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Gemma 2: Improving Open Language Models at a Practical Size

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.651746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.651746Z digest=sha256:56f73c74de0b0c08fb4442e8e72a18e86c5927cb6b8a9f2b62570d36abfef90a

Observation 733cb78b-219d-4772-a6d9-9a4b30d992f6 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.759896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.759896Z digest=sha256:44b86a321b5abeeb8ae8c61426bdf84507d4968da2179a4c0208250951f76235

Observation 86e8dac7-4bd3-4ca3-bfa5-b16cfe8053b8 · outbound

This paper cites Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.934810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:12.934810Z digest=sha256:8fa71da38c988580c455c6c18c992007f9bee35872923f420f0628bf845839ef

Observation 4c2ad085-1fa2-444a-9242-9cd64b26d309 · outbound

This paper cites The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data Annotators.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data Annotators

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.817946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:13.041107Z digest=sha256:20721554bf0a7f88f45048b8c770733ee07697f751bd86d6dcfd420f3bcfc136

Observation abc0dae9-433c-4b1f-96ed-5fa5cfe702e3 · outbound

This paper cites ScriptoriumWS: A Code Generation Assistant for Weak Supervision.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation ScriptoriumWS: A Code Generation Assistant for Weak Supervision

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:34:13.922531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:13.235686Z digest=sha256:112b697290e0b4633f2161f14cb9e8ff1ac65a6232be37eb4963fd6a783f076a

Observation 5e2df8e1-3ce8-4888-8d33-589c4ba98afa · outbound

This paper cites Autows-bench-101: Benchmarking automated weak supervision with 100 labels.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Autows-bench-101: Benchmarking automated weak supervision with 100 labels

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.616136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:13.376326Z digest=sha256:9920cebbfca593d566e5d94ee198fb8c8a62027aa26144df21cd0a797a16432e

Observation cc286cef-bb27-4d85-80b2-35448f585fb9 · outbound

This paper cites Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:13.508340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:13.508340Z digest=sha256:d103360de5d2ac8897af9debe9d333b6b6d6244df60ce38265e6878c469792ca

Observation 997eca7d-7afe-4568-bec9-8d04b8b2df7c · outbound

This paper cites ""Calculate readability metrics for response.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation ""Calculate readability metrics for response

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.331717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:34:13.658964Z digest=sha256:45d873a1884d6ce273a64f7d66595e92cd8bf8633d27fea0c15b48ba73500bd7

Pith citing papers

Observation 0f017e17-0894-4340-9915-1d2ef361703d · inbound

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation cites this paper.

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:50.709981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T20:02:40.980538Z digest=sha256:e4b8796739fbb086f5c3ad5b61fd9e5646f8b5467e971307bec46ffda3566228

Observation 286f4207-5d0a-40c4-b6ab-05c718a80828 · inbound

Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation cites this paper.

Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T19:15:44.221307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:15:37.744461Z digest=sha256:2f015559074f56f6d02d9302d7ea9a1a60e21076439468c142b388bd47b71f72

Observation b7d92186-e247-4ba5-ae91-d64872b60041 · inbound

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning cites this paper.

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T15:14:05.529138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T15:14:05.529138Z digest=sha256:e9a3ec1e7455daa89bb2e7cf3ff7e53196bcd180de20a9ef5201f30527803e03