Pith. sign in

Paper Citation Record · LEDGER

Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2407.10817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.10817 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.545529Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:27.459119Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82323832-097c-4040-9c37-f788075d9b54 · inbound

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set cites this paper.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.444884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.444884Z digest=sha256:c8778ba695d0214b75e3ccb132e658f717b07b88ca6a496cbf236e989a330dd1

Observation 2976c262-f0a4-454f-b32e-3bf84bdadcd8 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.575421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:29e6fc733fa45192cd51f93c868be3862e0d2c69a9a841b94d784e172fd57ed6

Observation 97accfb3-d797-48ae-9bd5-3cb04c4287d0 · inbound

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models cites this paper.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.707647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.707647Z digest=sha256:d03be3f4311457b735aa12481c0dd410390d14757414c7141460355341814be6

Observation a5d8d982-0432-454c-835c-fff7d9de49e3 · inbound

Atla Selene Mini: A General Purpose Evaluation Model cites this paper.

Atla Selene Mini: A General Purpose Evaluation Model Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.780674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.780674Z digest=sha256:4f8da0a8266b9ceb1bfefe9bc2f45b09e7d69b24ff2f24cccc2f35d8adf0fbbb

Observation af6ecb3f-2016-40f0-95d1-db9311c2b413 · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:06.997800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:06.997800Z digest=sha256:0889b1ecc81acc533509469bdf3d689294b1a1992b05eb05ffa649308753a65a

Observation 795ba532-f5e8-431e-a83c-0bb54cc36c61 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.469756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:18f6d79393fd2f75fedbaf8b36ae827b1457ef2a6f86266e08c332503287e68f

Observation 4396bc91-83a8-44e2-88cc-fb980e50cc79 · inbound

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators cites this paper.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.545529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.545529Z digest=sha256:6d62f274fd2237dfb98ad35b9f4689c3f2af97791bcfb4324e95591cda52baf8

Observation d495196b-ab18-4795-94f8-1b9ed58bd2d9 · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.076185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.076185Z digest=sha256:bfa1319f9a6d348c7e7270d1857e245cb7b40d3c7b183b26581082a2357681b0

Observation 23974def-50e9-40c3-b46c-5dc512e23a69 · inbound

A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety cites this paper.

A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:13.271655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:13.271655Z digest=sha256:06ad413fff845d499b89c45613dd46edacde56def492d632392496db8a479eee

Observation 240afa34-a1fe-46c9-a726-ea6ec02e62c6 · inbound

Do Biased Models Have Biased Thoughts? cites this paper.

Do Biased Models Have Biased Thoughts? Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T22:39:28.084796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:39:28.084796Z digest=sha256:7819b0354223fabdab11442879ebf7aec4a8d15b8fd0555e05f957cb9ec21c20

Observation 9da63934-2330-48fc-a9c6-0ad8e789d7aa · inbound

CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments cites this paper.

CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:36:50.269769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T20:35:29.338034Z digest=sha256:de4402c51026af5a50a6dfb7d9b176811d2981e1aee104d78915a17afe43ea2b

Observation 9fa42a41-edc6-4d17-ac23-407057dcd7dc · inbound

EvoSkill: Automated Skill Discovery for Multi-Agent Systems cites this paper.

EvoSkill: Automated Skill Discovery for Multi-Agent Systems Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:25:05.453288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T02:25:05.418500Z digest=sha256:7a13bbe33cc3e0a7dd92ecab2cbb031b9a2a2d4ab104e7eeb65cb12e7c452fff

Observation bb20f579-9ca9-4c12-84b4-8ce6d72d0dbc · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.460931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:d0376981b7197df38fa3d13c8cb587153bab01cffc97ceb859102419da8a07ad