Pith. sign in

Paper Citation Record · LEDGER

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2411.15387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15387 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:27:57.472919Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:43.460215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:12:44.210234Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2370a0a-672d-497f-8ea2-e4fd8fa89166 · outbound

This paper cites write newline.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.295402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.295402Z digest=sha256:1e906c0679ddd34219a345c772368bf2474e48691b5a47bd23aa588b3a9f8999

Observation aa36a654-e550-43e4-9213-79c5aad0fa78 · outbound

This paper cites GPT-4 Technical Report.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.301443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.301443Z digest=sha256:96224e22e696e356762f56f615fe55a4a8fa0223b459fa8670077f351697026d

Observation 765fb702-624b-4d61-8a61-ddbaaa0a1c59 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.306616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.306616Z digest=sha256:111e9a72e93771bde718be9804464f9b64848e4d7d6fdf1bc7cef70aeffbd5c0

Observation 3dcade8a-0a5c-43b3-8cef-98894326297d · outbound

This paper cites M., Kanojia, D., de Souza, J.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set M., Kanojia, D., de Souza, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.204094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.311576Z digest=sha256:8cf2fbcf007e3cf3728a0d541a9d75f4910ba39b0d696c44bc2f44a66391d212

Observation 4b980d68-92df-4fd5-9dfb-96c22b49e959 · outbound

This paper cites an unresolved cited work.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:27:58.188254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.316049Z digest=sha256:554dbfc2bdf1c45623500b94a7ac56a9600756cc676100bf858974d029242cd7

Observation 2b2bf236-a994-413c-a1eb-17e9eb7896a0 · outbound

This paper cites "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:58.041465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.320397Z digest=sha256:69b669f1e1e794d91ec91fd0378828fde36493580a6aaa70d52a19008f615c3f

Observation 39a3ad6b-604d-4c7a-8f84-3b60f3ed1628 · outbound

This paper cites Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.324998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.324998Z digest=sha256:5abf46f570764e7e98342f45834d41798f76dac9c1fc54879ddb1665aa77c552

Observation dd017cac-cb1f-4e0f-bfb7-b32a7ec95626 · outbound

This paper cites The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.330246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.330246Z digest=sha256:a413af13a967d0cb76d491b76d736d9d7b86c14e0ecc3a0b5f8595720499a0d8

Observation 73e525a6-2169-4b4e-bf02-4281df5ba786 · outbound

This paper cites Experts, errors, and context: A large-scale study of human evaluation for machine translation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Experts, errors, and context: A large-scale study of human evaluation for machine translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.172302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.334819Z digest=sha256:6b77720f548acf7e8467e729235d95c7b839f77ff4525499c359f9a7e9a782b5

Observation 80c61444-66a1-4897-8179-48de0c385e17 · outbound

This paper cites Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.157513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.339072Z digest=sha256:5c9129b0adb0063de0822b54c8c5803c14c2f833dbce8ac7878e26c489403a9c

Observation 3a98697c-e771-411f-9872-735787c7add2 · outbound

This paper cites Are llms breaking mt metrics? results of the wmt24 metrics shared task.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are llms breaking mt metrics? results of the wmt24 metrics shared task

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.142333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.343580Z digest=sha256:63a3c7c00e095a3a1066ef13a21f2d3d2454342595fd244e8d3981a5572b25d0

Observation 162cdb0d-e881-4c64-8604-04fcfeee94d6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Gemini: A Family of Highly Capable Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.347977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.347977Z digest=sha256:8b9c72c08d7d343ccf0e177095fbdcde13b42044bea3dbbb2b7f69922150fc4e

Observation 3d97bd7f-daf8-4e7e-8dd3-0d97340ee51e · outbound

This paper cites Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.352788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.352788Z digest=sha256:4953226b1dddc405df6a83d32879cf88a9a9d7fc2f2dd8c7ab808eb1276b1c1d

Observation e99e2b29-2927-4043-87dc-330d4cace675 · outbound

This paper cites Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.357217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.357217Z digest=sha256:bf0eb1119d940cc1afc9349863727fcda2ca616c9c6e1e732e75f7333ce3826c

Observation 41501d35-36a2-46ca-a0ae-7dbdd4ab5b33 · outbound

This paper cites xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.361653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.361653Z digest=sha256:7de368bb07068c7956827f9e2d91c2106bb3edf098ceebec1f1aa7c7afddda0f

Observation 8f166c86-3490-4477-ab0e-aed8fbd52560 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.366287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.366287Z digest=sha256:97b9ae5efb760e89d9a12d8b83f7ca77f5950c5431f1607d778aebe2704fc20f

Observation 0f58b68c-837e-4e60-a890-89530f504d13 · outbound

This paper cites MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.370889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.370889Z digest=sha256:b3b81c09c6947fceb359de29c4be53dbe1d0cf8a169639f1023b5aca8d076c98

Observation 9d47ae83-cd20-4f87-866c-0c52ca05b969 · outbound

This paper cites Evaluating LLMs at Detecting Errors in LLM Responses.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Evaluating LLMs at Detecting Errors in LLM Responses

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.375768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.375768Z digest=sha256:1429bacdfda0eae23afb97733c7377394fe54426372587a1fca35913470a9d3f

Observation 62cbb40d-b465-49bf-969c-f7f871e7ca8b · outbound

This paper cites The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.380621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.380621Z digest=sha256:1baeef25c8069ac1b6cecfe50f904e6f8bb208b0f8a8f7b961a831958b7ea179

Observation 7b5fb24d-e0af-416f-ba39-48c523efc5fe · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus: Inducing fine-grained evaluation capability in language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.127272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.386232Z digest=sha256:910d419ae6756f7988cf311da2296b7ec83faadf88026fda326bb3e7f7c04295

Observation e798bd68-e6a6-4124-8cd8-31a2e27dc89b · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.390500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.390500Z digest=sha256:e08f84b582ccd2cafe8a9869412e6fcc8456123439645f2d6ce446e1dade3e25

Observation 56011669-a2a1-470e-9f09-bd69fa38be1a · outbound

This paper cites GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.394882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.394882Z digest=sha256:95aea93b7a6506c8d264cb74b6091630ae98ee0b18f64dd9060764a1012d7eb6

Observation da9dd46f-6bce-454b-b4fd-afeb106f7cc9 · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.399246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.399246Z digest=sha256:1f316e7f1f8c68f683093a1da93bb16f3f6a746a40e9af893462de7dcf631648

Observation 7e24ccdd-f903-4822-aaaf-b41255033d4f · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.403790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.403790Z digest=sha256:a9fc2ad94b97014cfc03ac3816dbc3d630acb0f1d527611b894345f6f2287417

Observation af837742-f587-423a-9371-0a94188de7f7 · outbound

This paper cites Generative Judge for Evaluating Alignment.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Generative Judge for Evaluating Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.408571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.408571Z digest=sha256:0829a4e4445d8e9ff61e7dceae500a89e79e13b1e75954f76cd9d9cec1f15716

Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · outbound

This paper cites Holistic Evaluation of Language Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.413062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.413062Z digest=sha256:07c70e82039330783008d45685720438e5194224c0aa39a2bb1821d0ff86ba97

Observation 79d59b41-a065-4226-a157-4df6031187b6 · outbound

This paper cites Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.112367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.417856Z digest=sha256:8b957a24133de13c22c4d9836e0139ab44f64b60cad04092ba00c9cffa8e2a1d

Observation aac3a4ad-0c5a-45fa-8caf-873e72cc8048 · outbound

This paper cites Training language models to follow instructions with human feedback.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.422180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.422180Z digest=sha256:99b1732c1a7834586e0f56a156695259150400360a12b001e034a4f1bdb1bdb6

Observation ee483ebd-aa30-4ff4-8e25-2c27482c604c · outbound

This paper cites COMET: A Neural Framework for MT Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set COMET: A Neural Framework for MT Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.426608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.426608Z digest=sha256:c5b68165d591dd2c535893fa5ffa6a81855634ded8ad76b0fc7c3ab677085292

Observation 2be859e2-4de5-4729-930f-44b2ad56413d · outbound

This paper cites Finding Replicable Human Evaluations via Stable Ranking Probability.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Finding Replicable Human Evaluations via Stable Ranking Probability

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:57.772867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.430960Z digest=sha256:ce377c37faad7f5da7b54858abadc60a855aa2e0d16c93678f4dc19a9b1daf69

Observation d831d68e-f4cf-4eec-8d3f-f8c3b811329e · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set BLEURT: Learning Robust Metrics for Text Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.435735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.435735Z digest=sha256:343a40664c6aabe09a4430c188158a53af66744c5982db904241e0f949b388c3

Observation 64af8eac-5edb-4c99-9728-25da842f435b · outbound

This paper cites A Benchmark for Learning to Translate a New Language from One Grammar Book.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set A Benchmark for Learning to Translate a New Language from One Grammar Book

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.440359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.440359Z digest=sha256:1e6eb6ec6575fc80ff683434265368607207972d21c499a58515f8c24f3926e7

Observation 82323832-097c-4040-9c37-f788075d9b54 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.444884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.444884Z digest=sha256:311750b6e667140d98d15e2ffc72d11ad3d0d1780a0a37e22f8800f340b38204

Observation bb60349e-0bc1-4a53-b4dc-57677226650d · outbound

This paper cites Y., Li, L., and Freitag, M.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Y., Li, L., and Freitag, M

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.449578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.449578Z digest=sha256:a3bb642f26528d606ca4566ac43fd2d7d2f490129059fc8b1c1d0f012bc41de3

Observation 8c0a4f2c-9491-4ecb-83cc-ce198d6d09b0 · outbound

This paper cites Understanding In-Context Learning from Repetitions.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Understanding In-Context Learning from Repetitions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.454353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.454353Z digest=sha256:eeee41018454f998ecf94e4db827482f0c917a3061c60e3b17b5c04fc4e2bb65

Observation 49bb4659-453f-4f01-a427-8747bd00d0f2 · outbound

This paper cites Learning from others' mistakes: Finetuning machine translation models with span-level error annotations.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Learning from others' mistakes: Finetuning machine translation models with span-level error annotations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.459572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.459572Z digest=sha256:e093adf51a5ef81c443b7225c37dadb3d05a68f397ac2f79b731e2462f2472a8

Observation 177f4ef1-ec04-471b-8832-8d477b046815 · outbound

This paper cites J., Wang, Z., Hwang, J.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set J., Wang, Z., Hwang, J

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.463831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.463831Z digest=sha256:9f17b7e0818f0329436a49aa9871fd8cc49800225ee7d2cddbd781ed1a66fa9c

Observation 1412af09-f169-42d1-9937-190f21bf84d6 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.468381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.468381Z digest=sha256:7d02e191de9ac8adb437ae115f70fe291c86194a1d9eb387b65d47f38c4b6af2

Observation 0500f9c6-192d-415a-89e4-b9e6981223a0 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.088220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.472919Z digest=sha256:9692aa90454b5d2408a59a394e5148a704270242cf2d35a961fd30cf2f63aad8

Pith citing papers

Observation 07ce0915-a39c-4b1a-ade1-b61e9c2304cd · inbound

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress cites this paper.

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:44.214904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T23:12:43.460215Z digest=sha256:1e776f3ea47035e6a5669bb1c588da0576b3041c89931e82d81914cd85fb1713