Pith. sign in

Paper Citation Record · LEDGER

Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2408.08808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08808 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:51:52.750466Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:27.434263Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 851c67ec-feeb-488c-aea2-44d22c07a26c · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:43.951672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:1fe01ed372b9ef840ba3065e789b1ab6ee484fb7401a1f70a027f7f768a98ba0

Observation 4aec9680-629e-4283-b8bc-45e09261fb09 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:34.651161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:9897ca2a2657f3f2b5f3d5023260b14adfbd6d8604907dd3d8ba8ac35b1c616f

Observation 6c8d3e87-a505-4894-8225-8bfc32a0d4bb · inbound

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences cites this paper.

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:52.750466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:51:52.750466Z digest=sha256:8817242a3443c7e3df0ce09f40ad5e8d0caa41506d898938786dd7fd01988108

Observation 4f27e840-ed48-4081-b638-a4dbc9f94157 · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.479008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.479008Z digest=sha256:e29cb5c61b75a312d1a8f300edfc230b626a1bd9be6e3aa8666b0db9f9e01a2b

Observation 4fdc0035-c199-4ab5-a018-885e7c3b3cd1 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:34.674305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:34.674305Z digest=sha256:f24b88ac141cba8afe28a1212ab00eec2b93e4cb8a2a9d1c708042dd8c7cc248

Observation f6561b71-fc11-4d01-bc47-dcf706c93cf2 · inbound

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models cites this paper.

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:41:55.985356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T00:39:02.015913Z digest=sha256:a24eec5080b0a223480d0b649cee6deeef5a73e5dd11c6789ad925215ee509e7

Observation 615fee28-b7c6-4363-a9b4-0f3914f2d9c3 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.435589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:15d7f3c133e5f92847251ac2fc2d54ac3db663e1c627b20223309197a323aa52