Pith. sign in

Paper Citation Record · LEDGER

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2402.15754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.15754 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:47.053105Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11e491a1-2bd6-455c-abe7-d4b56addf389 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.341931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:9c432fc0dfdf8f3b1ce67cddad95138a872e321eb2443704ced4ed140958fc99

Observation 079f109e-1ebb-443b-b672-28b8b4f858e0 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.053105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.053105Z digest=sha256:742cad495b6ac9af5db30240392b03716bd710e9222a6a47db1b5d8b9918dc59

Observation 38b7b8a5-fa1b-472f-83df-95981658ea5c · inbound

Balanced Hyperbolic Embeddings Are Natural Out-of-Distribution Detectors cites this paper.

Balanced Hyperbolic Embeddings Are Natural Out-of-Distribution Detectors HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:07.186704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:07.186704Z digest=sha256:fcad6ef5d7c828b49b901b871e641ee759c3f1215f3ba705a24524c64efbd8a9

Observation 369c7027-b1ad-4fd9-98bb-0f23ec2edc62 · inbound

Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses cites this paper.

Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:15.092166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:44:15.092166Z digest=sha256:e176a0cc3c78c40f7e171779c98384f7e38501eb223813d27c08a6565be6fbaa

Observation 874e96cb-6304-4471-93d0-70d302bc59ba · inbound

LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers cites this paper.

LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:03:08.140918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T17:01:47.603706Z digest=sha256:95ade5cf397c93b1806cd6762f25901ef3350eb1d77f75193f3cc5054c56fab6

Observation fc689107-5c00-4232-8359-19832d7ab9e6 · inbound

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences cites this paper.

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:08:58.747674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T20:04:53.983505Z digest=sha256:79117a4d7beb7b6ca34b3f42e6b3977a3eca8a24eb1498828d1439adafd373b8

Observation f9173bd3-91a0-434b-99ee-354f45b8bb77 · inbound

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences cites this paper.

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:08:59.301022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T20:04:53.983505Z digest=sha256:53b50b8808413ae5dbddc1444cb148b62aa067260c49158c87ca69e79e84ba78

Observation b8eb6e69-c084-4c87-b914-610b0f4c24ba · inbound

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation cites this paper.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.686429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:3116a810bd254b23a81fe80b43cc98d085a266dbea19a183ca196eaa1a1825e7