Pith. sign in

Paper Citation Record · LEDGER

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2402.15754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.15754 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:47.053105Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11e491a1-2bd6-455c-abe7-d4b56addf389 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.341931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:13189edd201cfc7be13f75667a3ebb51bffdf85f04e3a6144974fed8da6a257e

Observation 079f109e-1ebb-443b-b672-28b8b4f858e0 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.053105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.053105Z digest=sha256:589f178fe39b15c1538e223e2dafddebd8e2ae8c518b1b3c8920dba0115281ad

Observation 38b7b8a5-fa1b-472f-83df-95981658ea5c · inbound

Balanced Hyperbolic Embeddings Are Natural Out-of-Distribution Detectors cites this paper.

Balanced Hyperbolic Embeddings Are Natural Out-of-Distribution Detectors HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:07.186704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:07.186704Z digest=sha256:75cd14b39ff5392879219bc12020081fe7fe1dde62467ea0706acd2bb6daed6c

Observation 369c7027-b1ad-4fd9-98bb-0f23ec2edc62 · inbound

Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses cites this paper.

Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:15.092166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:44:15.092166Z digest=sha256:5b60ec7e13765c628c477994ae3ace20e40dc0ca07ea21820178536a82dfd38f

Observation 874e96cb-6304-4471-93d0-70d302bc59ba · inbound

LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers cites this paper.

LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:03:08.140918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T17:01:47.603706Z digest=sha256:d61eb73abdd3c9532a663ff76b5fb24070486dfdda4e2b9a48d7780ed56c8fcf

Observation fc689107-5c00-4232-8359-19832d7ab9e6 · inbound

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences cites this paper.

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:08:58.747674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T20:04:53.983505Z digest=sha256:c50767fc690aec7bde3c16fc0a222600ff40dddce9f7a0b004aab74bab29c3ac

Observation f9173bd3-91a0-434b-99ee-354f45b8bb77 · inbound

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences cites this paper.

Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:08:59.301022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T20:04:53.983505Z digest=sha256:6eb5785d7eaaf45452844006ac4667d1b51e10074a615f63efa5e1b0f055676c

Observation b8eb6e69-c084-4c87-b914-610b0f4c24ba · inbound

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation cites this paper.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.686429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:b406e14c161144eddb38a0bc154fe365c5bc9e233694d47fd300c9d66579bd69