Pith. sign in

Paper Citation Record · LEDGER

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2504.18838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18838 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:32.780612Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c69955e5-9d2e-4952-96e6-1cd2ce2b0990 · inbound

Continuous Monitoring of Large-Scale Generative AI via Deterministic Knowledge Graph Structures cites this paper.

Continuous Monitoring of Large-Scale Generative AI via Deterministic Knowledge Graph Structures Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:32.780612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:32.780612Z digest=sha256:d7985cc754645de024792271078146c86968b85cbac0a80ed7d2a9b829719a18

Observation 13655fd8-9d32-46dc-b1ef-2da6a2738f79 · inbound

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users cites this paper.

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:23:39.925937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:22:42.431997Z digest=sha256:3b4005eb7bf80af409770cb8632188dbecba2dbb01a28232d60d3508f259280f

Observation 02aae78f-22e2-48c5-b8bf-d9153f9c2d30 · inbound

Security in LLM-as-a-Judge: A Comprehensive SoK cites this paper.

Security in LLM-as-a-Judge: A Comprehensive SoK Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:13:28.899936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T00:12:59.886619Z digest=sha256:c1a65b7bb0ee9e78209a883afedab0a1fbfec01cbcf1af2f1c43f4e9831f125f

Observation 96b15329-22c6-421d-a8ec-8d5a5c93106b · inbound

Evaluating LLMs on Large-Scale Graph Property Estimation via Random Walks cites this paper.

Evaluating LLMs on Large-Scale Graph Property Estimation via Random Walks Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:13:57.707885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T15:09:03.417040Z digest=sha256:83e5c5049f3e0ea7c1f744ea805d5d8c3443352781d08f208c3fab4296b4c79f

Observation c9163e06-9be5-4916-966f-63e6853f40a1 · inbound

Reasoning emerges from constrained inference manifolds in large language models cites this paper.

Reasoning emerges from constrained inference manifolds in large language models Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:01:15.372846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:58:30.713839Z digest=sha256:ae505b6e0f0c90fff495b22ba5d1f6b400d70faa9b2a994293f3f4f8932f7a31

Observation dd061787-81fe-4c96-88b4-32812e5ef625 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.358638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:54:13.146373Z digest=sha256:32c001002f85dc9304607974daa8d2e2568b4e7d930e1d6e08f366be024caa9c

Observation 0d5bc779-74eb-48b0-b695-1e0419257460 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.047592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T23:34:43.785312Z digest=sha256:2fde4c14d91aa1d44034ed8750dcd85cef975b2b0a604617a95b0f4cdeae3588

Observation 7975c132-25c7-49f6-b576-e7e34c6edcc4 · inbound

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents cites this paper.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:20:19.660105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:17:36.487954Z digest=sha256:e117d83f9b488eca7da909b52bc7d28e413bd7835505ca0be2aea24ad2dc8604

Observation 9b742e26-6c99-428d-a1ca-37d918b187c9 · inbound

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents cites this paper.

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.086828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T16:13:43.527362Z digest=sha256:db353a6d2e0acfb10c2836bc65e811c2285cb9c0110853f44a99263b27d4ced4

Observation 8bb11e71-cea9-4d42-bbe7-f5f9281947c9 · inbound

Discriminatory Compliance: How LLMs Answer Queries from Protected Groups cites this paper.

Discriminatory Compliance: How LLMs Answer Queries from Protected Groups Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-26T12:59:29.207594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T12:58:09.227648Z digest=sha256:92e9953dd275c6eaf3eccc356c32df7350e0993c92b636dd8ace456edb9f8bc2