Pith. sign in

Paper Citation Record · LEDGER

FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2404.06003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.06003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:40:13.907632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:10:41.139960Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3c7e143-51de-4d68-9ec0-e218f84b6280 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.142309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:f10b8502a1c0dc5dc7d02434b5defc20634d4efdc8e997a7a1516d389bf92fc0

Observation 7dca7866-e9b6-499c-960e-481489b1c060 · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.907632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.907632Z digest=sha256:c553cf5445c0221d7b9cc4a9846daa6caf93a9547be777c4f71cd7bfd62601e2

Observation aac12416-9c89-4f07-a085-38a7e4404567 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.109652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.109652Z digest=sha256:3217051b64c67f903fa4e0ac5e4eb75eeb398de98bb6e6f3aec9ba576396acc9

Observation 0c2c9326-3176-47d3-bd94-fa9397f04101 · inbound

Behavioral Fingerprinting of Large Language Models cites this paper.

Behavioral Fingerprinting of Large Language Models FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:04:08.940093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:04:08.940093Z digest=sha256:261098d9d4673c7314988aa04f3c3bed509d9c59fd12072bbc4cc04734553773