Pith. sign in

Paper Citation Record · LEDGER

Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2309.04369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.04369 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:58:46.629457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:58:14.016563Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 53f7a55a-5686-4c2b-97ab-c2bf45d9e5a9 · inbound

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios cites this paper.

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:58:46.629457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:58:46.629457Z digest=sha256:c76c5d1ca39c3089e655382b6d31388067dd85eebc7181293fcf5c92cae8afc5

Observation baaf731e-2d06-4501-9f47-67da1aceedf5 · inbound

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data cites this paper.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.648196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.648196Z digest=sha256:89595274a386324c3d6b44b2c6c78f7ad826e3556cabe7036b65f8dc2b6f5abb

Observation 436b2797-8b3f-4072-b92b-71a493b4faf2 · inbound

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps cites this paper.

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:29:00.614737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:29:00.614737Z digest=sha256:61fe4d0f58cf5358dffb0bd75b9c86100345279bc75a812b2f4db0141e2513c3

Observation fa83cae4-9498-4641-8f08-e74ac876e1c3 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.018329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:dc81284cfe6a92996a60d6bcaa048bbcf4ed6dd9b308cb11a7e8f7db292b310a

Observation 347972dc-3143-4ecf-9f33-67af60199b95 · inbound

The Evaluation Game: Beyond Static LLM Benchmarking cites this paper.

The Evaluation Game: Beyond Static LLM Benchmarking Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:58:07.772438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T07:56:01.638262Z digest=sha256:4c6a82325efb21654a0d74bc4949c7733c39881cc9185e1f2f36cc297af2c5a7