Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Foundation Models with Language-Model-as-an-Examiner

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2306.04181.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.04181 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:18:35.230093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:56.739789Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 709e196c-4968-4e26-adee-d265c8b598c2 · inbound

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate cites this paper.

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:03:18.797450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T13:03:18.765496Z digest=sha256:423c41e63af5c51354654dc8a2ab3feb66b234f2b34b0093acf3de58d8323ccc

Observation 32ecc13a-87c7-47d8-96bf-10d9804d0665 · inbound

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding cites this paper.

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:22:10.630306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T20:22:10.482509Z digest=sha256:ced92415a7867109503848dde503d01799f4e830462f6d574e97dc7fe70c6b9a

Observation 951e27cd-8ec9-4c66-b6be-7de0ba20a22f · inbound

AgentReview: Exploring Peer Review Dynamics with LLM Agents cites this paper.

AgentReview: Exploring Peer Review Dynamics with LLM Agents Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:38:37.146359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:38:28.005028Z digest=sha256:a94a63f4bed3199279e81721cbd58339fd0c6693611d15d80e6321fc50585237

Observation b89aacec-13e8-466a-acfa-de56a665e9d2 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.809748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:f113db7863bc83e69f3a8792461928ccf069ef0442857c5ed44e0c67c3fc6120

Observation 1bb3e0e2-d9dd-45df-adac-46ac38cdbccd · inbound

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence cites this paper.

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:59:38.269088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:57:32.072709Z digest=sha256:92526ab54abc4e74b2086c16a8f8a39291f0f901b9463cddcb82df649eb59f33

Observation 2ea56a0d-fd6a-41df-857d-faf580d5c496 · inbound

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence cites this paper.

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:35.230093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:35.230093Z digest=sha256:b676592cb19283527bf7bd673207898fdead92ef9652f435f3153b20a9ae1831

Observation 28a2f59d-009e-476d-9a58-83cfb2b69725 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 47

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T15:29:56.741596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:5c41669e8611e2887f6c67267680e05839318b40693153e8fe7551e1042a0d7a