Pith. sign in

Paper Citation Record · LEDGER

DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.17167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17167 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.504623Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:18:22.614907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 80abcfa1-0db8-4723-ad31-ad77ce6bf351 · inbound

AI Alignment: A Comprehensive Survey cites this paper.

AI Alignment: A Comprehensive Survey DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:28:49.009396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T14:28:48.987140Z digest=sha256:5f82a10bb321073e7c9db1706790eefad2cd8ed4d2b93cdf95a419642fbb0431

Observation d096e8aa-526c-456e-9285-3ca85a935ca0 · inbound

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey cites this paper.

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:16:41.701975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:16:41.679855Z digest=sha256:d0ea8b391a3dff940cd127bd8445d243ad621cc484de0b822775020e05f5df15

Observation 4d3f199f-97ef-4ba1-aebd-2a112b542616 · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.504623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.504623Z digest=sha256:e01e4a567fa924df1de4b6395255c97173d2f96c464472dadd9230f4b1ce6376

Observation ffb450d2-df90-4b18-9933-f04b013d0295 · inbound

Automated Capability Discovery via Foundation Model Self-Exploration cites this paper.

Automated Capability Discovery via Foundation Model Self-Exploration DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:55.420540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:55.420540Z digest=sha256:3b06b258baa7cb8366eb9e579b26c20ae5da2941538ca2ac4b1ad499f1fc80a9

Observation bcc4f788-8ab4-443c-b4bf-217c5acd20de · inbound

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation cites this paper.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.859281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.859281Z digest=sha256:2326ff508b18b282d129f4e4ff283a7fbfaac982eedfdfe3f3fd64ec5f6edc1c

Observation 733e74a4-d303-4d0a-9215-9bd1fdd42f1d · inbound

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era cites this paper.

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:12.862651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:12.862651Z digest=sha256:9445d97a47c6102bdc1b5ea8f9b6add4e4fdd1e1595fb59c73560a799dfff393

Observation fd6e6606-853a-45cf-9a72-7e3e07d768a2 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.575749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:307e0926284677a373c4ddd7d3fb81370733ba50a60d70972af2228a1115be55

Observation d8b1cad9-4724-4f5e-a61c-630045c9aaf5 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.745491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:89a2d39207cc153035af10ef2fd03382e4010a9961cd3764f8b2d10e4c0db4e6

Observation 72ba854a-499a-49dd-8308-cf390beb281c · inbound

From Text to DSL: Evaluating Grammar-Based Model Generation Using Open LLMs cites this paper.

From Text to DSL: Evaluating Grammar-Based Model Generation Using Open LLMs DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T16:43:34.739101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T16:40:08.788179Z digest=sha256:923fd295f2dd75fc349905a274cdc5d055435df1509a8d79acd7deae359ad54e

Observation d15d5841-55e1-4ab4-aac9-a2a998bfcdcb · inbound

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations cites this paper.

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:26:09.983958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T06:25:06.024542Z digest=sha256:7ff85b4e90afa654fc97bba6d1a91000682663122a0f93c4a54e149057681c33

Observation 1e0e4761-68dd-421c-af32-243da1e6ddd5 · inbound

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction cites this paper.

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T14:18:22.616286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:08:50.462880Z digest=sha256:154fd5c03c24fa5dc75ffd6bf8f08dd799bc7bb6991ad50c5087a63e0c9264e7