Pith. sign in

Paper Citation Record · LEDGER

DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2309.17167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17167 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:14:48.935323Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:18:22.614907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 80abcfa1-0db8-4723-ad31-ad77ce6bf351 · inbound

AI Alignment: A Comprehensive Survey cites this paper.

AI Alignment: A Comprehensive Survey DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:28:49.009396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T14:28:48.987140Z digest=sha256:4b83de77fca02c5720964b91fd7d4a916da1739c52d3e57e903faea38be34f3c

Observation d096e8aa-526c-456e-9285-3ca85a935ca0 · inbound

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey cites this paper.

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:16:41.701975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T23:16:41.679855Z digest=sha256:ca42cf8eed2459e39edb37a47e0b82179184807ed952ad9e048cd6a5f38f7ecf

Observation 10e396b0-2f39-4d67-a49a-4e228e7af54b · inbound

Neuro-Symbolic Data Generation for Math Reasoning cites this paper.

Neuro-Symbolic Data Generation for Math Reasoning DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:17:13.124865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:17:13.124865Z digest=sha256:8ec1b1539972bde3c275f484fb481fa54a367f24e24729d5c57e555782330263

Observation 7b10c5b1-eb13-437b-a2d3-e8c4a2892122 · inbound

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios cites this paper.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.762978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.762978Z digest=sha256:c5bf5f8060851ddfb64afcca6c0dd2af40d1abb1f67fb4c946445f7a01a4a2cf

Observation c4b9627b-f19b-4e02-a1a0-8274e95e2d9f · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.341691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.341691Z digest=sha256:5612d6218306c5a8cd59056146424ca26bb68aca0daec61e5e41be1b68a10573

Observation c9a55490-a49a-4bef-95f3-e7797aed594e · inbound

Dynamic Skill Adaptation for Large Language Models cites this paper.

Dynamic Skill Adaptation for Large Language Models DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T00:46:46.343201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:46:46.343201Z digest=sha256:0a6357df06641eb85ae4d1930d6ad4cd64ff6bf94be6c87922232480ed86d543

Observation 4d3f199f-97ef-4ba1-aebd-2a112b542616 · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.504623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.504623Z digest=sha256:f83af96a731c8cc214ba54bef84b3e9d25754139c28a55ccdf80d954aa88fc3d

Observation ffb450d2-df90-4b18-9933-f04b013d0295 · inbound

Automated Capability Discovery via Foundation Model Self-Exploration cites this paper.

Automated Capability Discovery via Foundation Model Self-Exploration DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:55.420540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:55.420540Z digest=sha256:1472b8a041a6a134d7853b619cf2dd3dd0daf6111fdc80442185217d023fc7bf

Observation 1cdeee7e-923d-4e0c-a08e-82d9643bf942 · inbound

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models cites this paper.

THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:14:48.935323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:14:48.935323Z digest=sha256:18f2ea0057970246e51ccdf68f97c4a71f7234aaf2abfa339ea0033eb514ddcf

Observation bc37ec53-3f00-410b-9698-170f711e3da8 · inbound

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks cites this paper.

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:31.461050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:31.461050Z digest=sha256:a7b19ad0a2c6094d76856a2bf1cb894d574ba6934775d5d33e615d310bd74050

Observation bcc4f788-8ab4-443c-b4bf-217c5acd20de · inbound

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation cites this paper.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.859281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.859281Z digest=sha256:c0ea98422fce2969fab5544a85fadd0fe03870dc6d466eccadd401bfa3e9ed9c

Observation 733e74a4-d303-4d0a-9215-9bd1fdd42f1d · inbound

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era cites this paper.

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:12.862651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:12.862651Z digest=sha256:0b42ed0750cb40cd322aa24e123ac158c8a5dae01ba771395264b9ec3de66ceb

Observation fd6e6606-853a-45cf-9a72-7e3e07d768a2 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.575749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:3e732b4428ad1d744feb81c26024cad2f997c16689d29d918d4a4808a46ebf84

Observation d8b1cad9-4724-4f5e-a61c-630045c9aaf5 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.745491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:953af97a135e92f66a9ae0956db3568602d2218d30890a78063a24c448e191a6

Observation 72ba854a-499a-49dd-8308-cf390beb281c · inbound

From Text to DSL: Evaluating Grammar-Based Model Generation Using Open LLMs cites this paper.

From Text to DSL: Evaluating Grammar-Based Model Generation Using Open LLMs DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T16:43:34.739101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T16:40:08.788179Z digest=sha256:ed6472b1d62fab7d59c1ee1cdda1d1858e0fdd7037739fa116bcfdc352c5bd48

Observation d15d5841-55e1-4ab4-aac9-a2a998bfcdcb · inbound

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations cites this paper.

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:26:09.983958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T06:25:06.024542Z digest=sha256:946a368c2369c6e3e2dfdb1a5e1faafc95e72c7609e0ee0485c2f7bf5c0e0750

Observation 1e0e4761-68dd-421c-af32-243da1e6ddd5 · inbound

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction cites this paper.

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T14:18:22.616286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-03T14:08:50.462880Z digest=sha256:7441784eb2c19f0c8f4092b653d61313e8a09464de83177741dd78b3975855f1

Observation 8a1ffbb2-14d1-4c9c-96a8-296db5c6fda7 · inbound

Ask-E: An Environment for Calibrated Question Generation cites this paper.

Ask-E: An Environment for Calibrated Question Generation DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T18:11:50.726572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:11:50.726572Z digest=sha256:2643488fee29de8fa8e9c491254051836944370c5551b0aa1a626fb23d504b9a

Observation 91bf9fc0-82d3-427b-aa8d-ebad0bdf38b9 · inbound

V-FiLLM: Verified Financial LLM Reasoning Benchmark cites this paper.

V-FiLLM: Verified Financial LLM Reasoning Benchmark DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:14.569773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:27:14.569773Z digest=sha256:1d5812ab4c9201d859922b3cca94089256b1cfd1aedebffaa1ef4608997bc38f