Pith. sign in

Paper Citation Record · LEDGER

Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2502.15840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.15840 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:31:27.304146Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e917c6d9-ae85-4aba-8155-8e4637ba99a5 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.485174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:15617302501e8140f529d49cd873fa77d3021f809b1492cc9ec71dbb390df593

Observation b4ffdfc9-7646-4a7d-b84c-26b6d3f8d008 · inbound

PyVision: Agentic Vision with Dynamic Tooling cites this paper.

PyVision: Agentic Vision with Dynamic Tooling Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:27.304146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:27.304146Z digest=sha256:7f93aa923ddabd8811034bdb1bc4bff9ce357d02fc0af29b7a9cc88171ec1906

Observation 4877a63d-1dff-4ce2-b53b-1d5ed6dc5a8d · inbound

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format cites this paper.

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:10.799080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:10.799080Z digest=sha256:2b61df7ef6a2fa4543d59e5f610136cea444894cae57cf68394b0f18b18b1875

Observation 32327f00-ae59-4aab-aae6-910381a70a9f · inbound

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare cites this paper.

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:55.517898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:55.517898Z digest=sha256:9f2348175344bbb443b2bffa7334c37c6d81e4f695c17b16a395b5dabbbfbd88

Observation c124a8f3-92c1-4ae3-b4e8-b0464d687c5b · inbound

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions cites this paper.

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T21:20:35.290693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:20:35.290693Z digest=sha256:df96e67cbae2c04db1e6836b50f00e1dfcd75a5179bd1eba1016ce78bd1f0dfe

Observation 7ee1af49-dc9c-4beb-9424-7cf3aa09b1e8 · inbound

LLM-SAA: LLM-persona Generated Distributions for Decision-making cites this paper.

LLM-SAA: LLM-persona Generated Distributions for Decision-making Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:46.074199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:46.074199Z digest=sha256:2cc137cecdf1d61eb07af980e6c3959a409e4256423497aa8083aae96e0b0663

Observation 9329eed0-28e2-4aaf-9c6b-9b73b2a1c756 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:57:24.662243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:839850afae3e5a822fb324103267d4f1ceff8b1e1214e5f01c90b7907aebea19

Observation 80edd497-0d36-49de-9428-7ab2ca384f26 · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:41.129911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:b173154bf3c1aa46fa1dc04f2f844f0db6505c8eb15152746f781e8643a8e237

Observation b5035e18-9b49-4d02-810c-bb368d148b38 · inbound

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents cites this paper.

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T15:09:29.436835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:09:29.436835Z digest=sha256:c4300353acc4dc57f81d62b4622d05d71cef1fce107a58696712ca6b3a46f1f6

Observation 70d12f98-1c10-4ba0-8fd5-7f5d50dd68f3 · inbound

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models cites this paper.

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:03:20.367050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T22:02:12.555644Z digest=sha256:ddb231b8848630fda620213f04eeb39a4c11c2e91db970fb3ac3a0d68db78177

Observation ed5787e0-d2e7-485d-893e-4613887c86c6 · inbound

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V cites this paper.

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:01.844569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:59:21.887765Z digest=sha256:0d1eb16381b478557752bf0b636465fc640d6fd4d8fdcde773ecfa307b381dbb

Observation 44113a13-f4f7-4fcf-af13-fdf52409b9ab · inbound

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration cites this paper.

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:53:08.341871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T18:49:15.486503Z digest=sha256:3ae1a625f6c642e8e0eae6b89d480b6f0ff9aa4449ecaac9c19f0603fe55987f

Observation 5e7f2a29-2cc6-41d3-9f64-8cf0c5e25773 · inbound

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas cites this paper.

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.701078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:27:02.965559Z digest=sha256:936cf2e92095c6a79f593d7102aa33288fbfa5b67f071a914c7739badce5cefd

Observation 094ade2b-8ef5-4766-b959-46212130cab2 · inbound

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas cites this paper.

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T19:45:34.692776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:45:34.692776Z digest=sha256:aa323c8d3659038b3190546b8280b26ce60125e46e3ce5d39ceb9139b9f54c83

Observation fcfb4317-89f8-4629-8821-98b268dfbbcf · inbound

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations cites this paper.

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:07.696827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T03:14:01.097146Z digest=sha256:24fa1c2210c2c3f75a96ddb580c094454d5922e8fc00e6d67eed9107a9ebba88

Observation 7f937a47-e7c0-46c2-9d87-c97643217658 · inbound

CL-bench Life: Can Language Models Learn from Real-Life Context? cites this paper.

CL-bench Life: Can Language Models Learn from Real-Life Context? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:27.025477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T09:42:33.635866Z digest=sha256:9fdd7f1257f7bd37ca7124b3625ba61a857625ceeea3f9b8268e56a1effc07ce

Observation f5b9b71b-94e5-4585-bf4d-7f13bc5405a0 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:47.873226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:762d83ca10663989b873df5ff2a32245be75951aca8e2683a4175e5b53836e0d

Observation cf28f79f-6def-4060-99f5-446639afa09f · inbound

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces cites this paper.

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:28:19.054534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:25:48.728238Z digest=sha256:545d737eee68f95588dac434244b5a6db6a6f37cf9de21bb302343579ff99654

Observation 5ad775ec-7465-4ab4-8e7b-85ec66537527 · inbound

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use cites this paper.

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:16:00.246835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:03:36.851403Z digest=sha256:7ce8bfc16d1675d9fd9499712137610a2790a5b08d5dfde05e99599960f31808

Observation 33b4960e-f1b0-44ee-9365-1bb344d6df77 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:38:58.693981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:85946eeb4b0e6cb3dc1b2e7d9322fa58175818f4e21b9c228149affea144f368

Observation ee76460d-f7df-49d9-a7a4-fdc5bcfc724e · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:16.111772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:16.111772Z digest=sha256:c4144b0e500bf5261359095ffc84361e7304bb95188be9d2e5e876a14f366eab

Observation a76bf20e-fa30-4e2f-83b3-ea5bd853df8e · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.490566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:648a66d20b46ed59b584ec7f12b46f8287e93f29ef6e49bfe472512d9db21922

Observation 0746d08a-99fe-4367-bc8d-99bb79d794f7 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.915566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:8cd107bd04b1078fe817aa6a0e6f3fea48563bcf8dc4e5372bf6415d95faba23

Observation 51fbd126-8dbb-4a01-ab59-49c10fb4055d · inbound

Graph-Enhanced Large Language Models for Spatial Search cites this paper.

Graph-Enhanced Large Language Models for Spatial Search Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:51.950322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T06:46:29.757188Z digest=sha256:6e12ef4b273228bfaeecafe1511881327e16a7ab5bb14e4dfb35fb6b7d1331b5

Observation 3526cecc-8a5a-425d-b370-31f8a8788bdf · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.678930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:3a5d4c68441d8272ca1cacc8f1becf31e60e4c79d6e20dfd15bccd776e190547

Observation d05376fd-dcb6-4499-8295-d0052e97d8be · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:4843c90972cc2ae2be8667be1e4bd5fabfbd565369553e22ce4246104b4cb3b9

Observation 1e3dff5a-2ea5-48cd-8372-785d5005be33 · inbound

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle cites this paper.

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:05.920354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:05.920354Z digest=sha256:924d4e879c4804e90fca74d1e33f2a77e7d9113f7db793b59f8e417d65967a89

Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · inbound

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios cites this paper.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.686510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.686510Z digest=sha256:1842b897c32c76489e7b4a8e00fee7403beb2695fc2248af17a35f1750f51528

Observation 854d0ebe-bbf0-4d65-87de-47cbcbb6674e · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.191363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.191363Z digest=sha256:e761de4eaeaceb6902cc77df4ea90dfcc92ec76c887fb0d072f7570e147cc498