Pith. sign in

Paper Citation Record · LEDGER

Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2502.15840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.15840 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.831296Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e917c6d9-ae85-4aba-8155-8e4637ba99a5 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.485174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:044e9ce587b35c761f7973cb2d6e86ecf51daa7e4108ae36b80656878a6d5e45

Observation b4ffdfc9-7646-4a7d-b84c-26b6d3f8d008 · inbound

PyVision: Agentic Vision with Dynamic Tooling cites this paper.

PyVision: Agentic Vision with Dynamic Tooling Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:31:27.304146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:31:27.304146Z digest=sha256:78b2d1841cb0f68f71c2b4983a1bf12045f716e8e594d8bab7686e42b3ae35f6

Observation 4877a63d-1dff-4ce2-b53b-1d5ed6dc5a8d · inbound

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format cites this paper.

BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:10.799080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:10.799080Z digest=sha256:b957b066201088676d2ad78c6307711e8b0d2aff25a1a588616d341c330ab326

Observation 32327f00-ae59-4aab-aae6-910381a70a9f · inbound

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare cites this paper.

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:55.517898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:55.517898Z digest=sha256:b9d5480e203747d2869e4634dcd9628dbfecabb4a4619b041983e6130e23c390

Observation c124a8f3-92c1-4ae3-b4e8-b0464d687c5b · inbound

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions cites this paper.

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T21:20:35.290693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:20:35.290693Z digest=sha256:c9514ef7c7358803b364f4d9f1bea86f4d654ed11a5b4943c30820a557c0bd84

Observation 7ee1af49-dc9c-4beb-9424-7cf3aa09b1e8 · inbound

LLM-SAA: LLM-persona Generated Distributions for Decision-making cites this paper.

LLM-SAA: LLM-persona Generated Distributions for Decision-making Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:46.074199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:46.074199Z digest=sha256:98c130caaaf710340ef8248e350d95f25f2e34d3d1218b377525c436997db1fb

Observation 9329eed0-28e2-4aaf-9c6b-9b73b2a1c756 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:57:24.662243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:5b2bf4c55367838d959bbe5b275323ca6c36da9bdeac492f327876d202b2e5a3

Observation 80edd497-0d36-49de-9428-7ab2ca384f26 · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:41.129911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:072c43ff51a326aaddebf055988c8a5b4defe89b15d39c55daf5c5661f73e8b3

Observation b5035e18-9b49-4d02-810c-bb368d148b38 · inbound

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents cites this paper.

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T15:09:29.436835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:09:29.436835Z digest=sha256:6928be07588cffb39c51b48481d2c4f076efa45842b8b5ff90c135c72582db19

Observation 70d12f98-1c10-4ba0-8fd5-7f5d50dd68f3 · inbound

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models cites this paper.

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:03:20.367050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T22:02:12.555644Z digest=sha256:57e1a8575b4c3196174b090ead78dca6f62e254c190a00a24f5ea97e63b54d6f

Observation ed5787e0-d2e7-485d-893e-4613887c86c6 · inbound

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V cites this paper.

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:01.844569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T17:59:21.887765Z digest=sha256:5a3e3e4099c5f56370681f88543d31f2c26e72fa27f469339f9ae12aa7063218

Observation 44113a13-f4f7-4fcf-af13-fdf52409b9ab · inbound

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration cites this paper.

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:53:08.341871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T18:49:15.486503Z digest=sha256:2ec049c1b916bd0f917ccbf62944e9bb637ccf5baa3239837914e5b1d45f7cc9

Observation 5e7f2a29-2cc6-41d3-9f64-8cf0c5e25773 · inbound

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas cites this paper.

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.701078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T09:27:02.965559Z digest=sha256:9193db681b522d42e1f9e334ec373cc2a8ebf59ea49a808e30a82cec5ca8c333

Observation 094ade2b-8ef5-4766-b959-46212130cab2 · inbound

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas cites this paper.

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T19:45:34.692776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:45:34.692776Z digest=sha256:7c6e7df8a6b9017d6cfc4d9a7465bf1671eb3381acbc6a83ffd9836a03b338ae

Observation fcfb4317-89f8-4629-8821-98b268dfbbcf · inbound

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations cites this paper.

Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:07.696827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T03:14:01.097146Z digest=sha256:d97c1a2e376837c6e8773b3f445b73f358913ec732a6f4a62074495b6fc044fb

Observation 7f937a47-e7c0-46c2-9d87-c97643217658 · inbound

CL-bench Life: Can Language Models Learn from Real-Life Context? cites this paper.

CL-bench Life: Can Language Models Learn from Real-Life Context? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:27.025477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T09:42:33.635866Z digest=sha256:ef5fe58f73ac283bf3a802fde48335d5c40022f1cbf508a659e6860a566888fc

Observation f5b9b71b-94e5-4585-bf4d-7f13bc5405a0 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:47.873226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:01a2e1bcc6f0f3873364ab155ff0337f2cbb512258dcc697b975e0d83206ff6f

Observation cf28f79f-6def-4060-99f5-446639afa09f · inbound

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces cites this paper.

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:28:19.054534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T13:25:48.728238Z digest=sha256:6defff328f580c6f508466431998fa1edde40aa46366997d4f462ea94ac61102

Observation 5ad775ec-7465-4ab4-8e7b-85ec66537527 · inbound

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use cites this paper.

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:16:00.246835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T23:03:36.851403Z digest=sha256:a2da24266a174a8836caf9e569e6c13fa43ebbe87c27d22f9e8c3ec276ffaa11

Observation 33b4960e-f1b0-44ee-9365-1bb344d6df77 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:38:58.693981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:c7159f5ddb5f9b0dc6dc3caf4160b278c9b4492dd4d9616a19b8cd821d6da757

Observation ee76460d-f7df-49d9-a7a4-fdc5bcfc724e · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:16.111772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:16.111772Z digest=sha256:ab63d19982115828a8d16e55383e528fe9d48a04943a48d3042aec1c75e109c9

Observation a76bf20e-fa30-4e2f-83b3-ea5bd853df8e · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.490566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:e6de21bfb574af4ea2f8e7a6ad5c4a455e8396161a33f528fe3c8d47784391d9

Observation 0746d08a-99fe-4367-bc8d-99bb79d794f7 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.915566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:46e59b6c405945415ce03749639be4ab5117d342f026752054d814d8357ba2ff

Observation 51fbd126-8dbb-4a01-ab59-49c10fb4055d · inbound

Graph-Enhanced Large Language Models for Spatial Search cites this paper.

Graph-Enhanced Large Language Models for Spatial Search Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:51.950322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T06:46:29.757188Z digest=sha256:fe8bf2a1302cf829a4f54fa33dfdd44ffa30e9ee816ebe1200a900353d417048

Observation 3526cecc-8a5a-425d-b370-31f8a8788bdf · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.678930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:8f31d71ca5ba938b63582b543ad8f342eae5137919fc20ad770926d49b0577b5

Observation d05376fd-dcb6-4499-8295-d0052e97d8be · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:e69f296970b0fd0caf7be6b8f6b676ba6bc74dd4fa07951a29fe3a00e89a3b85

Observation 1e3dff5a-2ea5-48cd-8372-785d5005be33 · inbound

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle cites this paper.

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:05.920354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:05.920354Z digest=sha256:22bd107c70a4663dcf56510463498dda962bea7fbc7095a5fdfe08c99d599335

Observation f92f1d4b-46d7-4c3f-81e9-094a084376a9 · inbound

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios cites this paper.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.686510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.686510Z digest=sha256:75d50a90125d15e14d4d172ed38b4dbd0aa46b8d6e6a56bbd8ded568724b0e2b

Observation 854d0ebe-bbf0-4d65-87de-47cbcbb6674e · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.191363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.191363Z digest=sha256:d735fbaa7abed039d26ea5268ffa7c1815c6de41838b3804a6e2f6daf3cba1e2

Observation 23a200b6-5d54-4401-bf6e-ecac8ee10a00 · inbound

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents cites this paper.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.831296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.831296Z digest=sha256:96ccc49dbf4d8afb0ece6688a72eb69edb8f61db28d14e23cc645bf9f7250ec9