Pith. sign in

Paper Citation Record · LEDGER

Exploration and Exploitation Errors Are Measurable for Language Model Agents

As of 2 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2604.13151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.13151 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:00:24.785343Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T06:01:28.311572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact9
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eddaaf24-f2a2-446d-af0a-49c46fcb1efe · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Exploration and Exploitation Errors Are Measurable for Language Model Agents gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.305466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:0e1189f609f3a06ba4a689f2b08a1b6089810ee42ab37ad7061aa1fd21ea0692

Observation 7dfcaba7-93a1-4d78-88e8-fd8e6ef58ddf · outbound

This paper cites Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:02:09.854882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:20ff65acc57fc12275ac189da3310bc7f1c23e571828c615007ebb07a522eadb

Observation 5427e764-595e-4a3c-8acb-547afa873d03 · outbound

This paper cites Anywherevla: Language-conditioned exploration and mobile manipulation.arXiv preprint arXiv:2509.21006.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Anywherevla: Language-conditioned exploration and mobile manipulation.arXiv preprint arXiv:2509.21006

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.275350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:f4fe9ddfd050347af781ae4cbeaebe4991efbc37e2d7daa89814b4d6cdbd880c

Observation f6a71493-9749-4afd-b47a-4bd7ccdef771 · outbound

This paper cites Should You Use Your Large Language Model to Explore or Exploit?.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Should You Use Your Large Language Model to Explore or Exploit?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:42.262884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:c5ffc74514d8394f5e2b1dc8f9101c3ff7d0f901aa23c9942b7bf60087ce7be8

Observation 1d10739c-1e8d-4d61-b8cc-aa51580c8936 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:15:58.398964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:baebafaf93b5b2887b7160ad8c82068a32d6520b3aa3c3eab31839acb6e61ee9

Observation 109a37ae-ff0a-4a39-a9ca-3e532bbdfab4 · outbound

This paper cites Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:21:02.341364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:04484dcc3c019eccda9daa112a8211c3e5438222a310c9c833fce39ad29be7fa

Observation b08a4978-6495-429b-bd9a-a600db9a573d · outbound

This paper cites AutoFlow: Automated Workflow Generation for Large Language Model Agents.

Exploration and Exploitation Errors Are Measurable for Language Model Agents AutoFlow: Automated Workflow Generation for Large Language Model Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.368384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:44ed1d2916522d210e66179ebe881119a0b1604c4ebc2afeb1ab0e8da13de326

Observation 342b5f46-0949-417f-b797-6b2d68b7ad42 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.382824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:2384ea0ab5f16734f6fbb52038c26e6aea93f316f184d6e1c1d44c4937c0b724

Observation 9f30f0fe-13a5-4269-9fce-7cffe6a444ce · outbound

This paper cites Expanding LLM agent boundaries with strategy-guided exploration.arXiv preprint arXiv:2603.02045.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Expanding LLM agent boundaries with strategy-guided exploration.arXiv preprint arXiv:2603.02045

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.296140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:312aee757637ae9d4d8de5334b6ab79248408a27912ae23d1364c962384072f0

Observation 5bb6699e-ae8a-44c3-8bfe-f21e5c7da09e · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.316967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:ed62be59e33de35fabc6b1c074a5d68d20228f8fe230eb9de4f44902d7d87894

Observation 623c3e4e-444e-419a-8ece-5adcc634f70c · outbound

This paper cites Under review.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Under review

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T09:46:13.716534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:9b00f0ed3539ac7367ce49fcda0a3baa9b2c337c692233faeb65231c26c79554

Observation d0bdd6f9-7ee4-4c65-92a9-ff63fa95865b · outbound

This paper cites action".

Exploration and Exploitation Errors Are Measurable for Language Model Agents action"

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T09:46:13.713882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:0325e28089e29b41783ce8552556e1acf0897a279cce5f5bd8d9ead6f0e1c485

Pith citing papers

Observation c56ea90b-128c-46fe-b281-7972fc5aade6 · inbound

Multi-Agent LLMs Fail to Explore Each Other cites this paper.

Multi-Agent LLMs Fail to Explore Each Other Exploration and Exploitation Errors Are Measurable for Language Model Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T06:01:28.311572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:01:28.311572Z digest=sha256:839d126aebe5cfb3cbaa1bca9ff8ff05a77f7ef1737c1047b8e030eddfa7f0db