Pith. sign in

Paper Citation Record · LEDGER

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

As of 4 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 12 inbound Pith citation observations for arXiv:2601.11044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.11044 v4

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T14:11:14.123806Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:13:41.885703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T13:13:26.970130Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact10
  • verified fuzzy4
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a028209-3bed-4162-9de5-2ce15bce74e3 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:12:58.683087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:3796abf31f6445fbe2ac625dbd6511e36ccaa4eb65e6a4fda507cb51f29063b5

Observation bde73516-2ff7-4988-a70d-eb795dfe8342 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:12:58.699728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:0c95621a44026e1adf8c5053734e232e0449e131e72f8f75398f47f2b7bbb48b

Observation 17782a1b-4dbf-40e0-8c4d-0020644b9e10 · outbound

This paper cites Innovatorbench: Evaluating agents’ ability to conduct innovative LLM research.arXiv preprint.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Innovatorbench: Evaluating agents’ ability to conduct innovative LLM research.arXiv preprint

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.690621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:7f3cce0e898a83fcb327907a208570dcc1561dda0ec706543249b61b3e6045dd

Observation c6af2593-b5b5-41ce-a215-477a0847e9f7 · outbound

This paper cites Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.686491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:1b7571d68c63c118043ba2882a7a24df28855acb247b1359f7d4092ab7350df1

Observation e9af553f-42d7-4bfa-8a16-ab96b077dbb5 · outbound

This paper cites an unresolved cited work.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:12:59.034084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:35aa96edc093e04bfbab9035d34cc1ba9bccd75ef597bf857aae28d13652eca3

Observation cf03ea51-78bc-4b4d-bf65-bcbed040aa26 · outbound

This paper cites Limi: Less is more for agency.arXiv preprint.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Limi: Less is more for agency.arXiv preprint

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.693840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:8c926b5d1607c13ebeb5486d23258ffec1143e1c8761090533a38abea92b5ff0

Observation 302e4c4a-b98f-41d0-b2c2-a4460ff1fa20 · outbound

This paper cites Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.696860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:2c3a2e338eb678ba35e529d327647da02f607b14e855a1e1581a65ba40f6c47d

Observation 9b458809-d36a-4fc3-a86d-fd3ac982ae18 · outbound

This paper cites Limopro: Reasoning refinement for efficient and effective test-time scaling.arXiv preprint arXiv:2505.19187.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Limopro: Reasoning refinement for efficient and effective test-time scaling.arXiv preprint arXiv:2505.19187

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.703590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:413912465b8cf78ad584c04844eb9469ea5addd3e32581a3b4c49654404d53b3

Observation 3fa6f6d2-ea48-439d-8e4f-45d2f14cfb0f · outbound

This paper cites an unresolved cited work.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Unresolved cited work

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.679925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:e74cb42431ff313273ecabfca008132479cae7b6f7c374675e05db050cc40714

Observation 208889a2-d2bb-4ee8-a661-03b1e730bcc3 · outbound

This paper cites ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.706797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:3c996502d4c1eaddff5bba034dca31a5955a4afced64b9ee2a786a8ccd656986

Observation 41c49cff-a209-4a65-b07d-a5b9fc47a8fb · outbound

This paper cites Qwen3 Technical Report.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Qwen3 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:12:58.709936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:c1ddc3406168776f519b97439bd5bab5cde9ed15069f0a13fbce6068d02c2bcb

Observation d9dad18d-5173-44d4-b58a-921952485b18 · outbound

This paper cites an unresolved cited work.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:12:59.031447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:68125b068a2068b759346b797313af3c698cbbd7f23aa9c66d67cee2ac4f3e6a

Observation aa0157a1-5141-4b7c-86e8-a50c930ac330 · outbound

This paper cites an unresolved cited work.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:12:59.047868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:f861b4544ae95ada7137726261d3a0dfaa5828d9a74923058c4c43460cf0c67d

Observation 8f7bd7f3-2c3f-4277-8917-1ed1dab5203b · outbound

This paper cites The ”Executors”:There is a striking divergence in how models orient themselves.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts The ”Executors”:There is a striking divergence in how models orient themselves

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:12:59.051600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:36e723700ef696dc8d1377c84334da1d5110bc88a80f4658d002b0800a958efb

Observation 8ae5fdd4-a877-466a-be8f-f8c8749cf1ab · outbound

This paper cites ”Rewriters”:The data reveals a fundamental difference in code modification philosophies.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts ”Rewriters”:The data reveals a fundamental difference in code modification philosophies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:12:59.044603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:7e16f63ed9958b14edd456d7f1cc8da280ddfb0eef699c684ceed3d108f4257c

Observation 61d522f0-f39a-4539-a665-2d4144a8ee70 · outbound

This paper cites It is the only model to record significant usage of update memory bank (22 times) and initialize memory bank (7 times).

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts It is the only model to record significant usage of update memory bank (22 times) and initialize memory bank (7 times)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:12:59.036469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:acad7bb207f0bc317b4dd67236b85f359e9709efc34ced31bed236965db4e038

Observation 99b10aef-6e15-413e-8db4-b27d48cd9f6d · outbound

This paper cites score": 6.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts score": 6

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:12:59.038886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:1390376a1461e83999107421ed5ab4e4888fc3245b423db4197ef363b56b42c0

Pith citing papers

Observation 9eb28bd8-1a4b-47e7-8d5f-f41e92880e3f · inbound

FileGram: Grounding Agent Personalization in File-System Behavioral Traces cites this paper.

FileGram: Grounding Agent Personalization in File-System Behavioral Traces AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:45:51.207473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:33:57.643549Z digest=sha256:f8ab45c931454657f678caccb8fdbfa7406da586b734b26ccee8f8a8f84ccd85

Observation 25f546a7-df3e-4b3d-b15c-5b6e6dc544ee · inbound

Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems cites this paper.

Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:51:28.126290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:25:01.741390Z digest=sha256:a51fa2f5c71d8cfb515c79ab438dca23bfef475f81a3ae562da28ff190a58330

Observation 25feaea7-7f7c-4832-b4c9-db0c1b886855 · inbound

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization cites this paper.

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:50:51.194696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:49:50.520222Z digest=sha256:a5ce05a0ae9e205e9fcd7b6c27bc928e355f0177b651b5ad845b38d9c10ffd48

Observation 11d440ab-8c10-4105-bf2a-f8cc43dba56a · inbound

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation cites this paper.

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:24.963224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T10:30:11.309910Z digest=sha256:337aae288bfaaafba8a6909349cbb1b7040d6a8d65c0d65ba051ecb1d5e72a90

Observation 063b7b09-54da-40bb-b4f5-722bd101cf7a · inbound

How to Interpret Agent Behavior cites this paper.

How to Interpret Agent Behavior AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:29:22.787555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T18:23:25.269217Z digest=sha256:29b20c5f6300e77eb4a77c24bbd1d6dd8ed765d8c0ec09ee3eed63ea27735608

Observation cca7bfc5-7a43-4f0f-a67b-683ec5c5cfbd · inbound

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades cites this paper.

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T02:33:32.341560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T02:31:18.183715Z digest=sha256:f215ecb622e4bddeed21293e601f03e86f1e1ffa0224e6ade3c75d5586ad94ef

Observation 3beacc29-90fc-4724-9f89-e59378a93e46 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T17:47:42.057431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T17:43:11.038874Z digest=sha256:1b420612bdc0187c6b79b3e9b53d39dbd944e7533f17d0899ead3e3a40a7aedd

Observation 9cbdcf98-2006-41b7-a88f-a34835be5e04 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T05:13:41.885703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:13:41.885703Z digest=sha256:b094ffcabda800dc7200f179cd8054aeb92dfe616c4547a5af88f0d6038ef12d

Observation 55c7fd09-43b9-42a8-a854-bf5e70a50742 · inbound

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents cites this paper.

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:13:26.973796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:26ec5036b6601e8c99a0a66a4350c60af589cf64437287c17d9cd09706e96811

Observation 0b953be4-78cd-4e97-a45a-f88b7bf53ccd · inbound

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems cites this paper.

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.159278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T00:05:31.780655Z digest=sha256:c5b45dda86df8725f3d7a9e840f02ada3beefc885479774ce4c8605fea536a0c

Observation 88d46761-5626-43db-bf1a-08836c3915f4 · inbound

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards cites this paper.

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T23:11:45.425149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:11:45.425149Z digest=sha256:c327045885814ad90614d9703f6f7913aad64feeec99a057013c350bc4f70182

Observation e87e7e36-1b36-4d65-b35e-eed4dc64a240 · inbound

LU-500: A Logo Benchmark for Concept Unlearning cites this paper.

LU-500: A Logo Benchmark for Concept Unlearning AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T23:06:24.702775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:06:24.702775Z digest=sha256:dacaddfe2574aad888b5de64ae4cae8406296dbe8f938ea983b43d05c4422285