Pith. sign in

Paper Citation Record · LEDGER

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

As of 21 August 2026, this Paper Citation Record lists 100 of 129 outbound references and 4 inbound Pith citation observations for arXiv:2605.12673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12673 v1

Coverage vector

measured 100 of 129 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T20:31:50.043920Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:35.095351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:12:42.362881Z

Reference resolution

100 of 129 outbound references displayed

  • verified exact37
  • verified fuzzy32
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 891ab41e-bc43-41bf-b548-03096873f857 · outbound

This paper cites Concrete Problems in AI Safety.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.730139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ae6ee4f9a8bd156d71dbddf7db4e73e450b88edcc186c62075b6bcf958c8ba73

Observation a7fba226-432f-4afb-9b10-69aeb0ab1979 · outbound

This paper cites Alignment risk update: Claude mythos preview.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Alignment risk update: Claude mythos preview

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.046492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1a902a0fd5795254e987b13fe7e1c85cb5e96da1e1211aa1c6a1c60c0ac3f4c8

Observation 8c849738-2c8c-4f25-950e-b6b7edc8ddef · outbound

This paper cites Claude code.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Claude code

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.050001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2745292f25bf535fa881649c66c91d83320daa70d1b565d95e7be0065ee723cc

Observation 971037cc-e3c1-4c2e-adfa-5b98ee9ae3a7 · outbound

This paper cites Analyzing and improving chain-of-thought monitorability through information theory.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Analyzing and improving chain-of-thought monitorability through information theory

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.724381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fd6c29a66aee08e862cbd7e7bb75dd5eb6e7456464d87f72d59e1ced2b093212

Observation 6db6f338-5a7d-4c70-b07a-8bd2834e63ea · outbound

This paper cites Rewardhackingagents: Benchmarking evaluation integrity for llm ml-engineering agents.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rewardhackingagents: Benchmarking evaluation integrity for llm ml-engineering agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.768744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d3d0b8a4976fb3fbbe62fa212f1e0513e059e4f822713837e73bcfa5fd2658da

Observation 6a176a57-f55a-4f09-acf9-c25e441b7e5f · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:24:13.179886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b1fbea67a678042b7c66d515a6d7043486536d56c3b1e0c84b023ffc590fda86

Observation d8bc6e99-076e-4134-b5b5-c55f67b5e467 · outbound

This paper cites Adversarial reward auditing for active detection and mitigation of reward hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Adversarial reward auditing for active detection and mitigation of reward hacking

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.897304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d820a0286b6c93da194645528a74b0862454fede3ac6c382037cde18a86f163b

Observation 8741e97a-d21d-4b4d-8341-40c8bcb021d2 · outbound

This paper cites What Will it Take to Fix Benchmarking in Natural Language Understanding?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack What Will it Take to Fix Benchmarking in Natural Language Understanding?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.901133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:c8679be71fa6630751d3def7948dde9881d62f1a0c50bc5353cfb2a2c8676393

Observation f1f887bd-a17f-4a92-b2ef-065878966ef9 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.909283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:691cc97af767a426a3ced20ee31247a1e0f81369a59e54a90ac842e90d3f144e

Observation 0ec8fc1c-28d0-4133-abd5-e6e0d7caed93 · outbound

This paper cites arXiv preprint arXiv:2502.17521 , year =.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack arXiv preprint arXiv:2502.17521 , year =

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.892065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:dea6e0da99892c79bb7dde761c2562ba0249fab704ccaf9ec1ac99ad9581c662

Observation b84fc53d-0b68-463d-b08b-8de0dee67701 · outbound

This paper cites Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.883357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0de8c6c8514e9d3a9a6ad7af767f95e51e807499974662203e8cd467c4f9ed5c

Observation 24b201f8-68b0-46b3-aafc-03a13b094649 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.763151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:bbb66921f7753f6b243c85e3a2d237ae55071fb66d8a8e982824e56b93d74d71

Observation 5b96a8e9-d100-49fd-9dda-f8e395fc79f9 · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.937213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5e9d81d87ec1486836be94434265b252cd2eaa1fd1bdf6aa86e7347e9604ea72

Observation 5f57ad81-487d-4945-ab52-ff5ed9280a0e · outbound

This paper cites The Benchmark Lottery.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack The Benchmark Lottery

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.887648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:9624adcd1ccd0bbd61a7304e11261e2edd58f0a6564b80e7bb4ed8be2fd25f12

Observation eb042689-a284-479a-8f12-f6a3a55d4fd3 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.742365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b261dd0012267162f57e9e7ac50f0795567ce0e748051be6e98a768f84c74128

Observation dcf2e410-a702-4925-9ad1-6e2cc7f80457 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b32b3a1106c72afbfcb0f0f2d7b068c75baa434ae24a400927d8956f0afd7c0f

Observation 688e7832-5307-4d62-912d-270b8db9bd2a · outbound

This paper cites Benchmarking reward hack detection in code environments via contrastive analysis.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Benchmarking reward hack detection in code environments via contrastive analysis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.872818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d170b183969bdb749bd09e848fac4f77fcb28a0b11bb5c824ee06545ab48fbb1

Observation 20557408-85e8-4ef3-bea9-6014f7709f51 · outbound

This paper cites Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.877863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:38a3ddaa42f080342d7e56a465161da2a52f0d2a1147549253711c9cdee07798

Observation b8231cbf-0c30-4838-ba88-0348a5a08f30 · outbound

This paper cites Generative adversarial nets.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Generative adversarial nets

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.963694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:24dfcba92215ce53f6e9556913b97664a3ff26fc469e2d450985af10a2dfdebc

Observation 970d21b9-f531-4e21-bc6f-1d3c77bd1dc1 · outbound

This paper cites Problems of monetary management: The UK experience.Monetary Theory and Practice, pages 91–121.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Problems of monetary management: The UK experience.Monetary Theory and Practice, pages 91–121

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.938448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:71b709ff7e9fbee8f2e284d40acfe14d94b8d7db799f115bc6938b8b13b9987a

Observation 55f6a03c-a705-493b-a3f9-0f7671f381af · outbound

This paper cites Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.914038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b4c5960a9ca7b44022fd13b636d4f15f8465252eb94456207aaa2613c67c0a99

Observation 91f824d3-4cc5-470e-a739-a6a1d5fe2d9d · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.970116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4385af3dfcc919ca1c4ca60ec7fce0de427df5b7cd6765984a0033aaf8ebf62e

Observation 95ab70cd-b430-4086-8a80-4793071f6217 · outbound

This paper cites Issue #14: Iquest-coder-v1.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Issue #14: Iquest-coder-v1

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.873321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1afe29fa71cd10e56b089d5a6b463a515287113498bf1482209958d7fca658b9

Observation 75d75803-1887-43d6-aa4d-943e7aa756fe · outbound

This paper cites Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.875651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2ba801bde77e3ebb938cdecc545caed04cc9d5d1b82c083d602aaad77e2e7550

Observation cff115f9-91d8-4e04-92c3-93e4e40e7104 · outbound

This paper cites Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.943868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:21b4738e6aa86df23129df6ff27b33ae95340490c7d51061bfdcb299ba60ea9b

Observation a4fecea0-9a15-419d-b761-be0e5ca96581 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.752707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5b856d5ab07257c4b94a1253b337939efecc904b7924e06f8656805db57be67d

Observation f30ec8bd-62c8-46a3-871c-e7a487cee1e7 · outbound

This paper cites Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.747821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:90d09223492bb9b3c80000b18684ba092ba757d8ad9b3d379accc1563a40a584

Observation bc9f9a25-82dc-4095-b9c5-322d5d6ad964 · outbound

This paper cites Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:19:45.019656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:36d66a8554cd44af701101fff4275be2c5db245a05823d8a1d13d371da340d00

Observation a88031bd-cb72-47c5-bb47-926e1d8b84e3 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.794085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:6f063dabce3b9dc1905bb7895848adcd0a3bf4ece835ee5ea5f6c6bab8f19d4d

Observation b4c7eac9-a96e-4e31-991e-2af5ca23690d · outbound

This paper cites ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.736935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:cdf16a389140e776c9d5e8c855c982577769982b320ae36df916893dc8366cbc

Observation ea3e3b66-fad3-4dd8-a431-b6b5a0afc290 · outbound

This paper cites Diagnosing Pathological Chain-of-Thought in Reasoning Models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Diagnosing Pathological Chain-of-Thought in Reasoning Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:05.474485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:98687b1b35529747556007cbdc4cb94323138de592f9298ceeac11e9371522d8

Observation 7c30f080-14de-473c-9660-a283785d9642 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack AgentBench: Evaluating LLMs as Agents

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.927996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:59a2af306385c0d81f17378c9ed92ef27e1345239d21337016256686f94e12cc

Observation 45f6f51e-7648-46fe-9c3c-cc0e84f7161f · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.933390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2010bf514975f3c4ee1d98a2bea0ab6c4bed6bf4fd1b01d61e4d519465958c1a

Observation b23de401-40ff-4c31-ae35-90626ce4c12d · outbound

This paper cites Gonzalez, Jingbo 12 Preprint FrontierCS T eam Shang, and Alvin Cheung.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Gonzalez, Jingbo 12 Preprint FrontierCS T eam Shang, and Alvin Cheung

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.938683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d6a1ea12013e6ab17449d8c14be9ce7953693cfbe037f90c439d8c8d3914aa41

Observation c32812f1-7fb8-42ee-940e-4ce580779882 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.949393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:260def909f75e23606410890ec2460350c6e5c7b7b312a2531a1faadf995ccdc

Observation 8b30275a-5e73-4e8c-a910-49f21a762280 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack GAIA: a benchmark for General AI Assistants

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.965337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:74138d99de42e95de57240af3988e949e21ed70bb7c5c6efa9b11dd5b64eddaa

Observation ae9e7d26-2de6-4bdc-8c2f-e27c4060e10a · outbound

This paper cites Introducing codex.https://openai.com/index/introducing-codex/.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Introducing codex.https://openai.com/index/introducing-codex/

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.877941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2dea7862327aefaa820286e68c3c592c0b59bfae40d6d7a27a31a98cf7840409

Observation 438b2977-8481-4dd6-a6d4-3c45a93a9207 · outbound

This paper cites Why swe-bench verified no longer measures frontier coding capabilities.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Why swe-bench verified no longer measures frontier coding capabilities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.866415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:3a4e64bc38c5632a10ba5d19a81bb03a17c3456149b9d758906cce5f61d53aa5

Observation 824532d7-400d-40ed-8919-f33015637cc8 · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Proving Test Set Contamination in Black Box Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:14071501c6de4c137d9e742e0b9c3c45237e744f5dab1d608e3ec66a0b60dd10

Observation 6637ef59-6403-4448-885f-50d32873c0e6 · outbound

This paper cites KernelBench: Can LLMs Write Efficient GPU Kernels?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack KernelBench: Can LLMs Write Efficient GPU Kernels?

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:55:02.220088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:93820d038010ced034f909f39f6a34b660c76165b4cbf8cda046cda409e0efe7

Observation ffc46a41-e1c3-40a8-a160-731b1f3714e3 · outbound

This paper cites Feedback Loops With Language Models Drive In-Context Reward Hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.960515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5f7532df25a1104df33cad558ccb2581ce1cf6d921258c79e5958e9e9d789667

Observation 9831abea-ee67-4253-a4c0-34ef70e079d2 · outbound

This paper cites Frontierswe: Benchmarking software engineering skill at the edge of human ability.https://www.frontierswe.com/.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Frontierswe: Benchmarking software engineering skill at the edge of human ability.https://www.frontierswe.com/

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.929623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fe9e8ffde0d34985390c1edfdb2fd19fda825b76787861c09b9593abe91ba2a4

Observation b7ef27b5-44f3-4959-a71d-bcadc872f6b8 · outbound

This paper cites Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.914374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0b18721ae56b25d8364e1887d09f06ba5bb6dd16499d891e693a2422cbc21184

Observation 13b2a7bc-488b-4033-9d77-e8a68d539c75 · outbound

This paper cites Posttrainbench: Can llm agents automate llm post-training?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Posttrainbench: Can llm agents automate llm post-training?

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.863338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:6747e3d88a6ce20f35f1692be6ed00b32cc87776fdc2a5ca3ec861b647a645f3

Observation fc2b3a61-925d-4124-803d-5b8dac7f352a · outbound

This paper cites PostTrainBench: Can LLM agents automate LLM post-training?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack PostTrainBench: Can LLM agents automate LLM post-training?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.858065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:16829898e85f485a6da30d6886eb32a973aa023f3eccaaa3bcdfa5a1781b8989

Observation d6b62d88-4d53-4aff-9442-31fa43fc5179 · outbound

This paper cites Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.848195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:f68a8b0dca4aa76468682a197028d8e5e598d4f08b32fb7379214fe377397228

Observation a1f7e64c-02e0-4783-8fc6-6403b05fd57a · outbound

This paper cites Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.900600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:e3e2cb90c008d368d84b3ff0a3e35e4a09e446de9458fe76f33743a719fbe431

Observation d4512d50-caf6-408d-91fd-53055ffd103e · outbound

This paper cites Defining and Characterizing Reward Hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Defining and Characterizing Reward Hacking

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.842198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:94e26b5023418379faf60d89d84b3df45cc653810d1ece73fa7b5d890d46f4fe

Observation 209f8941-b88c-4e25-860a-102c191ffd50 · outbound

This paper cites Detecting Safety Violations Across Many Agent Traces.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Detecting Safety Violations Across Many Agent Traces

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.820416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:a2bf797b1a0ac57b67469ee528088e3aa650bb662409f0b393cd1d0649334874

Observation aa31985a-c5c4-4a1c-8cf4-12c62ae7e30f · outbound

This paper cites improving ratings.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack improving ratings

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.909376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8399beecdc4778eb986b5f97b5badaee99d7798595fbc926b29b19135e5bd396

Observation 17b8a3d1-a9e4-40cb-9dbd-82c75f668e6d · outbound

This paper cites Recent frontier models are reward hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Recent frontier models are reward hacking

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.871075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5927c60b547aba1affe600120439b1ac60fdffcfb507eac72da53060b989d793

Observation c14b6047-9106-4091-a5f1-5cbb034ea167 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.921853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4798f8b93aa1945e4cbcdb14ebe3b68b80c9f3ddd3fb8b4c65a2e85d3bf31e9c

Observation 2ea1b0e3-a0d8-4623-8742-42e1973d8e79 · outbound

This paper cites FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.803955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4553111f29f757812659a9a7295b7c3a5c263381a554a5633650e3f06bdbbb8b

Observation f8358352-54c7-4d15-a484-b46f2d7741a9 · outbound

This paper cites BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.830494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fc274f17c7a33fadefdf5d17bc33d4acfbaefed77826ceb4c7bc8dbde445675b

Observation 905d6597-1ec7-40cc-b0fe-10b17eb2335f · outbound

This paper cites Detecting and Suppressing Reward Hacking with Gradient Fingerprints.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Detecting and Suppressing Reward Hacking with Gradient Fingerprints

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.825442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:668c37ed7887e19dd0a51c14d88a87378c7c43a10ea48b50ed8f1802e22634db

Observation bca79f29-383c-45bf-aeec-fa5c741cfd68 · outbound

This paper cites Reward hacking in reinforcement learning.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward hacking in reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.934517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:777c48b205ad4d6bd612822626189840dc272d9b784bfaee7231ea49d8a89a95

Observation fc031751-6045-4792-a429-c80663b4a523 · outbound

This paper cites Monitoring emergent reward hacking during generation via internal activations.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Monitoring emergent reward hacking during generation via internal activations

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.809401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:031dfaf8656b621368cc45b16957fe87fe725982ce649fa46d3a5297b6b3d1be

Observation 129bda9f-d70a-4730-97fc-d03263e9b001 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.835280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:7c3faedd5cceee59c23834fd9720057d34ffb4ea5f049ea547c8d3c29a069897

Observation 28f67635-80e8-445c-b423-2b71c8b42b35 · outbound

This paper cites Investigating cot monitorability in large reasoning models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Investigating cot monitorability in large reasoning models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.853082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:53b5f936407b17131c648b89aad5c96c38ff33119697bb87a92541b4d51ad1f5

Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.862291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:49637aafca6ec90f5ceb7d3b4b8972ddc827a4b67ee789c1c4ce6d4b8b77651e

Observation f7878bb6-d958-48d6-a10f-13d2cdb9def3 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.866671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2d92c6fead413d023808128b6e569ae7a8899e98647e80afb063ebb8155ee4b8

Observation 4f495794-702f-4bb1-8343-5752ebf469ed · outbound

This paper cites UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.814909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fa722ae96e0491934c74c5cc6565af48d56758636440bc0918d0bba4fdd01c4a

Observation b73a68c2-028a-43f6-8f0d-81d129b78d1a · outbound

This paper cites Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.854939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1ac64d522f00f6d8ef8c79153f27970bf75072c794703449fa235c1b7130e981

Observation 8b205cdf-96ee-41f0-87b9-3b0c39b8354b · outbound

This paper cites Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark,.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark,

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.954541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1813279e336fe0919adc739e6bd2df3d8541ba6ea9c8cd74b38a418f92491e1d

Observation 610552bc-b577-4555-af5a-0e0021d293a0 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.780884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:cca881ada924a793380b4f600444f948bd7928414e8fb53428fd78818834bad0

Observation 5545b94b-9fde-4482-99e4-00978f65bbff · outbound

This paper cites Netpress: Dynamically generated llm benchmarks for network applications.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Netpress: Dynamically generated llm benchmarks for network applications

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.788338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b37d66bf7db66a34ee00ffdcb24ba7b4c1acd69e7c755b1a05d540fa3fa0c6e2

Observation 9892f498-abd3-42e6-8b3b-f8dc35dc75ac · outbound

This paper cites Establishing best practices for building rigorous agentic benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing best practices for building rigorous agentic benchmarks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.907599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1a4d2e664dc1d84048d8bebfc467fee5b788105872e787fbfede2b9b9f103686

Observation 54fc18c7-13b3-4f01-ba8a-8a9115d6c13d · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.775266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0679297ade961ecdac4bb5337a1062f710d2e18068d1509df03a640ffd768464

Observation 1ea0ce51-533d-4e7f-bf2f-395262f6381f · outbound

This paper cites [exec(\"\.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack [exec(\"\

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.861039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:329aa5c27741b7220e7edaf00725e5b6a810c629886ddd44bb07e6e18425cadc

Observation ba73d78c-ee42-48ed-ac09-221cf49232f0 · outbound

This paper cites {name} benchmark github.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack {name} benchmark github

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.940010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:e4b14963241d2249fea668d56b533defcf3c494a1900d0aa48c65880d482011b

Observation 47673d6d-bb9b-4c2a-abac-182d16123638 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.878270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:192d4e3d852e7550cecf0b1c36aff04e8c9de1f89097ffc1a296575904093cb9

Observation f2e5def7-b746-4fa3-a8bc-47e48a4633f0 · outbound

This paper cites 7 8If the benchmark is well-known (e.g.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 7 8If the benchmark is well-known (e.g

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.911294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d8146e2d850fbc31fd703275cf615fab9adc3d5dcc1d712dbeb66aa71b021757

Observation d0c82d64-3b08-40bc-8dcc-a4486a5fe175 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.903132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:860cae7805cd91b348b6e8d9edb857a6226e9b80b8cf1f6fed1e079576f61086

Observation 5198de44-7cb3-4395-bbf6-441549a2d6de · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.844960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:811c6883bce9bb5e62c82c2ad012f4a7631d71a2786be77f32b14759d25a7216

Observation b85deeaa-af8f-4286-a1d9-333e7bd619cf · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.919417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b4a25cfa77ddafa39addeb98c4d130aa0eee0ef433fc9da24fca161fe48aaaa1

Observation f4714beb-dd3a-4113-af0a-3db315f5d17d · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.022699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ea9ef5e954476b32872edacd54c1dfae112de67528cf5570ec891cf48f5b63b8

Observation 6fc5e842-9b62-4f5d-b65f-929ecaf57aa8 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.977947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:26ffec98f56100d012ae0c171d72b8263f9cdbe0aacdd0612a95b3f0e0380592

Observation 8ec067b8-c9b6-4c7a-ab80-4e19bc931d68 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.955253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8401d40364a5cfbcd119179a51b382271059dcb05cfe2cf3d5cde1eaed0cda20

Observation e04210aa-ef86-412c-81eb-d8dbb5a4e75a · outbound

This paper cites task_id_1.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task_id_1

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.926837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:842563a27ed7501e6253b1cbb2869af4f6956674bd88702c11075adcc749026d

Observation ca9cc7de-6c67-42e0-a53e-e566e5933873 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.056466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0d94c55ed57bae0e2ab6eaa1fba0e6aca8a927a467afb46da49a16b9c27fd433

Observation 307d07ad-d9a2-495e-aa24-d94621d5a73c · outbound

This paper cites task": "<task_name>.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task": "<task_name>

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.059609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:61a8cbef06a5be450d506395ac0b1ae33580b4282d79272081d8c8fdd5cae781

Observation 2d2c5ca2-b64f-4587-8be1-c70625988b5a · outbound

This paper cites If it’s a URL or package name, clone/download it.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack If it’s a URL or package name, clone/download it

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.012830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8b2c17dae3c8d9efe6e9c2e20c67fd25d4ef62ebd7a219a342fbfcf6564a16f7

Observation f2dea01d-7d8b-438d-8487-73ad40b591fe · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.015995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:23c81535152ca3545562a97fa62512a8788ba8888ce86a5fb309ae09b9e5b0a6

Observation 3b1cea2a-a412-4bf5-82d2-cabfff1562e3 · outbound

This paper cites This is critical -- every point where agent output touches evaluation code is an attack surface.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack This is critical -- every point where agent output touches evaluation code is an attack surface

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.019458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5e46ddb6b17e4690809a754150b4aed53b3e0b6addb104963cda9fc14ada81c3

Observation ff6ad4f3-6093-480f-b0e7-4aedf0b2eb71 · outbound

This paper cites Report: 49- **Docker images / large files**: Does the benchmark require pulling large Docker images, datasets, model weights, or other heavy artifacts? Estimate total download size.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Report: 49- **Docker images / large files**: Does the benchmark require pulling large Docker images, datasets, model weights, or other heavy artifacts? Estimate total download size

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.026327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fb81196f3e9d8bcd18b192b88896d31cd465a6ab1e1e23181a4d229ad7c420e3

Observation 3bbee520-003f-40ba-bd83-ac8fb7da8e7f · outbound

This paper cites task_id_1.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task_id_1

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.029659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0214edd2a2c51bf142ef2f93555dea8d746a495b4f290170602a18c2756d7508

Observation 85201882-31d6-43ca-96ce-04efc05db9d0 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.032765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:10178972663589179598ccc0764ea19e7f534172cfdb664e318c10a9be6c623d

Observation 6a7ac56d-a637-4e35-a07b-3d16d8f49736 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.002636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4ed9b26bc0e73b0168909e243f84be402d9ea3e452ed131b4b36b6604d15b5ae

Observation 326e90fa-5212-4f40-bf44-0f1282fad61f · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.009720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:3e5260da362f873a5be5e98df639a8bcaef69a35d8afc7ec00a82fbed8be057f

Observation 233c198e-ac02-427d-93f0-db24c9e7758c · outbound

This paper cites 275- It should set up the environment (install deps if needed), inject the exploit, then launch the evaluation.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 275- It should set up the environment (install deps if needed), inject the exploit, then launch the evaluation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.992674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8ab284f7d2bba00c0148c7946bee37c7ee8e9cd4d57dd0273d946e3f23072657

Observation abd3adf9-501a-4497-a496-d68af95d6cbd · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.995547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:f620eb122044ea83cab5444794804f4dcd8cfd1517eaf92cd77798a4cca4af94

Observation 99c8d5b7-00ae-446f-8082-5d32de178fac · outbound

This paper cites 309- Examine the task’s specific evaluation logic -- some tasks may have stricter checks, different scoring paths, or edge cases the current exploit does not cover.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 309- Examine the task’s specific evaluation logic -- some tasks may have stricter checks, different scoring paths, or edge cases the current exploit does not cover

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.987704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:f5fc9eb9a597cd402dbf227781cff2a812df0717a63e796ef664477b253225c9

Observation 0150ac14-3ecf-4e65-9c62-c2da0613dc43 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.981517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:91485f7df1debd868ba1b20a98c7e0c52f7d3205d5e36e79d4107b089a427c12

Observation 6b4a04f6-1358-40af-9477-53587d0ece94 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.984613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:45c84bbb86a57e9b31186e015fd0e11c3a3b9fee9135c9df802546a90b03a144

Observation e7ae5a2d-1d21-4d41-b9f3-74c388754970 · outbound

This paper cites 316 317Each iteration should be a deliberate improvement -- do not re-run the same exploit unchanged.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 316 317Each iteration should be a deliberate improvement -- do not re-run the same exploit unchanged

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.006600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:766974eee3725227eac0f99348dcc56842e7f50d15934faa82688b79332a6b5d

Observation 9fdcdc83-c97f-41e7-880c-f4909279ad73 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1d8e99163b48db99952c3bc9bdfba42c6e1d6969d47aba520297f7c6d911d93a

Observation 6c058542-f296-479c-a109-35482807ce44 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.964024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8047b18a0b708019429a9e2dd5267a638f795d3992eaccab66387c2fdfb1a405

Observation e3b51c43-9b8f-4d8f-9617-5d161b24108a · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.985706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:bca8127a8ba64e2ebdac63e531c9a75f748e8d41dc14e989f55399b69dbb078c

Observation f1a41981-224f-4419-a8c3-2dfaaeb0ed35 · outbound

This paper cites task": "<task_id>.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task": "<task_id>

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.049423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:38b5aaaa93d8e5c27142db0b527e45608705832d49ab08ee160abf7147a0718b

Observation 3bcf1007-d79c-4416-a94e-aa11d6fb7d38 · outbound

This paper cites If test.sh aborts before reaching compute_reward.py (e.g., set -eon a missing file), Harbor still ingests the pre-written value.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack If test.sh aborts before reaching compute_reward.py (e.g., set -eon a missing file), Harbor still ingests the pre-written value

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.967441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:65aed1816863e70d1ad151e55ec997bd23dce952c2ba25c86558339078815bf7

Pith citing papers

Observation 6efe1fcd-1a88-46c9-bf30-345e8cf79660 · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:27.636193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:27.636193Z digest=sha256:1efe44d7af577ea0cc0e8a78e5e83dd6c4bf3ffe3f856bae418dd3f5cf03dffc

Observation 070048b1-064a-4cc1-932d-4c002e77f329 · inbound

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution cites this paper.

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:12:42.368939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:12:42.043626Z digest=sha256:907b212736a3b3f2bcba9c67e309612a359b1f50f383fe9c389a88450405f395

Observation 79f9d0c6-a0b2-4cbd-a4ed-ecaf38d7708b · inbound

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution cites this paper.

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:29:25.322682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:29:25.322682Z digest=sha256:1abbf517918145d3a4dd4e29da37ff0b7f6f16527095a326920e1904496685f0

Observation 35ee17cc-9a66-49da-ab10-9323b7abab82 · inbound

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference cites this paper.

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:35.095351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:38:35.095351Z digest=sha256:ef4ff9fc381246a29ca6f559dd3c70dcfc73027e2dc170794fb26ac544439412