Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T00:06:02.229517Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.25152.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T00:06:02.229517Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ba35ee0c-a4db-435c-9a67-67491996ea4a · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba18b0d-4372-4b9e-9146-a5d73a93a127 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb4a6551-b82f-4d42-a567-e2e50ff28a51 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3296c37-e76f-455f-9bad-bc9a8554d11a · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c371ef5f-2110-42c9-b13b-004fbc0ba343 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Deep Re- inforcement Learning from Human Prefer - ences,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e21d03-5745-4716-9552-e616413f705c · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops EvilGenie: A Reward Hacking Benchmark
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df82bd36-55ea-4ca2-a8aa-20f4bd708f46 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Scal - ing Laws for Reward Model Overoptimiza- tion,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d69b39d-f819-499e-8920-194b40d65375 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Large Language Models Cannot Self-Correct Reasoning Yet,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 912e8103-f563-4ad6-b9e9-e9a808abea03 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops AI safety via debate
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f926ccb-1d5e-4c05-9e73-e560f5900863 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Controlled Exper- iments on the Web: Survey and Practical Guide,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aba963f-f187-4d10-82b5-07710c08c0d8 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Scalable agent alignment via reward modeling: a research direction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0ccc12-c41a-48c8-98fb-179d494630ee · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7441adf-2310-4c39-991d-6033c8b791bc · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Let’s Verify Step by Step,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a0517dc-6130-461b-893c-de4ab8ddb17b · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Retrospec- tive Progress-Aware Self-Refinement for LLM Agent Training,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation effd6ed5-f217-46d1-9037-895159e7c9f4 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Self-Refine: Iterative Refinement with Self-Feedback,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f9bc9a-3029-432d-8d7d-9d4c6c2dffde · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Categorizing Variants of Goodhart's Law
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60514e2-ca06-4acd-b210-8d964cfcb204 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Training Language Mod- els to Follow Instructions with Human Feedback,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b084f1d-6454-4fbe-bb3b-74e0c7ff3cf0 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops The Effects of Reward Misspecification: Map - ping and Mitigating Misaligned Models,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0dc9f1-da99-491f-9b26-d10498874b10 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Spontaneous Reward Hacking in Iterative Self-Refinement
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b0bd32-9eb6-4620-8bb7-18cbf470cd0a · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops LLM Evaluators Recognize and Fa- vor Their Own Generations,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e51bc7-4821-4415-a6bb-a23c9b3c4d78 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Reflexion: Language Agents with Verbal Reinforcement Learning,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5889ad6-d8e6-4686-bbd1-cc028bf4d6d3 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Defining and Char- acterizing Reward Hacking,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300eacf1-8137-437f-891e-c2b8b8d2333a · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Solving math word problems with process- and outcome-based feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b51709-4ae3-4445-b4ac-b6c889540733 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cac049e-4bbc-4de3-acb2-ad957a09c817 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12ca15e8-2dee-4054-b2cd-3432047aef7e · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops SWE-agent: Agent- Computer Interfaces Enable Automated Software Engineering,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce56bc70-e032-48e9-b970-c352c3180329 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops ReAct: Syn- ergizing Reasoning and Acting in Language Models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd76697-74f8-4787-b2da-f22913978e9a · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f31ed6-b2f3-450c-81d9-0cda687c4198 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e525af4-cf19-4908-a32b-a58f072f4bd4 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb047815-d530-49fb-9be7-bb482fd10599 · outbound
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7bf308-a0e1-4d9b-9ac3-1e8ecb116fa5 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Building Effective AI Agents,
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceaeaecc-b479-4e83-98fa-9f6965791359 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb15f01-af43-4cde-b83f-623c58ad06e1 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops SWE-bench: Can Language Models Resolve Real-World GitHub Issues?,
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de3934e-92c4-4e34-bcee-718312eaf032 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Specification Gaming: The Flip Side of AI Ingenuity,
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7683982-5cd8-4534-b471-02c2786afd19 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290563ce-728e-4d18-917d-dcad6186e0c5 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Reward Hack- ing in Language Model Agents: Revisiting AI Safety Gridworlds,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27e3f52e-3632-40c0-a007-5f08b151d0ac · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Effective Harnesses for Long- Running Agents,
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1030f6f-d2ce-434f-9943-3b5aac86714d · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Reward - HackingAgents: Benchmarking Evalua - tion Integrity for LLM ML-Engineering Agents,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b1f4b6-1c8e-4072-b87b-7f93880075d9 · outbound
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Concrete Problems in AI Safety
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.