Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:28:31.474196Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.13705.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:28:31.474196Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a112a73-3bc9-4612-951c-4ead398886c6 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Claude Opus 4.8.https://www.anthropic.com/news/claude-opus-4-8, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 197f83e3-e884-4221-a905-adeffc39534b · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e756433-4acc-4421-a16a-e033447bb6a5 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Aider: Ai pair programming in your terminal.https://github.com/paul-gauthier/ aider, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0267d8-ec3e-49d9-bd85-4bbe61cacb1b · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Deepeval: The open-source evaluation framework for llms
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8cf8cc-22ef-4809-918a-9c251ef70962 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Opencompass: A universal evaluation platform for foundation models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb58207-0634-4b68-810d-6d1f5a17cd6b · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0621a85-e7c8-4429-bb8b-4a912848320e · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Benchmarking reward hack detection in code environments via contrastive analysis, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f28693-c565-41e0-b0f6-20393d4d0686 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc53bdf-3f4c-49eb-aef5-d0090c36e8c4 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Glm-5: from vibe coding to agentic engineering, 2026
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd8b8f3-f1e3-4fb0-9ecf-4919009c8635 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Gemini 3.1 Pro.https://deepmind.google/models/gemini/pro/, 2026
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9253900d-2b6c-4451-b062-8554366a7783 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Deepsearchqa: Bridging the comprehensiveness gap for deep research agents.arXiv preprint arXiv:2601.20975, 2026
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ede13a8-3b84-471b-b69b-284509b152dd · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Harbor: A framework for evaluating and optimizing agents and models in container environments.https://github.com/harbor-framework/harbor, January 2026
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2605cf5-3fb8-439a-a162-913c218c1c0f · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa3d58b-de38-402d-b2f9-55ec33559f62 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Waytowich, and Boyuan Chen
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778216b8-cfd7-42ba-9ad0-4a7ce3779317 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515bd38a-6e16-4164-94d7-361d9deb0406 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Langsmith: A unified platform for debugging, testing, evaluating, and monitoring your llm applications.https://www.langchain.com/langsmith, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eed8db8-3e40-415e-b90e-b0319d14ee78 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Camel: Communicative agents for "mind" exploration of large language model society
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff3998e-bff5-43ad-9dca-5334fd0dc904 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03511ce-8668-43c1-8e3b-9d94ad913832 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Agentbench: Evaluating llms as agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea47235-b6e9-4555-9be9-a5c3d6bf63c3 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Agentboard: An analytical evaluation board of multi-turn llm agents
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4fe4edb-6a58-43a1-a89a-dd60664ede31 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f239a4a5-de51-4a96-8ed6-93f86c96d445 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities GAIA: a benchmark for general AI assistants
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbbbfff6-7fa6-4e5c-b36a-0b8310de683c · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Kimi-k2.6.https://www.kimi.com/en/blog/kimi-k2-6, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d8c413-4466-4b37-b33e-aa99afdd43ba · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Introducing GPT-5.5.https://openai.com/index/introducing-gpt-5-5/, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926f82ef-149c-43eb-b510-5c63a8a3defc · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities OpenClaw.https://github.com/openclaw/openclaw, 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0821aa-585c-40b1-96b2-b63bd253c2f1 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Patil, Huanzhi Mao, Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13666849-670e-4f99-bd67-eded6936b09e · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec5f367-1567-4f62-8d14-1e4b669df13c · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Humanity's Last Exam
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecab5476-e8e8-477f-abeb-72eff4d96610 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Pinchbench: Real-world benchmarks for ai coding agents, 2026
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260d6c22-6de8-4797-8d2a-2c3fafcc2707 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Qwen3.5: Accelerating productivity with native multimodal agents, February 2026
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84cbb913-4bfd-4424-b91a-7bdc7c8f31fb · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities 2, 1, 4.1 10 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66914381-83cf-4a6b-bb9e-772d72f0424c · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities EvalScope: Evaluation framework for large models, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a09ceb-9aff-4413-8f89-761301debeae · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Terminal-bench: A benchmark for ai agents in terminal environments
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb6d0aa-92c0-461d-9201-071e9e450411 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Scicode: A research coding benchmark curated by scientists.Advances in Neural Information Processing Systems, 37:30624–30650, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff36f398-609b-4664-86b7-4fc02010694c · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Frontier- science: Evaluating ai’s ability to perform expert-level scientific tasks.arXiv preprint arXiv:2601.21165,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76426af5-bab2-4688-a7b0-9e3c42294efa · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Openhands: An open platform for ai software developers as generalist agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925e1fb1-34ad-4886-a120-3bc1b06c90c1 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9af247-75da-40ef-8ce2-a2f447e9aefb · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Autogen: Enablingnext-genllmapplicationsviamulti-agentconversations
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2f40a7-5177-4628-a744-fec252b6d472 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Mitchell, and Yuanzhi Li
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1ef50f-53d0-443b-86cc-7469c83d35f1 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Agentgym: Evolving large language model-based agents across diverse environments, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26cd8ed4-7d82-40ce-b02c-f3dc5927545e · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348, 2026
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc637492-3a41-4958-b1a6-694c775d4164 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c33621c-975c-4203-8a9f-f86446e16545 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Probing scientific general intelligence of llms with scientist-aligned workflows.arXiv preprint arXiv:2512.16969, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea311c2f-4163-47fd-89e5-73c79f2ffb0a · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities SWE-agent: Agent-computer interfaces enable automated software engineering
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966e94a1-5a41-4de1-9250-2658777dbadd · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2602e56-b100-4cdf-b869-33d3361eba6a · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be112a3-3113-4d1a-a1a5-cf830076da6d · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Maslab: A unified and comprehensive codebase for llm-based multi-agent systems
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f64c576-bd5a-4cd7-97a7-2df81a0ea21c · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Hle-verified: A systematic verification and structured revision of humanity’s last exam.arXiv preprint arXiv:2602.13964, 2026
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0555dc90-65cd-45b4-8ee0-76f1a24b4115 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a261d788-53f4-44ef-8fe5-4c0af4caab2d · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af9b981-42f3-4e70-b372-90d68eef9244 · outbound
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Intern-s1-pro: Scientific multimodal foundation model at trillion scale, 2026
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.