Pith. sign in

Paper Citation Record · LEDGER

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.13705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13705 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:28:31.474196Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a112a73-3bc9-4612-951c-4ead398886c6 · outbound

This paper cites Claude Opus 4.8.https://www.anthropic.com/news/claude-opus-4-8, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Claude Opus 4.8.https://www.anthropic.com/news/claude-opus-4-8, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.157020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.157020Z digest=sha256:7b34e9e03c0a3d1966a8f7df8311e115f8adc00a795009b93b95ce2ec2013272

Observation 197f83e3-e884-4221-a905-adeffc39534b · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.267570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.267570Z digest=sha256:e729e921af0fd5124317baf4fc4f6665a630df2d697bb35e511e08a665bb4160

Observation 5e756433-4acc-4421-a16a-e033447bb6a5 · outbound

This paper cites Aider: Ai pair programming in your terminal.https://github.com/paul-gauthier/ aider, 2023.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Aider: Ai pair programming in your terminal.https://github.com/paul-gauthier/ aider, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.341590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.341590Z digest=sha256:9dd88d3dc2191a83fe83ec36ff54845c06ce52f773581b4a6d5af80baad3472f

Observation 1c0267d8-ec3e-49d9-bd85-4bbe61cacb1b · outbound

This paper cites Deepeval: The open-source evaluation framework for llms.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Deepeval: The open-source evaluation framework for llms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.414940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.414940Z digest=sha256:70f13a728cdd8e154e9de25da8db26955ec0fb26173444fc57b988b57d2ff264

Observation db8cf8cc-22ef-4809-918a-9c251ef70962 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Opencompass: A universal evaluation platform for foundation models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.533532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.533532Z digest=sha256:a3c5030a2be56dc10a2ad7b70438739573d8c16e0ba4d722db7b177f11c0ad8b

Observation 8cb58207-0634-4b68-810d-6d1f5a17cd6b · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.606904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.606904Z digest=sha256:0708110295654348166542fd07a14aab89ee3ef27cd7080a46d0851265845af7

Observation b0621a85-e7c8-4429-bb8b-4a912848320e · outbound

This paper cites Benchmarking reward hack detection in code environments via contrastive analysis, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Benchmarking reward hack detection in code environments via contrastive analysis, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.727577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.727577Z digest=sha256:1f6780db49a4d1b272df9634a19f09c0dc9b40f24a8c5d626cac499dc296dc91

Observation 47f28693-c565-41e0-b0f6-20393d4d0686 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.849276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.849276Z digest=sha256:08d3373dd6abb63a9d1272a12bdc4c3b38ade64ea66f098d035be0131c694887

Observation 7fc53bdf-3f4c-49eb-aef5-d0090c36e8c4 · outbound

This paper cites Glm-5: from vibe coding to agentic engineering, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Glm-5: from vibe coding to agentic engineering, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.907095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.907095Z digest=sha256:fe5a5c76626eb91e1c4054f3d92518cf58096ee83f4a3c9302d93f71f408392d

Observation cfd8b8f3-f1e3-4fb0-9ecf-4919009c8635 · outbound

This paper cites Gemini 3.1 Pro.https://deepmind.google/models/gemini/pro/, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Gemini 3.1 Pro.https://deepmind.google/models/gemini/pro/, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:26.985290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:26.985290Z digest=sha256:a5b3baf2f2dc0b9ef439ebc55ea406f58e6dc53d1a5e67d076f7f664cd44564a

Observation 9253900d-2b6c-4451-b062-8554366a7783 · outbound

This paper cites Deepsearchqa: Bridging the comprehensiveness gap for deep research agents.arXiv preprint arXiv:2601.20975, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Deepsearchqa: Bridging the comprehensiveness gap for deep research agents.arXiv preprint arXiv:2601.20975, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.068755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.068755Z digest=sha256:823ea36732e2bf83346c2375ffd57f950474f8b98ca8a513aa757300a7f60f92

Observation 7ede13a8-3b84-471b-b69b-284509b152dd · outbound

This paper cites Harbor: A framework for evaluating and optimizing agents and models in container environments.https://github.com/harbor-framework/harbor, January 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Harbor: A framework for evaluating and optimizing agents and models in container environments.https://github.com/harbor-framework/harbor, January 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.178873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.178873Z digest=sha256:4df53e2924704e16c0a2990d99cea473d0e433944b26f76cf3528338fdd71651

Observation e2605cf5-3fb8-439a-a162-913c218c1c0f · outbound

This paper cites DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.333960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.333960Z digest=sha256:12c718a850a52cd8bc77b97259c8b3a1d64b6ab6fcf3b3b5f851269f9e98c615

Observation 1fa3d58b-de38-402d-b2f9-55ec33559f62 · outbound

This paper cites Waytowich, and Boyuan Chen.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Waytowich, and Boyuan Chen

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.464856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.464856Z digest=sha256:15926f4f3c0bc29c005da565491ca40b34cf88db09cd33c8733b656ead27c2a3

Observation 778216b8-cfd7-42ba-9ad0-4a7ce3779317 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.576474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.576474Z digest=sha256:92a1a30eb22566fa856fb6cdcc5b49b113c95870a4a1ca1e2b6680dc319b74a5

Observation 515bd38a-6e16-4164-94d7-361d9deb0406 · outbound

This paper cites Langsmith: A unified platform for debugging, testing, evaluating, and monitoring your llm applications.https://www.langchain.com/langsmith, 2025.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Langsmith: A unified platform for debugging, testing, evaluating, and monitoring your llm applications.https://www.langchain.com/langsmith, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.656925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.656925Z digest=sha256:5db212c6935e22981f22905ce3ecfc27e5a1947f38e991b0ed02af4f23f5853f

Observation 8eed8db8-3e40-415e-b90e-b0319d14ee78 · outbound

This paper cites Camel: Communicative agents for "mind" exploration of large language model society.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Camel: Communicative agents for "mind" exploration of large language model society

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.756312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.756312Z digest=sha256:effbd905447dd38261c24ecb05e267d0274216b7a270231827cbfbf34e57fc54

Observation dff3998e-bff5-43ad-9dca-5334fd0dc904 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.859425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.859425Z digest=sha256:f5f00914315cadcfc9e695267551b147c419885f98c4b93bad5ccc983917e1cf

Observation d03511ce-8668-43c1-8e3b-9d94ad913832 · outbound

This paper cites Agentbench: Evaluating llms as agents.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Agentbench: Evaluating llms as agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:27.968828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:27.968828Z digest=sha256:cc135c48cc59725c7f3caa38bae73754e7b269181757cd8bcac109c2e36981b1

Observation 7ea47235-b6e9-4555-9be9-a5c3d6bf63c3 · outbound

This paper cites Agentboard: An analytical evaluation board of multi-turn llm agents.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Agentboard: An analytical evaluation board of multi-turn llm agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.078610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.078610Z digest=sha256:34c9253f52d4f25ce3530be1b3126d896b216764088f7f3bda21bc4cea539062

Observation f4fe4edb-6a58-43a1-a89a-dd60664ede31 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.175736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.175736Z digest=sha256:0319ce14f4ee42078ad4de3cbb2ddb738ebb04a10dc46dbd0f5c0207df478026

Observation f239a4a5-de51-4a96-8ed6-93f86c96d445 · outbound

This paper cites GAIA: a benchmark for general AI assistants.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities GAIA: a benchmark for general AI assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.281304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.281304Z digest=sha256:1da85d4d6f3099b01fe88c77fe527f84fae83d74ec22346e48da2f7151ae7010

Observation bbbbfff6-7fa6-4e5c-b36a-0b8310de683c · outbound

This paper cites Kimi-k2.6.https://www.kimi.com/en/blog/kimi-k2-6, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Kimi-k2.6.https://www.kimi.com/en/blog/kimi-k2-6, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.368950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.368950Z digest=sha256:2f089d9475126078c387070b3acd6c43f79f7d6e3eb55a982f163ddfc2cf87d4

Observation 28d8c413-4466-4b37-b33e-aa99afdd43ba · outbound

This paper cites Introducing GPT-5.5.https://openai.com/index/introducing-gpt-5-5/, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Introducing GPT-5.5.https://openai.com/index/introducing-gpt-5-5/, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.455201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.455201Z digest=sha256:fbaf9d509db8b6b28e82924377f6975f528f28d5669722334826b8add7dc1bc5

Observation 926f82ef-149c-43eb-b510-5c63a8a3defc · outbound

This paper cites OpenClaw.https://github.com/openclaw/openclaw, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities OpenClaw.https://github.com/openclaw/openclaw, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.567579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.567579Z digest=sha256:392a807c9e7eb455d2e2bbde00edf438cc952316741dca65dcbc1ce7cf83ba5e

Observation cb0821aa-585c-40b1-96b2-b63bd253c2f1 · outbound

This paper cites Patil, Huanzhi Mao, Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Patil, Huanzhi Mao, Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.665658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.665658Z digest=sha256:e47941a39113d89d1b0dbb3297733206e67773baf601c9f20fdf9c728039e999

Observation 13666849-670e-4f99-bd67-eded6936b09e · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.774561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.774561Z digest=sha256:341ae46664527fe907759c5aee72e8940e71d0f301143c2debc0fc788f107ad3

Observation 2ec5f367-1567-4f62-8d14-1e4b669df13c · outbound

This paper cites Humanity's Last Exam.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Humanity's Last Exam

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.856099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.856099Z digest=sha256:a3a86743ad0413dbbf30bc871415f6aecb8508b5a020eca542bd6f7f891e5843

Observation ecab5476-e8e8-477f-abeb-72eff4d96610 · outbound

This paper cites Pinchbench: Real-world benchmarks for ai coding agents, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Pinchbench: Real-world benchmarks for ai coding agents, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:28.922495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:28.922495Z digest=sha256:b55f686e1d1632737174c764abcc02dd2e83057fb4cf40f45c2ff41121364bd4

Observation 260d6c22-6de8-4797-8d2a-2c3fafcc2707 · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Qwen3.5: Accelerating productivity with native multimodal agents, February 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.010111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.010111Z digest=sha256:51df5ac55bab1621268e19c77a2abd8014d7e97cf725111e1207338a1503bda0

Observation 84cbb913-4bfd-4424-b91a-7bdc7c8f31fb · outbound

This paper cites 2, 1, 4.1 10 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities 2, 1, 4.1 10 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.095446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.095446Z digest=sha256:88f7981477d219128c063bc7f07abb791c7c42a7eabee414ab61d982ae7aa5c3

Observation 66914381-83cf-4a6b-bb9e-772d72f0424c · outbound

This paper cites EvalScope: Evaluation framework for large models, 2024.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities EvalScope: Evaluation framework for large models, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.175292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.175292Z digest=sha256:4555696255d1782f93657be1c54096fd1d86486eee290520aecf568e89990a7b

Observation d7a09ceb-9aff-4413-8f89-761301debeae · outbound

This paper cites Terminal-bench: A benchmark for ai agents in terminal environments.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Terminal-bench: A benchmark for ai agents in terminal environments

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.244098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.244098Z digest=sha256:73aba8d0b0a31d93c70c3f43a76c5415accbfec99e836cadb30a8a750b84e0c0

Observation 2eb6d0aa-92c0-461d-9201-071e9e450411 · outbound

This paper cites Scicode: A research coding benchmark curated by scientists.Advances in Neural Information Processing Systems, 37:30624–30650, 2024.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Scicode: A research coding benchmark curated by scientists.Advances in Neural Information Processing Systems, 37:30624–30650, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.359750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.359750Z digest=sha256:1082c9b49258dd017f66b55b98d7bb23216c919e6636254b88f349840f7e27e2

Observation ff36f398-609b-4664-86b7-4fc02010694c · outbound

This paper cites Frontier- science: Evaluating ai’s ability to perform expert-level scientific tasks.arXiv preprint arXiv:2601.21165,.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Frontier- science: Evaluating ai’s ability to perform expert-level scientific tasks.arXiv preprint arXiv:2601.21165,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.466375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.466375Z digest=sha256:7d673cd547c70c61b620ed0766e83929380a01bc09ce5f08451588046a04c51e

Observation 76426af5-bab2-4688-a7b0-9e3c42294efa · outbound

This paper cites Openhands: An open platform for ai software developers as generalist agents.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Openhands: An open platform for ai software developers as generalist agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.570199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.570199Z digest=sha256:d133d3fd11ca61caf3018f5f1f3be7935de6e22c7b5f1956afe398252b2b44b1

Observation 925e1fb1-34ad-4886-a120-3bc1b06c90c1 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.673169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.673169Z digest=sha256:a2b11a5914a6f6581c18324533e17c6307219558294aebd6d2e1a0c6d902c3fc

Observation fc9af247-75da-40ef-8ce2-a2f447e9aefb · outbound

This paper cites Autogen: Enablingnext-genllmapplicationsviamulti-agentconversations.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Autogen: Enablingnext-genllmapplicationsviamulti-agentconversations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.820203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.820203Z digest=sha256:a600e2e2624d4691d16fac2135dde6dcf8732b306bd9b4c2c9e4a67354eda814

Observation 0b2f40a7-5177-4628-a744-fec252b6d472 · outbound

This paper cites Mitchell, and Yuanzhi Li.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Mitchell, and Yuanzhi Li

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:29.916914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:29.916914Z digest=sha256:dbec25e4bb33dfa72313983e53104a78cf570d5e98f0780f95077811e1e9833a

Observation 5a1ef50f-53d0-443b-86cc-7469c83d35f1 · outbound

This paper cites Agentgym: Evolving large language model-based agents across diverse environments, 2024.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Agentgym: Evolving large language model-based agents across diverse environments, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.000535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.000535Z digest=sha256:3ed27a89b9654792374a620d0a37e3cf16f4723343419f4b87cca45d9725d5c5

Observation 26cd8ed4-7d82-40ce-b02c-f3dc5927545e · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.094777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.094777Z digest=sha256:dc111f6fb4503f79e83e043839bd387e0f687e94f1abd93cea3337130a135433

Observation cc637492-3a41-4958-b1a6-694c775d4164 · outbound

This paper cites ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.208799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.208799Z digest=sha256:536104d2f2c16521cf025a6bc87b5326a83c745e468c95eb6ab8558de03d92de

Observation 0c33621c-975c-4203-8a9f-f86446e16545 · outbound

This paper cites Probing scientific general intelligence of llms with scientist-aligned workflows.arXiv preprint arXiv:2512.16969, 2025.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Probing scientific general intelligence of llms with scientist-aligned workflows.arXiv preprint arXiv:2512.16969, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.343249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.343249Z digest=sha256:8922e9a8eb366ccf3d4e02592cf3bfe8f93d219a784fafbd28144b3dd4c11519

Observation ea311c2f-4163-47fd-89e5-73c79f2ffb0a · outbound

This paper cites SWE-agent: Agent-computer interfaces enable automated software engineering.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities SWE-agent: Agent-computer interfaces enable automated software engineering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.483827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.483827Z digest=sha256:4d4e7b16395ea87956112b0d19fa9b081910912f907c263b1ce1080dcf165b7a

Observation 966e94a1-5a41-4de1-9250-2658777dbadd · outbound

This paper cites Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.579259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.579259Z digest=sha256:a2afec897f92f704abcca38f31d8b1721efa03b040c2bcbd3d96ac430e60001c

Observation c2602e56-b100-4cdf-b869-33d3361eba6a · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.672460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.672460Z digest=sha256:3472a8caaee5144c72f69a3f41555ca56c69e470ef4fb6cd82247938f5666245

Observation 8be112a3-3113-4d1a-a1a5-cf830076da6d · outbound

This paper cites Maslab: A unified and comprehensive codebase for llm-based multi-agent systems.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Maslab: A unified and comprehensive codebase for llm-based multi-agent systems

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.792653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.792653Z digest=sha256:ef9210fbb74597b1fee2253b0ac061f5d92fcee38f5d3008eccb9d1c3640f534

Observation 6f64c576-bd5a-4cd7-97a7-2df81a0ea21c · outbound

This paper cites Hle-verified: A systematic verification and structured revision of humanity’s last exam.arXiv preprint arXiv:2602.13964, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Hle-verified: A systematic verification and structured revision of humanity’s last exam.arXiv preprint arXiv:2602.13964, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:30.979644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:30.979644Z digest=sha256:c5e5fb9a3cf474a393b4577215ce8e0e8723ee2da5b758e2b03016ac10b485b7

Observation 0555dc90-65cd-45b4-8ee0-76f1a24b4115 · outbound

This paper cites BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:31.131799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:31.131799Z digest=sha256:5057bea668dec575420984e6dbe347df07e6938443dbc283cb314a04515203f1

Observation a261d788-53f4-44ef-8fe5-4c0af4caab2d · outbound

This paper cites MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T04:28:31.310957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:31.310957Z digest=sha256:cbf30651cbeae7b79660c0396f6f68fe975184d2225a77e2ec0fd20d9d9d54d8

Observation 1af9b981-42f3-4e70-b372-90d68eef9244 · outbound

This paper cites Intern-s1-pro: Scientific multimodal foundation model at trillion scale, 2026.

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Intern-s1-pro: Scientific multimodal foundation model at trillion scale, 2026

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-02T04:28:31.474196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:28:31.474196Z digest=sha256:a22842db1c1ed9e282770df79fcc593a1821258c0c80f112361e87b63bdcfb6c

Pith citing papers

No inbound Pith citation observations are available.