Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T06:36:52.202847Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2604.23781.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T06:36:52.202847Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.758892Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:37:42.725645Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c745d5cd-947d-45b1-b23e-3be81c733c16 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43f1aa7d-cb1e-489a-9b73-05e2b84d18a4 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5408acbf-6df5-4556-a7e8-0321276c6aab · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3412cfb-7adf-412f-bf95-825aa7babade · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1396d387-47a3-4eb3-b6a0-a488517f3dad · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e170de81-13c1-4c27-8264-99aa5a626e3e · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4525fe0-769f-415d-9c41-d2dbc0339314 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9df3b27-bbb1-4b22-a4f1-1df1e37fed2a · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 889a55d8-3cc0-4ed2-a100-c62caa90a053 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec1f690c-0094-4a98-a97a-41eb93f7fb8a · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 77bd5cdf-05b7-48af-a50d-d0e1c6ece492 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df550c02-dd52-459a-aed6-c25951f65292 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b1ea6ea-9559-4c50-b4e6-2ca1d7d62910 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ca1a61a-1aae-4786-ae12-899f465d57d1 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents AgentBench: Evaluating LLMs as Agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53ed90ca-aa7b-4bdc-a015-497c18def0b8 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Gaia: a benchmark for general ai assistants
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b21eccad-d261-4c1d-aead-73eb8ef2d31e · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents ClawArena: Benchmarking AI Agents in Evolving Information Environments
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22252659-e486-4d29-934b-cc202fd46ae7 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4aaa25e5-9a7c-4c13-a291-38dbfdc45c9f · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Autogen: Enabling next-gen llm applications via multi-agent conversations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79e1908f-7616-4582-af63-ef7707d9240f · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f9702e2-a8e0-442e-94f1-e1f5dff331f7 · outbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Day 1 / Day 2 / Day 3
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e670151-85ff-468d-a7c7-b5031919b87b · inbound
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 903180f3-6562-4001-8539-c9c014c8b650 · inbound
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b12aacc-012d-4920-bf42-26f09ecb7724 · inbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b3611f80-5451-40c1-96e3-10bec46c0c40 · inbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eade2889-5875-4bbc-859a-1e97167440ab · inbound
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35d0ddb8-f4d3-4ac4-920f-1a718e7be0cb · inbound
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.