Pith. sign in

Paper Citation Record · LEDGER

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

As of 6 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2604.23781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.23781 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T06:36:52.202847Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:36:15.758892Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:42.725645Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact12
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c745d5cd-947d-45b1-b23e-3be81c733c16 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.841344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:54706dbd7073f34744d603b69d7260e9fa838e01142f4b3c59fc606c07a5a0cb

Observation 43f1aa7d-cb1e-489a-9b73-05e2b84d18a4 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:48:05.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:e68083345b1bc30442748c2bad7c39efbdccbf763d70a4bdbdf0254d3de94561

Observation 5408acbf-6df5-4556-a7e8-0321276c6aab · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:31.128039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:b702479f9e5f6d73f48a475b7ab301efa1fe1266a53a5546ff0a6be470b3e4f4

Observation c3412cfb-7adf-412f-bf95-825aa7babade · outbound

This paper cites Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Visualwebarena: Evaluating multimodal agents on realistic visual web tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.480542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:9b69717c3e4b8a2b3845a226ba438433fc6bdf46673de0a7aa23426cd728217c

Observation 1396d387-47a3-4eb3-b6a0-a488517f3dad · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.472468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:06b9231f9902c2b0aeba9a4b4d729d40eea8abb8961530195f82a17dc84c8a7e

Observation e170de81-13c1-4c27-8264-99aa5a626e3e · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.918407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:0ef6915829f56690f610335368b70d2868cea6e632cb8ef6e5261b55d2f76e80

Observation a4525fe0-769f-415d-9c41-d2dbc0339314 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.854057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:2ff8762c302462ac8823f769ade0e48d669acb3b83a30aad187cd6566d0e0537

Observation d9df3b27-bbb1-4b22-a4f1-1df1e37fed2a · outbound

This paper cites Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:10.898917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:754a6862e2818173768b5e4ba66333b40df3d45bc34faecb6cd4c41fbac29158

Observation 889a55d8-3cc0-4ed2-a100-c62caa90a053 · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.795764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:eca6aad4b4adf90fe005c37c23bf8a2ed4343cc497cc4da5bf62d1b3e4bcd071

Observation ec1f690c-0094-4a98-a97a-41eb93f7fb8a · outbound

This paper cites ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.876875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:405e758b9b4a1ef5a4ef68925f7a963a93b6516382a6bbb9e2273b5af59c6afa

Observation 77bd5cdf-05b7-48af-a50d-d0e1c6ece492 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.484838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:3e4e131c65e35f53e3152d6b20c75ed84e349825ea1c0c003e09739d0a46f151

Observation df550c02-dd52-459a-aed6-c25951f65292 · outbound

This paper cites MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:10.779830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:5d1bff6c4c8621d2da689a4aa5ca98c9c4c82ace5356c6871a28ef4c130f3a76

Observation 2b1ea6ea-9559-4c50-b4e6-2ca1d7d62910 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.882978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:aee158212115c83f40c8b8b8edbd42fc5440b21914f20e5d401accd4f24aa4bb

Observation 0ca1a61a-1aae-4786-ae12-899f465d57d1 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents AgentBench: Evaluating LLMs as Agents

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.925570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:74376ae68cb8d7b87ab4db7caf3b9148675c40bf45a77a60cf01e9e4f6d1fe13

Observation 53ed90ca-aa7b-4bdc-a015-497c18def0b8 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Gaia: a benchmark for general ai assistants

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.476310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:0464acbc35518b2918656dbf4773a21b0cab06c6656b30b9872259a05480f67c

Observation b21eccad-d261-4c1d-aead-73eb8ef2d31e · outbound

This paper cites ClawArena: Benchmarking AI Agents in Evolving Information Environments.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents ClawArena: Benchmarking AI Agents in Evolving Information Environments

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.906594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:52d35f38013c61620d552c8994859d67df36108e533a43c92b7d16245110f1e6

Observation 22252659-e486-4d29-934b-cc202fd46ae7 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.488632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:dbce1473b740a674c12f83c71e49593ce118a28005204d26722e7193929f08b3

Observation 4aaa25e5-9a7c-4c13-a291-38dbfdc45c9f · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversations.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Autogen: Enabling next-gen llm applications via multi-agent conversations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.465276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:85ee5d6efb435bc8163e45190d47f253a16bdebc7d685ed308d093dc18337593

Observation 79e1908f-7616-4582-af63-ef7707d9240f · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.806436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:d9cad9d93447dcdc205ae52da3f5fe5aa59803e27886c47423ff0dde02d4200c

Observation 3f9702e2-a8e0-442e-94f1-e1f5dff331f7 · outbound

This paper cites Day 1 / Day 2 / Day 3.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Day 1 / Day 2 / Day 3

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T18:38:14.468512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:ad990b1d4b3a1ba7af36f9a2a9c7120f15881da61ea889650340b2da1c96bff8

Pith citing papers

Observation 8e670151-85ff-468d-a7c7-b5031919b87b · inbound

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild cites this paper.

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:23:28.378969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:14:07.487577Z digest=sha256:684de249fa0988546a27dbe98c8a6c21934da17c9d351e704152c866ea1aaa07

Observation 903180f3-6562-4001-8539-c9c014c8b650 · inbound

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks cites this paper.

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:57:47.643579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:39:51.630229Z digest=sha256:34d064d152ea4fa3844c21d6a7402e6a9031120b628d56b65e9cc6383a7099d9

Observation 6b12aacc-012d-4920-bf42-26f09ecb7724 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.094932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:8286d6c05b36009b1ed379314e92155c0369fdb32d25a79ff2b076be8c73e3a2

Observation b3611f80-5451-40c1-96e3-10bec46c0c40 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.757410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:dfe5a3d46cd326f04a1ccaa6b860b56a99f6a0289821fa9a31b9c37bf5b18540

Observation eade2889-5875-4bbc-859a-1e97167440ab · inbound

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks cites this paper.

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:41.270682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T01:39:57.741032Z digest=sha256:6406376e1b3ccfa7cec55b021e4fdf50d3df65be57a743b11979d20c12d8b5e7

Observation 35d0ddb8-f4d3-4ac4-920f-1a718e7be0cb · inbound

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios cites this paper.

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:36:15.758892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:36:15.758892Z digest=sha256:08e864c6794201f5c1d668d8707aeb2c718011fbd6a934b3e38c8970493f8773