Pith. sign in

Paper Citation Record · LEDGER

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

As of 7 August 2026, this Paper Citation Record lists 100 of 101 outbound references and 0 inbound Pith citation observations for arXiv:2605.14133.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14133 v2

Coverage vector

measured 100 of 101 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T20:19:21.824216Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 101 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved4
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch48

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c506d781-5242-4333-ad29-2a91399bc01a · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents A-MEM: Agentic Memory for LLM Agents

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.232826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:00d64a42a025a9b3e856cef187a240cac6f88538713f6f90d17aeb167a516481

Observation 9d552512-599c-43ed-b6fa-b14fe3918dea · outbound

This paper cites Zep: A Temporal Knowledge Graph Architecture for Agent Memory.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.266224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:094291c2cc227273dada59a87591d81a603b579c72f3115a383ae4160217b253

Observation 6dadaf98-a18c-41cb-9423-f9c534de58ca · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.617096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:e56c7a976faaf23756efbeec6e50800198c9a329c3a0f49efbea743334016b2b

Observation b852cc17-53c5-45ea-b847-94f550b98c3a · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.615410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:4a9b83fa721989842bd164c2da1f536001bc9940d2b71889baab76c0ec8b9daf

Observation ffbfee9d-f7f3-4273-8c74-1ce8b8c3bbca · outbound

This paper cites author=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents author=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.619234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:6a2228d1636f2b61b39cc9b7b1e700d367322625f77b1d09de4b062380b0d432

Observation 98b5c3c9-90de-4059-9c5a-926d7ecce04c · outbound

This paper cites ClawEnvKit: Automatic Environment Generation for Claw-Like Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.268782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:4a1dff6f8b0d39de5fefa23d0102b0652b3d8d7b77142e694da94954a430f3c5

Observation 0a46bae0-9e2d-4c6e-9985-dcac2444c112 · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.205856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:e31a6dfb6641e6070ac5bed692b019d1566764f97ade13eba01b802c82ee83fb

Observation 6ea8b4e0-dd3e-471f-9a61-85c3ec0be6e3 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.298529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:d4255cc97cd04bcca8487421c0743bf4b99924abc9e419865eabd18392ff1304

Observation bca6a4ab-ed63-4795-8e11-88a06f7c86b8 · outbound

This paper cites an unresolved cited work.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Unresolved cited work

Reference 9

Resolution
parse uncertain
raw_fallback, observed 2026-05-20T20:23:43.571077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:e9fe3733f65a2862cf77229bc521b04b1b3d778897773e57a44b9c2da62863fd

Observation 22bdc956-75a4-4be0-96a9-6eba6d84e67b · outbound

This paper cites ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.305322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:269f2b8ff89a1a844efb181e930593b9eb9647972112cf3c25c4b3de71f5f09e

Observation 52cb93c9-9be5-468c-9956-a1a6de76fe6d · outbound

This paper cites ClawArena: Benchmarking AI Agents in Evolving Information Environments.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents ClawArena: Benchmarking AI Agents in Evolving Information Environments

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.252073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:5a1c693a54e66fa769b204713f029060a1371d47ed92308966aef04f63e2dcfc

Observation 2ab436f9-2e60-425c-8d33-3fc0ae587a5b · outbound

This paper cites Advances in neural information processing systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in neural information processing systems , volume=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.600675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:8ee7cb6d976e121d8be645b7ad3684972111297c29027fb30cbe6ff092ca7533

Observation b9857001-9f9c-4855-a558-7e3b46a98960 · outbound

This paper cites ACM transactions on intelligent systems and technology , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents ACM transactions on intelligent systems and technology , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.611326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:f9fbbbfb124c2e651061a1370fa46265fe5a3dcf90f8cb2c51ae409795834058

Observation d0f7884a-c7bf-4bb4-bc59-92c32473ec89 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.257814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:bb1f25c80c5c6967b4a39ace4d0aa390d25366b3521a54a1c825771e6c7facef

Observation 19f9bdda-d2e2-4ca3-a8ff-cd8a95522b96 · outbound

This paper cites Advances in neural information processing systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in neural information processing systems , volume=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.588923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:ae142e8891e6b662453c25ec7864ea06f994914ea4ff15f37558eff7ea0a2463

Observation 599ed05a-99de-4a04-b221-93ae81f370f7 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents The Twelfth International Conference on Learning Representations , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.592691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:9ff8fe05c4f5100265453607ebb71f4847617fa0444c29fc11b1b19c9282f078

Observation 57098780-5d3c-4871-b7d9-281a9bef4d20 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.584867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:2a1078dba357f511a164e6499dcf4381de1472a4022e7c4c26312942a8ccf92a

Observation b1e0d0b2-4ee6-4a1d-a586-c5fcec4ba5d2 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.285765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:995566a0c8a8539d6ba821c6e35fe0d441f0cbc690276bd7a4c984d3645b188c

Observation 849e6e08-40dd-4f34-b595-e7dbbb06b952 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.293321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:8e8ef060d63b7d07d2cc3647138c4abc64ddc68704e71e7524ed244489d8c603

Observation 4875af07-62ff-4316-b2bc-290c472ab716 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.260600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:ea8a0a78331a7a83430c490b577fa3c8f96287a9d14c933808012823884e6afb

Observation 680341bd-9a72-488d-ba5d-48c85062a6f3 · outbound

This paper cites Proceedings of the National Academy of Sciences , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the National Academy of Sciences , volume=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.582909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:9dcacb7515361feb7c1d6d211c60ee0ef1f3d332315d7b06c80561dc26ef2c7e

Observation e66d6e89-3311-4752-92ee-d61e00d58e49 · outbound

This paper cites 2026 , organization =.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2026 , organization =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.587127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:52899289e8bbafb4621866083723c544a45925df6270e2cdec0a23e1c52465c0

Observation 118a9a4d-f76f-411d-8306-82a721ce0e51 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , year=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.590580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:588d97d6943b78194b7c8170a85d0342a605ea6ab04a5bb1aa4cf9a1809b8bc7

Observation b6726224-d691-4041-a418-b413cfdf3889 · outbound

This paper cites Efficient Lifelong Learning with.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Efficient Lifelong Learning with

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.607726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:790b64edb47908746fa26d076b463eeca749c2739e97451c3e475acc8a567b4d

Observation 72d81367-6f0d-404e-ab86-e04a6be6f732 · outbound

This paper cites International Conference on Machine Learning , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents International Conference on Machine Learning , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.628374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:0794d8b5773377b4f2d8f3ef1d523c33f68ec598ea20bad291090b1e85f47479

Observation 12348fa6-adf2-4fb6-9276-cc8b6c8f9c71 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.219405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:079196224de75f853023d7890096f58b177d0324a70aa110fbb2ba40d582781f

Observation 80cac00d-ba5f-4267-94ff-f122e2b12672 · outbound

This paper cites International Conference on Machine Learning , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents International Conference on Machine Learning , year=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.574444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:2dd359969e7b0707e4be729a9a2a5d97f34ac13b0013e2303ed78bb1652a69a5

Observation daa4816c-5381-4dea-b699-d67064d6fac7 · outbound

This paper cites International Conference on Learning Representations , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents International Conference on Learning Representations , year=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.572776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:2e0a25a26fb1645110d46fb8be6add4adca3b42169ce61b557a8a8052125f3f0

Observation 72cdc3e9-4b3e-4afe-ab5a-d1242d278fdc · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , year=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.576140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:c763221226ebe7cc1b2dc13113bbecedf7a1110a8219f108696053068ada9486

Observation 99b41d65-ae1b-46bd-8666-2c9832c853f0 · outbound

This paper cites an unresolved cited work.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-20T20:23:43.577919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:88faf3b99cabfcc56276bfe0c3ec45682ebd51a06f6f93d632ca9b8818d1541d

Observation 1a422a7f-9bca-4147-bd62-267212708230 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Kimi K2.5: Visual Agentic Intelligence

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:23:43.288369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:78eb0dfab82ecfe4a3e30eaaf16c7756f5873446780d41d837508b4eca415c09

Observation e470b44a-a27a-4dd8-a77e-b05ed3ae569e · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.569366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:36d72659f2c878d82f153ee06660589c647da725f5c81275184aeda1db611341

Observation 93e6a1d2-d111-47bc-bbeb-7937f2d5bec5 · outbound

This paper cites The eleventh international conference on learning representations , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents The eleventh international conference on learning representations , year=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.579687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:03cd5b0017cf4713ce31372060f2c1c7373c81407b84a65151ca3a7465a2ebe0

Observation 60aa9139-216d-4e74-993b-2c095952c349 · outbound

This paper cites MemEvolve: Meta-Evolution of Agent Memory Systems.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents MemEvolve: Meta-Evolution of Agent Memory Systems

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.224817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:c34176553c48b67496447c7609723a82655f5cb83efc8990642ce7069d1dfbcb

Observation 030660b5-612f-4ee4-96be-096bd092e41c · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.199913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:3d8a38285000ac9375c17ac22039e1b3e06e7e711f76a4897dfbdf520714c7c7

Observation 5a331ca2-640f-4a46-b4d0-4a1f09a4559d · outbound

This paper cites Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.202908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:52930389eb2cf5aacd88bb21c5e2db9c7f2efd36c01d78a0a6ae5027a226beba

Observation c391b1c4-9c86-48b9-9792-9e83ce110b92 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents LoRA: Low-Rank Adaptation of Large Language Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.275938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:2f228286d46eb5f6c7019ae02011ff1af7444d3e1f36ceb339be284d8504b0b7

Observation 7ad20e6a-54a8-4071-800a-b47c1866e2e0 · outbound

This paper cites Agentic Reinforced Policy Optimization.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agentic Reinforced Policy Optimization

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.281072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:f8cf6d87cab69d752785ecb2b7fa0191a059376f59945e89eb7b033be7c62299

Observation 8205dc0c-86cf-4630-a061-98a061310350 · outbound

This paper cites Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:36.165646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:42fbb15cf32dba913ed8a33f57d232917b2aa1eac269e3187c40f68d4d6de513

Observation eba12bb7-94ab-4eb3-a90d-d3a23ea2ba79 · outbound

This paper cites Advances in neural information processing systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in neural information processing systems , volume=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.622899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:64a66dee02bf0aff50c490963deb81773305429760094fc9b2fffebd49f5d778

Observation 7875fc84-9951-4897-afe2-84d764bf24c7 · outbound

This paper cites Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.295748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:ed98d148e52b2d0a929127068f0c3d8276bff2919aa7e651a8e652bc4305b056

Observation 878f693c-75aa-41ca-8276-28e6bff1591b · outbound

This paper cites International conference on machine learning , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents International conference on machine learning , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.647243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:10242ece279b6eb6c6d6116b7b288bde5619659bc54a2d424915fe4aff732440

Observation 28066ca5-8486-4685-822c-4ca9c2fadb81 · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.565462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:6db9df966ec35ad0abec163e761aaaa2e52fbde2ec0439a803f3e2eb9cd046ad

Observation 15682fd4-f984-49ca-afbe-b41b76c83dfd · outbound

This paper cites Agent Workflow Memory.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agent Workflow Memory

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.303034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:645d80ad3f6d5837e9b89a9309ec161c33f7ea3ada151a9031ee8587f7f9e589

Observation e812a28a-df11-4ad8-84db-a08f174907fa · outbound

This paper cites G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:23:43.263414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:ae34d3e641f8ea4ae6982b43408cd7595f18a242442cad55ff30db67bd9e68ff

Observation 5f66ffa7-ff6a-4a72-a512-7a056785a20f · outbound

This paper cites Advances in neural information processing systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in neural information processing systems , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.567230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:6438d93a8b01ec5e0248eb1e73ef8a1a9e4764f6381fcb7a674fe155d1b1f97b

Observation 2eaeec4f-a6e7-4d62-addc-9ef1969c9cce · outbound

This paper cites A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.213760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:c0426494edcd3f8adbfbd7f0e5a08fb2b87c7bf756f5dfe72fca7ddd7db7ef80

Observation 45d3bdba-a0fa-48f3-a05a-49fa57729609 · outbound

This paper cites Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:23:43.194138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:89d7213ebe47e23c7843a69f3f711847762c14d48fe6229cb29d2102963cd2cd

Observation d14f72ef-54cb-4916-8734-16a0957f4d5a · outbound

This paper cites Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:23:43.254983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:d6c9d32be9f7713705c8aaa6466ba1fefb42186ee72b87a33a81734ba9842fbe

Observation 2de9dac4-8635-48b2-8ffd-f1451a045931 · outbound

This paper cites Neural networks , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Neural networks , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.645568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:a15d5f973da8af2e92182ae956a075fc3ac44695bede1e2d621634e5eb1aae5a

Observation c8f9b9b9-79e2-4aaf-ac5f-112e18ef8d2d · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.649194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:0a0a06608a75daf973c49b0876ab8c9d4c48b8fccb307f896c5d39f2130ac9f9

Observation 64b4efdd-ebd0-495f-acb6-e8cdf66fcbfe · outbound

This paper cites IEEE transactions on pattern analysis and machine intelligence , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents IEEE transactions on pattern analysis and machine intelligence , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.641628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:dcb5ade90baba7086a7e7a53aa5f20e4187af356b3bfe285fbf9a0470cd4034c

Observation 509e0e6e-f60c-4ce1-a9ba-11566927791b · outbound

This paper cites IEEE transactions on pattern analysis and machine intelligence , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents IEEE transactions on pattern analysis and machine intelligence , volume=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.639796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:8801c06aa5574cf6a5c7fc722926050d6574e6d88d23bb0f9c0f9cee8c96cc6d

Observation 2df3e228-70cb-4a80-8f44-3187ac8101fd · outbound

This paper cites On First-Order Meta-Learning Algorithms.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents On First-Order Meta-Learning Algorithms

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.283425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:a73b97d4fe3934c094c2fdfe6340fafbf684b1a93476a3d2dc3271a3f662997f

Observation c108e4e6-efd1-4716-97d5-eaabab27e103 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.235537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:5aa386cb605b8936d8b2b79693097edb85cc9bb417af915876b2dd6a36550725

Observation 1176fd6b-9a16-41f8-9334-7257b7525361 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.310355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:a8c65aa30cf29d508c6fec083b04508f5ab3aec1faffe0716f89ce22a7ddad3a

Observation 0395589b-2b48-4624-97e7-66e7201ca3f8 · outbound

This paper cites International conference on machine learning , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents International conference on machine learning , pages=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.637699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:157b6f47c6bcd1ccfcb52ebf0a055103c6680f9633e4254144006df11a34c594

Observation 76eed3a5-7e5a-4cb8-8725-a308cc4a8fff · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.643491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:63612db76c37ef8bdbd8cc0a325110e5ac81bcbbc8d91b6ef86aabd742db4181

Observation 656a3345-2b1a-44a7-97c9-33cd94bfb0d4 · outbound

This paper cites Group Sequence Policy Optimization.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Group Sequence Policy Optimization

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.182269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:a235f707e1e923ada7824d74c4ec3e11d5a47e0b47f3483fbbca3b08835dc7f3

Observation 5dd3e6b2-8875-4803-9d4f-1cb3b55e528f · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.273474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:dea53ace382c4ac8b851170364501bfe0b97d1917a647956784b900e0743c944

Observation f6ffcddd-0919-425d-a233-6a54e2861649 · outbound

This paper cites SimpleMem: Efficient Lifelong Memory for LLM Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents SimpleMem: Efficient Lifelong Memory for LLM Agents

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:10:45.412758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:8d79019c00d9ca5b8c36133323c1c497de968874342a690286cf720fa730a8f7

Observation c480a274-e797-43b3-a8ca-7a35bca8a40a · outbound

This paper cites Testing Language Model Agents Safely in the Wild.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Testing Language Model Agents Safely in the Wild

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:23:43.230284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:cf452ed8708eb4375dd299cb84b5e6b8aca89dd3b4429fb88b3dffa22489ff8d

Observation f0aa1872-a5ff-4330-88c7-3eee64bbcc57 · outbound

This paper cites an unresolved cited work.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-20T20:23:43.635925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:d3d7649350000507e72fe5bccc1bc7b9fefce1edc1c7efea969e60518d2803f9

Observation e7a1938f-aca4-4f4f-bf34-0c20b1df9780 · outbound

This paper cites AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:23:43.307883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:64eb5d3438729c062a569bf3724ae95576dc485baf501fd1acbbd6788b711e6b

Observation 80e583a1-106c-492b-b2ff-080ba6a5b621 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.300937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:a219ffebcf7b12df61ac801646542addc4a4ec046eb76140567d0dbade7a1757

Observation 2ea9022d-7cf4-4d9f-bce8-bcdd5975f5ad · outbound

This paper cites Memory in the Age of AI Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Memory in the Age of AI Agents

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.185292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:8e16d7f3e282441ab625fa4864b5a42faf33dcfbf972392668db603a7c162a47

Observation 570e4bb4-7db3-4ad4-a78a-88c6062da656 · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.290691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:b577f7d72d5ded0c639e962f4c6f99ef7f79f76ad812bf19e7fb0711a2284cc8

Observation ccb9b716-cdf5-4de6-ae34-619a021625fe · outbound

This paper cites Agent kb: Leveraging cross-domain experience for agentic problem solving.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agent kb: Leveraging cross-domain experience for agentic problem solving

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:23:43.242300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:564a8f8b498463fec28558fc35a62e95e2d0feae2216a0101c78afc91f8a0444

Observation 8b8471d5-1755-4bce-992d-417782c54952 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.630219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:d780b253f2b86128d4eabd08e7d57afdf166871e53a7b48809c026cf9476db34

Observation 589e422e-f02f-471d-9d21-4ddbe6f33688 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.632124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:0cc66a85f04abe51539181335e2f080f1a88561eccb3b98b27e58a26283da734

Observation 2989f9d4-7382-4301-891c-091ae6f1f227 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.191304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:ff48e0f3f899bbe890f37a77a98b98f8d21b6bde80b1fd1cf9340031d27d9092

Observation 5d6359df-0977-4115-a374-4cb799265576 · outbound

This paper cites EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.312709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:0d1b8611433b21bc403b3d4df41d675ec2556309d13e521ca1b56fef2914afd7

Observation 1c0bcbfc-01d5-4e1e-8ab1-378d8d9d23e7 · outbound

This paper cites Memp: Exploring Agent Procedural Memory.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Memp: Exploring Agent Procedural Memory

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.197029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:076626ad77fd66366a20e3127bbe4db883a3d9fb0874cbfd62408920c3607d9d

Observation d1141c40-cffb-498c-9546-d3f32f126f6d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.188388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:68b68029af4a640427d3a36f0b026913911b24674c0a19f5a259aea7315d2279

Observation a1ccfae9-bab3-4898-93f1-3289fc5327f4 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.211084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:ad268ab4e2eb21727c32335db479343db4ea255c5f9ce31c0feb34a211ea19ea

Observation 0896cfae-2f57-47b8-97e2-77b05a8387e9 · outbound

This paper cites 2025 , author=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2025 , author=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.624591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:2abfdd8acba30460731b2223e73f9ab5fe31c9fce136a63787d4ada27bf0b240

Observation 66ec27b4-e54b-4a71-941d-e67d1dfd3472 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:3044dc1444f12610e65d6d1ed1c5b60af8e506312cbcfbb8c27f0c59516e7e04

Observation 0954337a-a67d-4eea-97f8-edae9dc98a5c · outbound

This paper cites 2024 , author=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2024 , author=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.620875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:661214ec3d5f81a3e06971e505289e010a0a3bc463b91fde5714c534207adc17

Observation 193767dd-1662-4075-929c-f4675355fe6f · outbound

This paper cites 2025 , author=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2025 , author=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.563601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:db2ab704e523f9c6c7c3e07df967a1c56e8c9bf2dc7220765b222bb99091bbba

Observation 5e10ad2d-a3ec-40a2-8227-d31963b42d11 · outbound

This paper cites 2025 , author=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2025 , author=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.581238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:6e15676c28d6d89e25d1bd9cf6fdc7702ac33c59ea587809128505b34217c58a

Observation b503fd6e-69a1-43d2-b184-fa67ac06084a · outbound

This paper cites 2024 , author=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2024 , author=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.613147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:f0dd23c7184fbb6a01279e76b2a238eec818adcbdedc48b9f27062c8a6ff8326

Observation 078254ff-ecf8-4552-b090-1c1fc1e866db · outbound

This paper cites Tongyi DeepResearch Technical Report.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Tongyi DeepResearch Technical Report

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.271148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:255f83e6566676f7a43d7c3f96213ed1c83c279cabc5d91c1d50a0321b7ebdbd

Observation 8e158ed5-693e-4b2c-a033-26d60cf63a92 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.278367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:b06396fa2b739f3a7225b7dbc7b4b0967f679b654bcc37c7eafc74b5196dd81d

Observation 093e8c82-8f7e-418b-9b27-d9c567c06f94 · outbound

This paper cites Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.238254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:de3abdd0e6649d735c528d571ad55e08e3e8798c205d821f96edd481866bd21c

Observation 0c003b22-810e-46b5-85f5-02c42d3b06d0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.626388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:4bab944a13ba72cccf3fc09385cb8e81668b9a7951720109b3c2163da1ad22b9

Observation eb5338d2-f0a9-41a7-bb42-4b6fad8d1578 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proximal Policy Optimization Algorithms

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.208503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:a293f8a1074cc8092e429a06033f8088fa3766f414e62a0fbb1d9b047a326e3d

Observation 03da1012-3696-4a93-90a2-e6fa14d1f4d9 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents FireAct: Toward Language Agent Fine-tuning

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:23:43.246128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:3767ed97904f6243a8477cab710c731326d520b46941abcea4c0586a96202d3a

Observation b6e3a6a4-92ad-4296-bd33-8f6fb192b888 · outbound

This paper cites Qwen Technical Report.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Qwen Technical Report

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.227288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:31bfd396ebf2f5e48d75496591005727a748c29ce9b2d0251b4f8a28b298d440

Observation ef44d215-85a0-4fbc-95f9-19c312d1161a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.605880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:2fa9c0a6ceeb2950704af21f28b6cfd7f6e5d0c3197487aa7c652f8c81772429

Observation bd352baf-f58a-421a-a73e-8fe55beb5764 · outbound

This paper cites an unresolved cited work.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-05-20T20:23:43.561919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:c15147d8b3774ae1160fb28e46407738e7d151cad1a8b591d178a4172990a184

Observation 3f53375a-8356-48f2-b947-a85ca6ede89e · outbound

This paper cites 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.602470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:26d2b76e115a337db599ea91ce63831be097127c4e765f12e1306cb6dc707641

Observation 7b6f92bc-4103-4c0e-99b1-1c25a22b2b70 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.604189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:c74863e5acfd075e746f32011e9d96238418a295e7a16febbf5a2d906c5b8e45

Observation 201985b5-69db-472d-8c9b-d4c3abd33e7a · outbound

This paper cites Artificial intelligence , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Artificial intelligence , volume=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.597041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:937597e9ed468cf06a404b21c0e6d910cbc34c6c8bc7a2fe0d67cccfd3bd8b78

Observation 2b626a41-123d-40dc-be3b-156607899dfb · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.560119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:3579089db639433d4b35ed9be88fab0e2a959fa4890960a700fb2ee6ff0e3801

Observation 7268c921-ab63-407b-86fe-f30b413d5e4d · outbound

This paper cites an unresolved cited work.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-20T20:23:43.599011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:37bc16e0b9c319c3d305d1e0eb0463d4c42b96d222aa20686efb06dea4730cd4

Observation ec6519f3-ab46-4151-aca2-fd4c2f6d94e7 · outbound

This paper cites First Conference on Language Modeling , year=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents First Conference on Language Modeling , year=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.558320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:fd26009885676a68881a2c376f30798be80965aaf92d24ecfd8234057501eff2

Observation d80640d4-a9b6-482a-83d1-d9abae162931 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.633974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:80f3535bc26b512c6a5ed4e4114e8c1bc55d20c364a5640903387fe41502bab0

Observation 512679d3-2900-45f1-9bb9-b348ed991bbf · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Advances in Neural Information Processing Systems , volume=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:23:43.594966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:aa49337be41cf034268ec46e6ebfe3f2d5cf4b92ab85b6b425cdc95a36de08c1

Observation dead71bf-aece-4abb-8e00-e3923bc08082 · outbound

This paper cites MIRIX: Multi-Agent Memory System for LLM-Based Agents.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents MIRIX: Multi-Agent Memory System for LLM-Based Agents

Reference 99

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.222323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:b2d26133f2c2517ea73bf728c153bf0b374ffd581a99bb152f683cba8918a200

Observation 0959e00a-2a79-4d39-bbed-43218083d110 · outbound

This paper cites Agentic Reasoning for Large Language Models.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Agentic Reasoning for Large Language Models

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.249036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:05e39cdfb867d1ad741753f374d2552f2b2ac9fd495563b465de31fed2bed362

Pith citing papers

No inbound Pith citation observations are available.