Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.06008.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ba6f5ccb-7b29-4481-aebd-b213a5031950 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Mle-bench: Evaluating machine learning agents on machine learning engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9cadee4b-edbb-4b17-9701-12f9ce70257e · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0dab83c0-d197-4afe-b396-dfe62bb4f293 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ca1e77b6-5512-4903-95dd-b814270c0cf5 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Measuring Massive Multitask Language Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T01:08:06.256034+00:00.
Observation 0d3b979c-dc46-4c93-8180-60abeb157686 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Maps: A multilingual benchmark for agent performance and security
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 77e47f8a-b7d3-4513-a9bd-9127697abfd5 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0f85dd62-8a26-47f3-bc12-5ac9d935ec30 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 71407461-773a-4b33-add0-da937d2de379 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c58c1416-b373-45ee-82f0-70fd990fc1fa · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b05fc453-493b-4738-8cfc-449cbc814a46 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Agentbench: Evaluating llms as agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b3611f80-5451-40c1-96e3-10bec46c0c40 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e42fabf6-20f8-4ef8-8515-51907686bed7 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Gaia: a benchmark for general ai assistants
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 59d84f9f-5299-44ed-8dbb-bcf40d61c9ee · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Language Models are Multilingual Chain-of-Thought Reasoners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f27d869b-e975-4c8d-93d0-e06cd117d0da · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Coffeebench: Benchmarking long-horizon llm agents in heterogeneous multi-agent economies.arXiv preprint arXiv:2606.16613,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f397c96c-d193-4806-aac6-ff5a0bebca5a · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Mirage-bench: Automatic multilingual benchmark arena for retrieval-augmented generation systems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3d6e37f4-3f48-4ee5-9232-60d5934d89f5 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ac55d9ce-f3a7-4083-b841-78885a022dff · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ccd35c02-7096-43ef-8624-5d94d1e5a477 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4144ca50-3033-4005-831b-e75ca6ee40f0 · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 76d7ff12-5a3a-43aa-883a-6144cc88164d · outbound
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents Webarena: A realistic web environment for building autonomous agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
No inbound Pith citation observations are available.