Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:36:02.687612Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2607.20911.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:36:02.687612Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b6516c9-2241-4dd7-ad8d-9024af362ed3 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650f846d-4b37-4b51-a235-a9645a146205 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Introducing SWE-bench verified
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c6ba681-e33a-4809-b64f-156ad88b51bc · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa691d00-fcf7-46ec-874d-4f770eddab61 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb23a183-e6f0-4462-ac72-bd3fb0a7bb06 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction How we compare model quality in Cursor
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 406da3f5-0052-4ebe-8d4f-43c552b4f8c0 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Harbor: A framework for evaluating and optimizing agents and models in container environments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4d42de-a987-4681-a33d-05815eb392e9 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Commit0: Library Generation from Scratch
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b43127-8028-46a7-9eed-18e299593ff0 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af780340-34d9-4dc1-bd71-02dea49c4db1 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9253df45-1629-43e1-81a3-21d46acd79e2 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Aider polyglot benchmark
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 54fbd8e2-e65f-433b-b9f8-241fc400176d · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Terminal-bench: A benchmark for ai agents in terminal environ- ments
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d9c0d408-5be4-42c8-8c2e-c806aedab33a · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Vibe Code Bench: Evaluating AI models on end-to-end web application development
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9c2460f-fe8b-4e76-94a8-78aefe0da60c · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486503ac-cb80-48c5-997b-d6488756e01f · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 220a86e3-00de-491c-af06-d2f0da5a29a7 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48816bfd-db39-4be8-9079-d9dd24e348bc · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7354a6cc-f6ba-4479-98e6-f46bec310083 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5bce11-45c1-4c3c-bf0e-d09550334045 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757db498-8eef-4455-a0c0-234e7c9f69af · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction SpreadsheetBench 2: Evaluating agents on end-to-end business spreadsheet workflows, 2026
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 365145fe-aa15-4a18-8967-cceae3af76e2 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89753231-60b7-4d4e-bbb6-a4eec87cab5c · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd19d687-a554-4838-8eb8-40c39eb67ebb · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b9270f-cf53-42fd-96ab-6c1e058371f5 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7a8392e-6257-4529-abeb-a9af6c27a4a3 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291bfb6a-69c1-4a23-a44c-26698bfae635 · outbound
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CyberSecEval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.