Pith. sign in

Paper Citation Record · LEDGER

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2607.20911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20911 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:36:02.687612Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b6516c9-2241-4dd7-ad8d-9024af362ed3 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.595063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.595063Z digest=sha256:ad6ae0928475a4fae967d559fecc1a6acc17e7d53c87c4fe8ae7d99bffd02e12

Observation 650f846d-4b37-4b51-a235-a9645a146205 · outbound

This paper cites Introducing SWE-bench verified.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Introducing SWE-bench verified

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:03.140598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T15:36:02.599654Z digest=sha256:33618d163b8ec6f55bb3a6b729bd884e39afe4e7c441e54c60ff0a4e79d14946

Observation 6c6ba681-e33a-4809-b64f-156ad88b51bc · outbound

This paper cites Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.603024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.603024Z digest=sha256:ebe1451543321cf85fc1d615e249d87407593f149147d18ccc96d75651d8564a

Observation fa691d00-fcf7-46ec-874d-4f770eddab61 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.607537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.607537Z digest=sha256:ed57488f5e5477f736b4c37220016f9c51ee38d18b34f7b11f0e1c8a83e2e17b

Observation bb23a183-e6f0-4462-ac72-bd3fb0a7bb06 · outbound

This paper cites How we compare model quality in Cursor.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction How we compare model quality in Cursor

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:03.129434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T15:36:02.611617Z digest=sha256:0a4099c63fc05df5fa83a418e38d03642e650f1d315377f4b36aeb11a47eacad

Observation 406da3f5-0052-4ebe-8d4f-43c552b4f8c0 · outbound

This paper cites Harbor: A framework for evaluating and optimizing agents and models in container environments.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Harbor: A framework for evaluating and optimizing agents and models in container environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.615447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.615447Z digest=sha256:b0f71fb3ee8a2948b65855d92b57c2915af70fa6e6691868969ff86012ea2ee9

Observation 1c4d42de-a987-4681-a33d-05815eb392e9 · outbound

This paper cites Commit0: Library Generation from Scratch.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Commit0: Library Generation from Scratch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.619120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.619120Z digest=sha256:f8414d8e2abc72a44fa9f7d01cc07eb3c54c0e6c61c48f4f5cd2f2749ac89dd1

Observation 53b43127-8028-46a7-9eed-18e299593ff0 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.623468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.623468Z digest=sha256:30bb5cf170ce4f44ef7afd6a1274b23c841483faa447ad71cc2a90d04f2c76b1

Observation af780340-34d9-4dc1-bd71-02dea49c4db1 · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.627067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.627067Z digest=sha256:ff2fb1e94585005854c988d6dbd11bc5a8f40e062706ee73a3dd5cbcc5a0e673

Observation 9253df45-1629-43e1-81a3-21d46acd79e2 · outbound

This paper cites Aider polyglot benchmark.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Aider polyglot benchmark

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:03.117955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T15:36:02.630625Z digest=sha256:aeb80f31bfe245f4ee153e1e2c93f518d6c8ce1ea94e562cc57a9e256440178e

Observation 54fbd8e2-e65f-433b-b9f8-241fc400176d · outbound

This paper cites Terminal-bench: A benchmark for ai agents in terminal environ- ments.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Terminal-bench: A benchmark for ai agents in terminal environ- ments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:03.106693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T15:36:02.634631Z digest=sha256:76786436c2d787a85fdd9f3f0cadb01d9a7fa0f03fff3183649ec0da316520bf

Observation d9c0d408-5be4-42c8-8c2e-c806aedab33a · outbound

This paper cites Vibe Code Bench: Evaluating AI models on end-to-end web application development.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Vibe Code Bench: Evaluating AI models on end-to-end web application development

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.638718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.638718Z digest=sha256:84ed82dacb8020c6df590666397c5c0ba3d0123d2bc1003b654ce412cc24278f

Observation f9c2460f-fe8b-4e76-94a8-78aefe0da60c · outbound

This paper cites an unresolved cited work.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.642322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.642322Z digest=sha256:196057fce7244118442563b5fb9ab42f590444587b2bcdf0db163f88f8aa944a

Observation 486503ac-cb80-48c5-997b-d6488756e01f · outbound

This paper cites FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.645495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.645495Z digest=sha256:c0eeea9454dac6e67a4da0ed036f1764e5b73da4e71f193aeaeb1adf52fe2d85

Observation 220a86e3-00de-491c-af06-d2f0da5a29a7 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.648855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.648855Z digest=sha256:ed6775d8cec0213ac4b752252f407365d1c41a6cc315c125f89794d68bc28733

Observation 48816bfd-db39-4be8-9079-d9dd24e348bc · outbound

This paper cites Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.653126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.653126Z digest=sha256:59d90f25de02dc7e5dadc5e53cc7416fbeaed776e22cfd540b12c7cb0c23bf4b

Observation 7354a6cc-f6ba-4479-98e6-f46bec310083 · outbound

This paper cites ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.657755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.657755Z digest=sha256:981e6b522e2fd0b96b35c3c44c01be301d47cd15fb441a606d3b3f3cdb8c1bd7

Observation ab5bce11-45c1-4c3c-bf0e-d09550334045 · outbound

This paper cites OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.661415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.661415Z digest=sha256:6dafa68272d69a40065c5d2bf03185a9f3f4e8e12fb34eb4d036a66f97f5f2fb

Observation 757db498-8eef-4455-a0c0-234e7c9f69af · outbound

This paper cites SpreadsheetBench 2: Evaluating agents on end-to-end business spreadsheet workflows, 2026.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction SpreadsheetBench 2: Evaluating agents on end-to-end business spreadsheet workflows, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:03.095635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T15:36:02.665405Z digest=sha256:70fd02204e0651dbb7e73de4a3be8c2e5f535055512e74c5eb3c003a1eaff69a

Observation 365145fe-aa15-4a18-8967-cceae3af76e2 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.669356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.669356Z digest=sha256:70989d4410bbaf10c0c352a3a674e3d7264821229457b20775a619d38959dc2f

Observation 89753231-60b7-4d4e-bbb6-a4eec87cab5c · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.672780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.672780Z digest=sha256:229b4c8997712a6d2008962758bf6c07bf70eaa90db9cf182f09550985a28af3

Observation fd19d687-a554-4838-8eb8-40c39eb67ebb · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.676898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.676898Z digest=sha256:a52bd4d409c74e9d4f502f722326f920faa6674e1f944403f91fc340f4570a23

Observation 27b9270f-cf53-42fd-96ab-6c1e058371f5 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.680411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.680411Z digest=sha256:b0a1f26b7cd2c1af3b6ed1c6e5de19bae178d98cc919272bc86f92a18a34352f

Observation e7a8392e-6257-4529-abeb-a9af6c27a4a3 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.684347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.684347Z digest=sha256:31e81ba0719b28a9ba65dd4a0100039cadea85fcd982098b7c4b713b9cda7e15

Observation 291bfb6a-69c1-4a23-a44c-26698bfae635 · outbound

This paper cites CyberSecEval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models, 2024.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CyberSecEval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:03.084462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T15:36:02.687612Z digest=sha256:4ac33cd27a6ec005dc6086d7bd4637f1ea75823e9be0229a39e7ccbe9765fbc9

Pith citing papers

No inbound Pith citation observations are available.