Pith. sign in

Paper Citation Record · LEDGER

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces

As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2606.01317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01317 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:44:18.994680Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved21
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0a007ea-a870-4cc8-8a61-7d1dd199d7fc · outbound

This paper cites Red Teaming Language Models with Language Models.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Red Teaming Language Models with Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:36:15.071498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:24d1e75039880082be98ae6b82add7f5a2b3006a3a66c73dd1bc86d60c229051

Observation 4c814a72-2663-40e0-ab91-eb609f048e7b · outbound

This paper cites Proceedings of the 16th.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Proceedings of the 16th

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:4c591e9c9b81a0db883ef40e8f1ab21b3d5507b06a893884bd0d7f5e647a5e08

Observation 7b90fe33-20b8-4630-90f4-95fefb774868 · outbound

This paper cites Extracting Training Data from Large Language Models , journal =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Extracting Training Data from Large Language Models , journal =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:f0d68cea6ec8f98b8d7eed202e90cb0598b402e5e11a68aa42f3ea184d0edbfe

Observation d7459563-31e4-4fdc-9422-086420ed3837 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces AgentBench: Evaluating LLMs as Agents

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:36:15.069087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:e67960e87dd4c1cd71db34a80877d59b61f61b689170ed42b63d89100751b678

Observation 9f57470a-be26-4ef8-97cd-cbd67dd9186b · outbound

This paper cites an unresolved cited work.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:45d90d9e4004038811b45a2be4a30ed684202553ab9f89d065d4a4a0e8bc35f0

Observation 1e53ebb4-1dec-42d3-bd00-6ed89d372d3c · outbound

This paper cites Forsyth and Dan Hendrycks , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Forsyth and Dan Hendrycks , title =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:cad2146a67326e8df3b630a7cd9b20a49b62115839fe00c281de9a8b3e68d4e4

Observation 3f34189b-fc60-43d2-b1bb-6419cf2d3b57 · outbound

This paper cites Chasing Shadows: Pitfalls in.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Chasing Shadows: Pitfalls in

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:183c17677cd3146185d939d1f450645e0b021231e0098b733319928174cfe457

Observation a2f23be2-0ab0-49d4-812b-3b7bf5847f92 · outbound

This paper cites Findings of the Association for Computational Linguistics (.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Findings of the Association for Computational Linguistics (

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:3d1c06806e5f8b4b8af19cf22efb9d6ea75a5b4dd4e53eff49e087d26e67b8f6

Observation 8c3ecaf5-4212-4eee-81a3-8fb7f4900011 · outbound

This paper cites Zico Kolter and Matt Fredrikson and Yarin Gal and Xander Davies , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Zico Kolter and Matt Fredrikson and Yarin Gal and Xander Davies , title =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:fc7138c077d14ece9149b0224ee5df55ad14d90f5fc7799de473403d9c68ccc4

Observation 4a39f42e-d611-4a64-b9ae-ca9df611ec69 · outbound

This paper cites Advances in Neural Information Processing Systems (.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Advances in Neural Information Processing Systems (

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:22c5e3706e3a2fb5236b383e92dc57280fbf3fd378238911d21b2a6b3a5492ce

Observation 13bcc762-ad98-4574-9b08-55292b7454a4 · outbound

This paper cites 2026 , eprint =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces 2026 , eprint =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:bf9becb21cb642b28182d2c2ecb767ec871d2ffcdcc7cb6c654a2ff3447ddd93

Observation 32c96666-bce1-412f-a7ea-35e976b6acc0 · outbound

This paper cites Findings of the Association for Computational Linguistics (.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Findings of the Association for Computational Linguistics (

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:144a02a7a5d649cc1877cb484e79d6e3027d026ec9762e158b6a6ac792094ef5

Observation d2321bcc-c45c-4a05-8b44-86d74c8aca49 · outbound

This paper cites Qin, Y ., Liang, S., Ye, Y ., Zhu, K., Yan, L., Lu, Y ., Lin, Y ., Cong, X., Tang, X., Qian, B., et al.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Qin, Y ., Liang, S., Ye, Y ., Zhu, K., Yan, L., Lu, Y ., Lin, Y ., Cong, X., Tang, X., Qian, B., et al

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:36:15.074645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:8dbed564d4f2d201a94685e65d91145350b2b2c47d93dd376581b4b070a37100

Observation b2878297-f368-4b5f-ab3b-26ac00db564f · outbound

This paper cites an unresolved cited work.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Unresolved cited work

Reference 14

Resolution
parse uncertain
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:f635097eb825d49821127e7468b4ab20c9e6d1080c908293310f79f66ab0a882

Observation 9c7489b0-7248-47a7-8e96-96eb6eb8252b · outbound

This paper cites 2025 , howpublished =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces 2025 , howpublished =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:eeab25862572f263d570f7e51906a64518d1a2f40178821efbc46ecfb1293950

Observation a553fda0-5d2c-4d78-9968-6217012cc40a · outbound

This paper cites Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks , journal =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks , journal =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:261359a829a8a800489de42af27654d6c59ff411d2f6e8ce219c799fb45acdb3

Observation a0a44ef4-f4a6-4b81-9857-e86861ec49e5 · outbound

This paper cites AgentDojo:.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces AgentDojo:

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:c75db8d51d11d4e5de9dc30f6aed28668f7e763933071bb6619e0bd7ed782e60

Observation ca446d45-f632-4055-8ae5-f677e5036214 · outbound

This paper cites Zico Kolter and Matt Fredrikson , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Zico Kolter and Matt Fredrikson , title =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:593aaef2616859b0dfa6f3d185f0a35bca7c86ffb528741828960a25d0e8cabf

Observation f1a0435e-8438-4dd0-8844-e0b1829426ff · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:bc595c55d90f1cdf7c9cf872782c1a19c68b751f40f775501d050a194a3fe078

Observation 88d93696-3e02-44b6-bf41-15f7cc886619 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models , booktitle =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces OR-Bench: An Over-Refusal Benchmark for Large Language Models , booktitle =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:179237df219ab88ead2d815dfab9e98cc2952b34baae0cdd31fa78630dd62cbd

Observation 1cc0622c-f20f-4d7e-91c4-cbbbfeea2273 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces The Twelfth International Conference on Learning Representations,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:90304859b180a1f5d5b2747a6dca02c7cae9f6d5e9d5324973414f031f551d6d

Observation ca95514f-71b0-451c-a019-a58b65b88753 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces The Thirteenth International Conference on Learning Representations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:a8205f76d8b3173d26d12adf44794f077dc5423459cdaa1131b930f3592b077b

Observation 63957f8f-a573-41e9-90bc-526d4c133cef · outbound

This paper cites Maddison and Tatsunori Hashimoto , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Maddison and Tatsunori Hashimoto , title =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:8407fbf813050e1492cde5b02c20ce005747e684bc9edcd7b7d09d33e7f82cfe

Observation bf7be383-ee5b-430c-98b0-968c893b5343 · outbound

This paper cites an unresolved cited work.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:c8810185e46a6abee8ace603495b84c1c03274a279fe17041b6fbe39b265b104

Observation 0e47e741-ab83-4a5a-aa3c-4c83b18a6e4d · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:15.066149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:3bb702286afcc9d97f735dccc7273d78456e72447332b874434a7f93088701de

Observation 71153747-9fc8-4aaa-b81c-dc3add28371f · outbound

This paper cites Findings of the Association for Computational Linguistics:.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Findings of the Association for Computational Linguistics:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:8e30bf245f0d78ad22093d0262954c69a1ff900125164c4f23d971fcbc0b798b

Pith citing papers

No inbound Pith citation observations are available.