Pith. sign in

Paper Citation Record · LEDGER

ORCA-bench: How Ready Are Language Model Agents for Oncall?

As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.28545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28545 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:40:33.122717Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a67d0831-c5a3-4c43-8f87-13ac6660e72f · outbound

This paper cites Long Code Arena: a Set of Benchmarks for Long-Context Code Models.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Long Code Arena: a Set of Benchmarks for Long-Context Code Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:31.765264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:31.765264Z digest=sha256:6d435d8ed21d46105bad544a7efb0cdeefb62383754691c5a9ef33bc6bdd2bb4

Observation b6c465db-c1d9-4df1-92c7-714a303b96f3 · outbound

This paper cites SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios.

ORCA-bench: How Ready Are Language Model Agents for Oncall? SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:31.853540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:31.853540Z digest=sha256:9a8241160e536273f2cb9ba2751f124d202b8e2a287766683558d7c14e415b91

Observation 7c631941-4708-4774-8bf1-f51af23ab502 · outbound

This paper cites Itbench: Evaluating ai agents across diverse real-world it automation tasks.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Itbench: Evaluating ai agents across diverse real-world it automation tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.585210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:31.932782Z digest=sha256:3172e8c6f96fcef6e51d4b60fe7ad2b852745d61620e63dfb72a48688ac5a70f

Observation d4f1571b-17c5-4f18-9ce8-1fb6f9658f37 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107--54157, 2024.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107--54157, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.571700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.007547Z digest=sha256:d3ef3d0d4a8ade75ce27e6cbb141e84e099d45a783013c28442e8c0f9c53bed5

Observation c1f0203f-d55f-46db-a1e5-0f68d07d964b · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:32.178531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:32.178531Z digest=sha256:6d2bd9c398ac5c8e4ffde2ab0b624b6f48f38bb7d2f0303cd7c73c748faf9fef

Observation 3caebaa7-a09c-4176-83a1-95f96065a7ce · outbound

This paper cites Rcaeval: A benchmark for root cause analysis of microservice systems with telemetry data.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Rcaeval: A benchmark for root cause analysis of microservice systems with telemetry data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.556265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.274533Z digest=sha256:360218478598d44249b2eb2eb4f8c5c168d29535c5911bfadaee8a5da8d17149

Observation 703cdbae-f913-497b-85eb-0dc67a3bfe5e · outbound

This paper cites Building ai agents for autonomous clouds: Challenges and design principles.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Building ai agents for autonomous clouds: Challenges and design principles

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.412446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.340183Z digest=sha256:edbd4f628475755079829cd123a2db8a4c5523c2db5bbfc377621358ac0eb6a4

Observation 5b613cb5-f74a-422f-abcf-d769fe4da5d5 · outbound

This paper cites Openrca: Can large language models locate the root cause of software failures? In The Thirteenth International Conference on Learning Representations, 2025.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Openrca: Can large language models locate the root cause of software failures? In The Thirteenth International Conference on Learning Representations, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:34.158418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.442926Z digest=sha256:4bced18a17ee877d97101b46ff7d4224dd64ce07f5231224bf84740b1870a513

Observation 29c43866-cf73-42dc-badd-965ef8c563a2 · outbound

This paper cites Swe-bench multimodal: Do ai systems generalize to visual software domains? In The Thirteenth International Conference on Learning Representations, 2025.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Swe-bench multimodal: Do ai systems generalize to visual software domains? In The Thirteenth International Conference on Learning Representations, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:33.982615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.524970Z digest=sha256:d9c5abdf4b58feb38772773ce2705c9e3ce9d75414bdc9af0651059def2378ba

Observation 47f386de-8d47-4459-919b-1bd9a2799903 · outbound

This paper cites Swe-smith: Scaling data for software engineering agents.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Swe-smith: Scaling data for software engineering agents

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:33.716878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.658578Z digest=sha256:2cfc68fadf3f2d4b814e3817af7afb4276faba35ba5b07579ae6a56ea8475d41

Observation 0bdee2f2-9d83-4c48-81fd-52effa5b812e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:32.768418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:32.768418Z digest=sha256:bf4a752ec0d7c5942e25b178e750c6e8f1973b7a971f82e757c799c5f4b452e0

Observation b1b018bd-c6e6-4698-8063-b7a62fff3da5 · outbound

This paper cites Graders should cheat: privileged information enables expert-level automated evaluations.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Graders should cheat: privileged information enables expert-level automated evaluations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:40:33.507625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T04:40:32.836167Z digest=sha256:065c8ca271db272529e0e4fb8d2093048287fec1e5baa9c2bb53f88d5b85746b

Observation f595175f-81f7-4df1-b643-eb90207a3f55 · outbound

This paper cites @esa (Ref.

ORCA-bench: How Ready Are Language Model Agents for Oncall? @esa (Ref

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:32.940042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:32.940042Z digest=sha256:9024261f1ff6c21afa54ba6924fdd6effe3eea48a8d6bcab560189dcf4fe6bab

Observation 136b50f1-8939-4830-8fd4-305350631b9a · outbound

This paper cites an unresolved cited work.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:40:33.014494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:33.014494Z digest=sha256:8e79dd571c0a70f921fba11d7f19cdead46408936a6bca8a93304c298e46cb0d

Observation 2397f58b-10e4-4d29-b909-fa110d9a7746 · outbound

This paper cites an unresolved cited work.

ORCA-bench: How Ready Are Language Model Agents for Oncall? Unresolved cited work

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-06T04:40:33.122717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:40:33.122717Z digest=sha256:2eed1e52305a2d30ec8b491fcadd45ac737220e502ea04503ea8e1c9ce1c727e

Pith citing papers

No inbound Pith citation observations are available.