Pith. sign in

Paper Citation Record · LEDGER

AlphaEval: Evaluating Agents in Production

As of 6 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2604.12162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12162 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:30:51.886471Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:09:38.698353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T08:16:48.330790Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact17
  • verified fuzzy1
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c71a63eb-01d1-4d29-af89-fcc9c93579a7 · outbound

This paper cites assignment.

AlphaEval: Evaluating Agents in Production assignment

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.840459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:63762a4ac4804c29933ff296c55b2ccc7acc30683823e5d56e5c09c431e46d91

Observation 4ee44383-207c-46ed-8d0a-5c6610146b1f · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AlphaEval: Evaluating Agents in Production BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:2c69136fb5575452fb7baf0503519e6f8bc59e6d8b070f4f274d9f781cf6d48b

Observation dc2d78a1-692c-445e-be93-18b5d2e7cc07 · outbound

This paper cites Evaluating Collective Behaviour of Hundreds of LLM Agents.

AlphaEval: Evaluating Agents in Production Evaluating Collective Behaviour of Hundreds of LLM Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:35.937490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:6ebcb5d022cb38b0f7c28c76ebe4d69eda379328dd4a0aa2b0c8793c732bb70c

Observation eaaccacc-27fb-4488-b716-c5d265e77e37 · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

AlphaEval: Evaluating Agents in Production xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.832941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:51ed520969b61a9275458d7c1620c131b1052fa72c6aebcf2cf51de7ed2a18ad

Observation c9169c1d-f4b1-46c4-8379-c9c115584ba8 · outbound

This paper cites Webworld: A large-scale world model for web agent training.

AlphaEval: Evaluating Agents in Production Webworld: A large-scale world model for web agent training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.856521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:78b617222dd23b8734764bd4dd21aa7caa87393e35c8c4bc5b82c19f477555ca

Observation ee85e413-d8c0-4d3b-8823-d78918b9f5f4 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

AlphaEval: Evaluating Agents in Production OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:56cd699218295fab53ce654c67327bb257214da193336db97bbbd518d5de32cc

Observation 6f6ff325-28bf-461d-8b22-c7bf70e4c284 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

AlphaEval: Evaluating Agents in Production TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:31.128039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:5501228899b1a0502fc376fefb602495274a732d700a5099b1404336db2f11f7

Observation 6e31f8ef-6b5a-44c4-8f5d-6d244db49fe9 · outbound

This paper cites SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?.

AlphaEval: Evaluating Agents in Production SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.877213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:88cfb6d1bff6a068dacbe4fbcd7659482f75cad72f737e96c0d8e51047821076

Observation cb0cb7a0-d26c-4ac9-a6af-8138c5ca62f9 · outbound

This paper cites $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980.

AlphaEval: Evaluating Agents in Production $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.885606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:9d40017345b6cc7c16d581eea0994ee7f8ef31cb20aa04082ca502dc519626a3

Observation 648d44ea-c9dd-4f93-ad56-2609760777da · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

AlphaEval: Evaluating Agents in Production $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:45:58.844635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:30043438c85163497dd45c6978e5362ec2ed7f182eef154bbeb03ff376912141

Observation f9f8fc77-4832-46ce-b469-f8754df43f45 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

AlphaEval: Evaluating Agents in Production MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:42.029646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:5926c09c94f281f06d8862e68857f5cc10640e25fc0a75ad77adf892a837a8a9

Observation 274d648c-9e66-442e-9fcd-957718ce5dec · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

AlphaEval: Evaluating Agents in Production Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:48:50.578231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:2ec3e54d7d1bbfc19cb083c21e996d4b61f0889ca71e9165e8eba65e8ec4aabf

Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.817569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:aac4353c4415b41e7f941b3af39bd44f12f1c7f8e2f45ac7d67d689a56bfadf2

Observation d4c1823a-3c93-420f-acf8-92088040d1e4 · outbound

This paper cites Zhang, Joey Ji, Celeste Menders, Riya Dulepet, T.

AlphaEval: Evaluating Agents in Production Zhang, Joey Ji, Celeste Menders, Riya Dulepet, T

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.811047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:c860ccc2a6139f9e100fbab8fe5341b51c753d6ed231a66c6f6ee898f2feb58d

Observation 29b4da49-2e8c-4344-8692-ed656bd1e3d2 · outbound

This paper cites Browsecomp-v3: A visual, vertical, and verifiable benchmark for multimodal browsing agents.

AlphaEval: Evaluating Agents in Production Browsecomp-v3: A visual, vertical, and verifiable benchmark for multimodal browsing agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.779600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:97deb8ebf6fd1a57bc76b091f058c66e008517f32a8af31e271153b00dcf4dab

Observation 0d493818-4289-42c2-b83e-756e62e965f3 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

AlphaEval: Evaluating Agents in Production Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:45:58.821265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:8c032f090d3fee75bcfb7365e40f646dcd80b002074f8e1c0577eb11b8222190

Observation d97ccce5-150f-411a-a200-04f6515d7e9b · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

AlphaEval: Evaluating Agents in Production WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:45:58.801295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:046c714dd79e1907dcec4d18882a9be4b344b814b65f2400e4438cda5bba2b0c

Observation 33fa85b9-3db3-4e29-814f-ffbaeec43fce · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

AlphaEval: Evaluating Agents in Production Agent-as-a-Judge: Evaluate Agents with Agents

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:45:58.864563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:f25ee8ac2097f41bb725cf519409343e133290eb4bb4b20f72f5e22866e9f67d

Observation 1ac7c2af-3768-4470-bb5f-00a94afaa875 · outbound

This paper cites an unresolved cited work.

AlphaEval: Evaluating Agents in Production Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-17T14:59:48.843258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:86d57257a685ca9e097c0df4369a91c06d9af54eb870f5243e841ba411515073

Observation 3ea5f684-78a2-4f75-8530-27b5a8ef1365 · outbound

This paper cites an unresolved cited work.

AlphaEval: Evaluating Agents in Production Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T14:59:48.840636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:8aaf40b418d1089a91ff8ddec875ccd8ddda42336e901e53ef3b22ec1c717354

Observation 826eda7f-4d7b-445b-b426-832aa1e3d216 · outbound

This paper cites no criteria.

AlphaEval: Evaluating Agents in Production no criteria

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:59:48.837685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:f0208628d5c57917e98fe49e24e029afa02d6ed588138d9c903c3a3629e863ea

Pith citing papers

Observation bd08c571-b324-4482-af18-467b6f476f9f · inbound

SentinelBench: A Benchmark for Long-Running Monitoring Agents cites this paper.

SentinelBench: A Benchmark for Long-Running Monitoring Agents AlphaEval: Evaluating Agents in Production

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:16:48.332342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T06:09:38.698353Z digest=sha256:e8cb9a756ac9286d08ee679301e0f4399bbeae8b9162ef0e1b13a80fb5e51248