Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:30:51.886471Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2604.12162.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:30:51.886471Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T06:09:38.698353Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T08:16:48.330790Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c71a63eb-01d1-4d29-af89-fcc9c93579a7 · outbound
AlphaEval: Evaluating Agents in Production assignment
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ee44383-207c-46ed-8d0a-5c6610146b1f · outbound
AlphaEval: Evaluating Agents in Production BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dc2d78a1-692c-445e-be93-18b5d2e7cc07 · outbound
AlphaEval: Evaluating Agents in Production Evaluating Collective Behaviour of Hundreds of LLM Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eaaccacc-27fb-4488-b716-c5d265e77e37 · outbound
AlphaEval: Evaluating Agents in Production xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9169c1d-f4b1-46c4-8379-c9c115584ba8 · outbound
AlphaEval: Evaluating Agents in Production Webworld: A large-scale world model for web agent training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee85e413-d8c0-4d3b-8823-d78918b9f5f4 · outbound
AlphaEval: Evaluating Agents in Production OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f6ff325-28bf-461d-8b22-c7bf70e4c284 · outbound
AlphaEval: Evaluating Agents in Production TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e31f8ef-6b5a-44c4-8f5d-6d244db49fe9 · outbound
AlphaEval: Evaluating Agents in Production SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb0cb7a0-d26c-4ac9-a6af-8138c5ca62f9 · outbound
AlphaEval: Evaluating Agents in Production $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 648d44ea-c9dd-4f93-ad56-2609760777da · outbound
AlphaEval: Evaluating Agents in Production $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9f8fc77-4832-46ce-b469-f8754df43f45 · outbound
AlphaEval: Evaluating Agents in Production MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 274d648c-9e66-442e-9fcd-957718ce5dec · outbound
AlphaEval: Evaluating Agents in Production Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · outbound
AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4c1823a-3c93-420f-acf8-92088040d1e4 · outbound
AlphaEval: Evaluating Agents in Production Zhang, Joey Ji, Celeste Menders, Riya Dulepet, T
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29b4da49-2e8c-4344-8692-ed656bd1e3d2 · outbound
AlphaEval: Evaluating Agents in Production Browsecomp-v3: A visual, vertical, and verifiable benchmark for multimodal browsing agents
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d493818-4289-42c2-b83e-756e62e965f3 · outbound
AlphaEval: Evaluating Agents in Production Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d97ccce5-150f-411a-a200-04f6515d7e9b · outbound
AlphaEval: Evaluating Agents in Production WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33fa85b9-3db3-4e29-814f-ffbaeec43fce · outbound
AlphaEval: Evaluating Agents in Production Agent-as-a-Judge: Evaluate Agents with Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ac7c2af-3768-4470-bb5f-00a94afaa875 · outbound
AlphaEval: Evaluating Agents in Production Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ea5f684-78a2-4f75-8530-27b5a8ef1365 · outbound
AlphaEval: Evaluating Agents in Production Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 826eda7f-4d7b-445b-b426-832aa1e3d216 · outbound
AlphaEval: Evaluating Agents in Production no criteria
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd08c571-b324-4482-af18-467b6f476f9f · inbound
SentinelBench: A Benchmark for Long-Running Monitoring Agents AlphaEval: Evaluating Agents in Production
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.