Pith. sign in

Paper Citation Record · LEDGER

REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2504.11543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11543 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:38:32.994746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b62f5eb3-f9dc-4bd6-bfd0-b76d32dfccb1 · inbound

WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents cites this paper.

WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:31:52.822244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T22:31:19.752228Z digest=sha256:1fe2f5164cfa7ddd757e3b36bb40bdb99104a016686d593cad58dfdf51583c5b

Observation 23a83ecb-31b9-40e9-87a5-5ca57130290d · inbound

Real-Time Procedural Learning From Experience for AI Agents cites this paper.

Real-Time Procedural Learning From Experience for AI Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:24:04.658536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T05:21:35.308929Z digest=sha256:2172ab50cb4a286dd7086fd47842baad195a66819ecb57f788aae0522e5b3e41

Observation cdfa8c19-b8ca-43eb-91a3-f9f72d5e4b90 · inbound

MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report) cites this paper.

MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report) REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:03.359460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T04:45:01.799128Z digest=sha256:a74d236ecacc4af7c4c6db98ba3c979102ccb1659b1b7b035b7f2cd78d6a095f

Observation 870bf046-e6c3-4605-a0dd-5deb4aa14fee · inbound

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents cites this paper.

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:01:20.695632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T22:59:16.413568Z digest=sha256:6d78d448e8487cddf7da70980e24a71f26a5f53168dee19169513b91af6a7fa7

Observation ce019e2d-cfce-4943-8f1c-2d91b59ec7ba · inbound

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents cites this paper.

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:38:32.994746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:38:32.994746Z digest=sha256:1b9e455a9d02708fe842e56feae30a835628e6b073e0219c79b8c4d709b57ab3

Observation f1320e27-f604-45d3-ac57-9e0aed5d4ae9 · inbound

WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents cites this paper.

WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:10.677991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T16:50:02.527821Z digest=sha256:ab73109f8d3e6f27e918df7418c398e1ea62c1aa1715e3488b6098d2aa77a9f4

Observation b9047a07-228a-4669-96c5-51f3f49b0bef · inbound

ClawBench: Can AI Agents Complete Everyday Online Tasks? cites this paper.

ClawBench: Can AI Agents Complete Everyday Online Tasks? REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:01.314440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:34:49.746922Z digest=sha256:de2c28b9af0b95f47cf31e99c8f88ca54f994c273ec8cce98e18d980cb4893d6

Observation 46f7b5cf-32e3-4f6b-a698-d4839af35e6c · inbound

ClawBench: Can AI Agents Complete Everyday Online Tasks? cites this paper.

ClawBench: Can AI Agents Complete Everyday Online Tasks? REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T16:32:51.938035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:32:51.938035Z digest=sha256:fc8f2d9ff1887ec75247573235a6351308d8f1f65a4c07e09cc43430fdce1337

Observation b31ef8b3-f27a-4504-b0ee-2daf9c58dd37 · inbound

HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks cites this paper.

HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:21:00.432980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:41:36.066333Z digest=sha256:1a179c6770b406d20b55391021239ff9e14e1bfdd487d0b3a0074d36bd72ff43

Observation 38b806a4-ca78-4599-897f-8828f0125e57 · inbound

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? cites this paper.

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:10.399622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T20:09:11.566825Z digest=sha256:3dd30fbc1639d3cc3c39d7948705acf89e08cde7d920f684a3aa91800bcedb3d

Observation 0e854b0b-0cb1-4f13-9ccf-146d2fa7594f · inbound

FlowEval: Reference-based Evaluation of Generated User Interfaces cites this paper.

FlowEval: Reference-based Evaluation of Generated User Interfaces REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:07.353668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T17:42:13.097342Z digest=sha256:3342ca2188c61df834546380fc447b12009bbbb72251e0cc04f7877c264a429d

Observation abbafd9f-7cce-4808-9ccb-b851eda4705d · inbound

Computer Use at the Edge of the Statistical Precipice cites this paper.

Computer Use at the Edge of the Statistical Precipice REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:24.677327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T01:08:15.736001Z digest=sha256:b8c1d1f26e6fcf0ed9eff597879513008ed3674b49808854645857cae97bca01

Observation b5fa657a-fcd3-4cd8-b3f7-b17b5413e902 · inbound

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents cites this paper.

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.205463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T10:29:17.065496Z digest=sha256:5e3b694f2f16228231d6636df2dc18261bd4ad0cca32443ac89700961a500809

Observation 54a287eb-57c8-4240-8dee-e0122e0c76a5 · inbound

Signal-Driven Observation for Long-Horizon Web Agents cites this paper.

Signal-Driven Observation for Long-Horizon Web Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:31:29.166504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T01:25:11.149977Z digest=sha256:7ca982e2b3054e39c3a3cc7f9c0d9b99f7982ba8be83395790372eba38afd2c8

Observation a1efb831-1483-46fe-935d-97051ea45731 · inbound

Uncertainty Decomposition for Clarification Seeking in LLM Agents cites this paper.

Uncertainty Decomposition for Clarification Seeking in LLM Agents REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:21.342388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T20:44:06.027685Z digest=sha256:0e342217fe84e4b134e14accbdc511cad0856f5eeb31a4c21afb0155793f966d

Observation 6dc4e38a-c8f0-4bd9-a5f7-1d1a654824fe · inbound

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation cites this paper.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.149393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:c1d70669119bc22088c95d3003a9bee3d0c259a03b31f549bace9134d57e2329