Pith. sign in

Paper Citation Record · LEDGER

WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2505.03733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03733 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:32.361437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:27:04.294250Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0afcb623-f615-4c20-98b2-6e90b041a284 · inbound

Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing cites this paper.

Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:32.361437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:32.361437Z digest=sha256:6bbed26af1503ce28b2f880efff9021a2ffd432c1c92b2d718da102cad594c8a

Observation 13694a49-581d-4ed8-8a05-2ccb34338de0 · inbound

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software cites this paper.

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:44.114748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:44.114748Z digest=sha256:8f3e6e643fbf0d24fee562bc00376eb0bc85f77732422b2989a4506870b1a1e7

Observation 01469e5f-6421-464a-afdc-22348cfa216c · inbound

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation cites this paper.

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:40:19.232153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:36:04.456245Z digest=sha256:28fb48cf71b1622ca3181d597216968f1525faa995ccfd133db41d011188439c

Observation 9e1457be-082c-483c-a7e2-1045ebf47852 · inbound

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning cites this paper.

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.729894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:18:18.340391Z digest=sha256:ef68444dfece676688e313360a545233bb0874c0853aaaba29586964e66c8423

Observation 4ff1504d-f614-4b47-8044-781daa8d3c33 · inbound

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies cites this paper.

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:08.931833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T16:24:49.710882Z digest=sha256:947562b0ae172fe36dfef3548aae86945b0a1e00bafe775321cc55eced120993

Observation 3a3d8263-d702-4510-999f-85fc96de1a2e · inbound

GameGen-Verifier: Parallel Keypoint-Based Verification for LLM-Generated Games via Runtime State Injection cites this paper.

GameGen-Verifier: Parallel Keypoint-Based Verification for LLM-Generated Games via Runtime State Injection WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:55.335095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:13:39.714918Z digest=sha256:7fb63c667eddc28a049cb0240170fb96828d7214b0693cbb743a8196b029aef8

Observation 6a1925d2-145a-4166-ba3c-b26ec5afe4cc · inbound

From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements cites this paper.

From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:22:52.285629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:22:10.879878Z digest=sha256:d18a78423430137dafc99ae136c1e0498011dc34e553af064fab1a842516a6c5

Observation 4ea8f743-b231-449e-9281-4049a6cd736d · inbound

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents cites this paper.

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:12:51.593866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T23:09:56.130802Z digest=sha256:51781f21b696efb405703829dd631c9b507b3120ccffb23e49002835d6e2a08b

Observation 00ca0e8f-e230-4e13-8c8d-e9176e509aa0 · inbound

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents cites this paper.

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:08:17.896885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T13:05:38.485058Z digest=sha256:b4b6e624277987c29bbe957397b7879af207330283a8e6f37e71e6abbe41b88c

Observation 2aa4923b-c732-4508-ac48-fd53dfb44ba2 · inbound

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering cites this paper.

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:43:14.965403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T22:43:13.251812Z digest=sha256:7fe653368f770ae2ed89f6d49797fce573ae47a98ec41b9c3fad572bf5975bf9

Observation ed6e5465-653c-494b-a89a-067a6d0a7888 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:17.235684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:9911d1e678d8ca46f93df733a5ab42f80a0898407b68514d4c4c14271e0d8039

Observation 1e345fc4-68bd-4168-b85d-95addb9972d9 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.245448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:15c77b1606582945f804d14ede14258f2ddf654d9f45e0d58794ba09a2c29b8b

Observation 5de4f68f-7581-498e-836f-ae8fa8be7db2 · inbound

HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML cites this paper.

HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:03:34.059095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T15:56:24.222373Z digest=sha256:6cf9afb368a0d56370b6c6f0f9548277e8ade1fd23cb5bfc0160b9d7551d25c4

Observation f7039c47-6d8b-4b1c-9bb6-b6fde9cde118 · inbound

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation cites this paper.

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.533156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T07:24:06.863653Z digest=sha256:f3218c2d79d2248b773703e5afddeec2ed7b9e024e681cd7bc362648c17c4731

Observation 3bcb69ed-e434-4381-bd5c-bc7b0d099c7b · inbound

QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits cites this paper.

QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:35:33.308090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T08:35:24.708864Z digest=sha256:7715c6613759d488bb6f109060c8d25f253f32c8b50ed8ddf2fb8bea82a5013d

Observation 4484d5b1-5e5c-41fc-bbfb-cb229759fd2c · inbound

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications cites this paper.

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.199841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T18:53:18.645984Z digest=sha256:8f02c6b8d092a9d454d1ec656f4d2ba883df4e4b2ed7ab640995267310838232

Observation 6f1ad68c-70e0-49cf-8b10-a98aaeb1b1b7 · inbound

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement cites this paper.

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:27:04.296226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T00:26:22.041924Z digest=sha256:d7039327a498689a18036c677d10af7d09823c2bff42b64d75091e8c5310ec2b

Observation df41b1aa-0ac2-48d4-9ca5-8afa6e89fc61 · inbound

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation cites this paper.

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T05:40:26.231504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:40:26.231504Z digest=sha256:1375f47e00c7ae10abf17ca156d21163560bbf727c200350f414d79a324f658e