Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:40:26.603500Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 8 inbound Pith citation observations for arXiv:2506.13832.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:40:26.603500Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:41:00.426393Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T14:27:04.297349Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f38cf3b7-8d95-4403-a9e7-a21208e8ad1e · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebb72d8c-dc26-4950-a9b9-e01f1ba9c800 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d113b5e-29f4-412a-92c2-f9831bda95f4 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3ce9ea-1ee8-41ad-a4f5-c54d6e0b283e · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation McEval: Massively Multilingual Code Evaluation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5e4b73-9850-4cc1-a373-24fd1f44151e · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c25eaa2-48b7-4843-860d-aa1b44dd5769 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cea8bf26-3f69-4b31-a11b-13874dc09bcd · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Measuring Coding Challenge Competence With APPS
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450bbcb9-5346-4fbe-9f4e-89d6a55f07af · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a611a773-c084-4928-a60e-ded7b64f1ecb · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8b8e56-cad2-41f2-b87b-77e607820ee8 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f003ea3-1196-4f8d-8718-2befa1349008 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9deef4d0-2bfa-43ae-b3fb-4e519aa81c4c · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c44362ac-b167-4263-9f06-b2c9384a52f5 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af0d42e-2235-4ec4-97b4-8b75deb804d6 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation DeepSeek-V3 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33f063d-aae9-4600-ae1c-1d6e12f31b2a · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5ba76ff-83e9-48af-9e76-e5f4482bc5d7 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a5a90b-6bc2-4134-a3d7-a46da2a880fa · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4947d0e8-ed72-4ae7-b606-ee953078acab · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2b3f76-d544-44ac-80ac-878b3712f9b1 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c92ee6-2ec0-4389-83d7-3708f1d96df3 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd56bf7-d9c0-4399-ad76-fcc4772a3d94 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1bf8f35-5a93-4ffa-bc67-23a84c5565bf · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Gemini: A Family of Highly Capable Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31c68b0-1d1c-4ae7-9254-938acc54e847 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation CodeJudge: Evaluating Code Generation with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a646339-9277-478c-981d-11078d93e809 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59c06fc5-b529-415d-996a-b916e6029b66 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b4fa840-a32e-4322-9dd9-737b774a5a7f · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf31d789-8049-4b07-ab27-d25861929a81 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4886c08f-ad2b-4fd5-9f6d-722a3b402c16 · outbound
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18a88b6-65fa-4483-adfc-5fa674fc85d3 · inbound
GameDevBench: Evaluating Agentic Capabilities Through Game Development FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71363743-5b87-42af-bc7c-e07ff116eb50 · inbound
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db5ee38-68e8-4650-8364-e0b8193d97e9 · inbound
OpenGame: Open Agentic Coding for Games FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5551a81d-9c35-462b-bda8-dcd614cdd1ff · inbound
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a52155c3-6b6a-49db-856e-a13fa93a26cb · inbound
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d5e393e-9640-433e-b82f-0fce3a287e62 · inbound
WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83b9c506-ab94-427d-8033-6c05aaa18cb0 · inbound
Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee6f492a-af4d-4e67-84f0-eedd510bacac · inbound
Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.