Pith. sign in

REVIEW 15 cited by

Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.20098 v2 pith:WIH2RIX7 submitted 2024-06-28 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords webpagemllmsunderstandingcodedatasetcontentevaluationframework
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Multimodal large language models (MLLMs) have shown impressive success across modalities such as image, video, and audio in a variety of understanding and generation tasks. However, current MLLMs are surprisingly poor at understanding webpage screenshots and generating their corresponding HTML code. To address this problem, we propose $\texttt{Web2Code}$, a benchmark consisting of a new large-scale webpage-to-code dataset for instruction tuning and an evaluation framework for the webpage understanding and HTML code translation abilities of MLLMs. For dataset construction, we leverage pretrained LLMs to enhance existing webpage-to-code datasets as well as generate a diverse pool of new webpages rendered into images. Specifically, the inputs are webpage images and instructions, while the responses are the webpage's HTML code. We further include diverse natural language QA pairs about the webpage content in the responses to enable a more comprehensive understanding of the web content. To evaluate model performance in these tasks, we develop an evaluation framework for testing MLLMs' abilities in webpage understanding and web-to-code generation. Extensive experiments show that our proposed dataset is beneficial not only to our proposed tasks but also in the general visual domain. We hope our work will contribute to the development of general MLLMs suitable for web-based content generation and task automation. Our data and code are available at https://github.com/MBZUAI-LLM/web2code.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

    cs.CV 2026-08 conditional novelty 7.0 of 10

    MT-Web2Code is a 102-page, 16-domain multi-turn benchmark that measures how coding agents reconstruct missing web regions and fix localized defects while preserving the surrounding page.

  2. LiveEvalBench: Toward Open-World Evaluation for Web Generation

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A multi-agent, adaptive benchmark framework for evaluating LLM-generated frontend projects, with a 100-query benchmark, an 11-model leaderboard, and 86% agreement with human UI ratings.

  3. WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    WebMMU introduces a multilingual, three-task benchmark for website understanding and code generation, and finds current MLLMs underperform on reasoning, grounding, and functional code editing.

  4. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  5. FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

    cs.SE 2025-06 conditional novelty 6.0 of 10

    A new benchmark adds 148 interactive front-end development tasks with automated sandbox tests, reporting a 90.54% agreement rate with human evaluation across four LLMs.

  6. WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code

    cs.CL 2025-06 conditional novelty 6.0 of 10

    WebUIBench is a 21,793-question benchmark that splits WebUI-to-Code into perception, HTML programming, and cross-modal understanding, and it ranks 29 multimodal LLMs on each sub-skill.

  7. FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

    cs.CL 2025-05 conditional novelty 6.0 of 10

    FullFront adds a three-task benchmark for webpage design, perception, and code generation, and finds top MLLMs still fail at fine-grained layout and interaction implementation.

  8. P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark

    cs.CL 2025-05 conditional novelty 6.0 of 10

    P2P is a multi-agent framework that automatically generates HTML-rendered academic posters from papers, backed by a 30k instruction dataset and a 121-pair evaluation benchmark.

  9. UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage Designs

    cs.SE 2025-05 conditional novelty 6.0 of 10

    UICopilot generates webpage code from screenshots in two stages: a trained structure model predicts a coarse DOM tree with bounding boxes, then GPT-4V generates per-region code and refines global styles, improving vis...

  10. ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

    cs.AI 2025-01 conditional novelty 6.0 of 10

    ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.

  11. MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs

    cs.SE 2024-12 conditional novelty 6.0 of 10

    A resource-list representation and a 500-site benchmark let multimodal LLMs generate web code with real links, images, and routes, lifting resource matching from ~0% to 66-80%.

  12. VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    Merging a coding LLM into a vision-language model via task vectors yields an open-source multimodal coder that reaches near-GPT-4o performance on the authors' new benchmark.

  13. SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SlideCoder converts slide design images to editable python-pptx code and reports large gains over prior baselines on a new difficulty-tiered benchmark.

  14. Multimodal graph representation learning for website generation based on visual sketch

    cs.LG 2025-04 conditional novelty 5.0 of 10

    Adding a graph built from OCR text boxes and segmented visual blocks to a Flamingo-style vision-language model improves HTML generation on the WebSight benchmark, but gains do not generalize to the Design2Code benchmark.

  15. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.

Pith tools