Pith. sign in

REVIEW 1 cited by

WebApp1K: A Practical Code-Generation Benchmark for Web App Development

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00019 v1 pith:5IADEXEY submitted 2024-07-30 cs.SE cs.AI

classification cs.SEcs.AI
keywords benchmarkwebapp1kcodecode-generationcorrectnessllmsmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce WebApp1K, a practical code-generation benchmark to measure LLM ability to develop web apps. This benchmark aims to calibrate LLM output and aid the models to progressively improve code correctness and functionality. The benchmark is lightweight and easy to run. We present the initial version of WebApp1K, and share our findings of running the benchmark against the latest frontier LLMs. First, open source LLMs deliver impressive performance, closely trailing behind GPT-4o and Claude 3.5. Second, model size has strong correlation with code correctness. Third, no prompting techniques have been found to lift performance either universally to all models, or significantly to a single model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation

    cs.SE 2025-07 conditional novelty 7.0 of 10

    CrossPL, a 1,982-task benchmark built from GitHub repositories, shows that LLMs achieve at most 79.74% pass@1 on cross-language IPC code generation and struggle with low-level protocols like Pipe.

Pith tools