Pith. sign in

REVIEW 2 cited by

Can Large Language Models Write Parallel Code?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12554 v3 pith:CGS4HEZB submitted 2024-01-23 cs.DC cs.AI

classification cs.DCcs.AI
keywords codelanguagemodelsparalleldifferentgeneratelargeevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are increasingly becoming a popular tool for software development. Their ability to model and generate source code has been demonstrated in a variety of contexts, including code completion, summarization, translation, and lookup. However, they often struggle to generate code for complex programs. In this paper, we study the capabilities of state-of-the-art language models to generate parallel code. In order to evaluate language models, we create a benchmark, ParEval, consisting of prompts that represent 420 different coding tasks related to scientific and parallel computing. We use ParEval to evaluate the effectiveness of several state-of-the-art open- and closed-source language models on these tasks. We introduce novel metrics for evaluating the performance of generated code, and use them to explore how well each large language model performs for 12 different computational problem types and six different parallel programming models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors

    cs.DC 2025-08 conditional novelty 5.0 of 10

    An LLM-agent pipeline with profiling, binary analysis, and SMT simulation automatically parallelizes latency-critical benchmarks via the Relic framework, reporting a 17% geomean gain after excluding failures.

  2. Cross-Model Cross-Language AI Coding Agent Performance: Accuracy and Speed of Parallel CLRS Algorithms

    cs.SE 2026-07 conditional novelty 4.5 of 10

    Coding agents write correct parallel CLRS code with little prompting, but meaningful speedups are model-, language-, and algorithm-dependent, with Sonnet strongest and GPT producing none.

Pith tools