Pith. sign in

REVIEW 1 cited by

Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19715 v2 pith:W4UO4CJZ submitted 2024-09-29 cs.CL

classification cs.CL
keywords codefeedbackcoffee-gymeditingmodelsdatasetenvironmenterroneous
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents Coffee-Gym, a comprehensive RL environment for training models that provide feedback on code editing. Coffee-Gym includes two major components: (1) Coffee, a dataset containing humans' code edit traces for coding questions and machine-written feedback for editing erroneous code; (2) CoffeeEval, a reward function that faithfully reflects the helpfulness of feedback by assessing the performance of the revised code in unit tests. With them, Coffee-Gym addresses the unavailability of high-quality datasets for training feedback models with RL, and provides more accurate rewards than the SOTA reward model (i.e., GPT-4). By applying Coffee-Gym, we elicit feedback models that outperform baselines in enhancing open-source code LLMs' code editing, making them comparable with closed-source LLMs. We make the dataset and the model checkpoint publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming

    cs.AI 2025-05 conditional novelty 7.0 of 10

    ELABORATION provides a four-stage human-feedback taxonomy and an 8,320-problem dataset, with experiments showing human-LLM collaboration improves pass@1 by about 7 percent.

Pith tools