REVIEW 4 cited by
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Developers often dedicate significant time to maintaining and refactoring existing code. However, most prior work on generative models for code focuses solely on creating new code, overlooking the distinctive needs of editing existing code. In this work, we explore a multi-round code auto-editing setting, aiming to predict edits to a code region based on recent changes within the same codebase. Our model, Coeditor, is a fine-tuned language model specifically designed for code editing tasks. We represent code changes using a line diff format and employ static analysis to form large customized model contexts, ensuring the availability of appropriate information for prediction. We collect a code editing dataset from the commit histories of 1650 open-source Python projects for training and evaluation. In a simplified single-round, single-edit task, Coeditor significantly outperforms GPT-3.5 and SOTA open-source code completion models (bringing exact-match accuracy from 34.7 up to 60.4), demonstrating the benefits of incorporating editing history for code completion. In a multi-round, multi-edit setting, we observe substantial gains by iteratively conditioning on additional user edits. We have open-sourced our code, data, and model weights to encourage future research and have released a VSCode extension powered by our model for interactive IDE usage.
Forward citations
Cited by 4 Pith papers
-
Pull Requests as a Training Signal for Repo-Level Code Editing
Mid-training a 32B code model on two million verified Search/Replace pull requests improves SWE-bench Lite by 13.6% and Verified by 12.3% over the instruct-tuned baseline under a simplified agentless protocol.
-
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
An LLM agent that reads a repository, gathers setup hints, and iteratively runs commands can build and test 33 of 50 popular projects across 14 languages.
-
Treefix: Enabling Execution with a Tree of Prefixes
By iteratively generating and refining code prefixes with an LLM, Treefix reaches 84% and 82% line coverage on two Python snippet datasets, exceeding prior learning-guided execution tools by 25 and 7 percentage points.
-
From Legal Text to Tech Specs: Generative AI's Interpretation of Consent in Privacy Law
An LLM pipeline that flags and fixes non-compliant consent use cases works imperfectly: it catches about two-thirds of relevant cases with reasoning prompts, and most of its fixes are legally sound but logically inconsistent.
Discussion (0). Continue with ORCID to comment.