Pith. sign in

REVIEW 1 cited by

Suggesting Code Edits in Interactive Machine Learning Notebooks Using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09745 v1 pith:ME633UFX submitted 2025-01-16 cs.SE cs.CLcs.LG

classification cs.SEcs.CLcs.LG
keywords notebooksjupyterlearningmachinecodeeditsmodelsdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning developers frequently use interactive computational notebooks, such as Jupyter notebooks, to host code for data processing and model training. Jupyter notebooks provide a convenient tool for writing machine learning pipelines and interactively observing outputs, however, maintaining Jupyter notebooks, e.g., to add new features or fix bugs, can be challenging due to the length and complexity of the notebooks. Moreover, there is no existing benchmark related to developer edits on Jupyter notebooks. To address this, we present the first dataset of 48,398 Jupyter notebook edits derived from 20,095 revisions of 792 machine learning repositories on GitHub, and perform the first study of the using LLMs to predict code edits in Jupyter notebooks. Our dataset captures granular details of cell-level and line-level modifications, offering a foundation for understanding real-world maintenance patterns in machine learning workflows. We observed that the edits on Jupyter notebooks are highly localized, with changes averaging only 166 lines of code in repositories. While larger models outperform smaller counterparts in code editing, all models have low accuracy on our dataset even after finetuning, demonstrating the complexity of real-world machine learning maintenance tasks. Our findings emphasize the critical role of contextual information in improving model performance and point toward promising avenues for advancing large language models' capabilities in engineering machine learning code.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CRABS: A syntactic-semantic pincer strategy for bounding LLM interpretation of Python notebooks

    cs.CL 2025-07 conditional novelty 7.0 of 10

    CRABS combines AST bounds with an LLM to reconstruct information flow and execution dependency graphs for Python notebooks, reaching 98% F1 on 50 curated Kaggle notebooks.

Pith tools