Pith. sign in

REVIEW 1 cited by

Source Code Data Augmentation for Deep Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.19915 v4 pith:RTCYJXSZ submitted 2023-05-31 cs.CL cs.AIcs.SE

classification cs.CLcs.AIcs.SE
keywords codesourcedataaugmentationcomprehensivedeeplearningmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasingly popular adoption of deep learning models in many critical source code tasks motivates the development of data augmentation (DA) techniques to enhance training data and improve various capabilities (e.g., robustness and generalizability) of these models. Although a series of DA methods have been proposed and tailored for source code models, there lacks a comprehensive survey and examination to understand their effectiveness and implications. This paper fills this gap by conducting a comprehensive and integrative survey of data augmentation for source code, wherein we systematically compile and encapsulate existing literature to provide a comprehensive overview of the field. We start with an introduction of data augmentation in source code and then provide a discussion on major representative approaches. Next, we highlight the general strategies and techniques to optimize the DA quality. Subsequently, we underscore techniques useful in real-world source code scenarios and downstream tasks. Finally, we outline the prevailing challenges and potential opportunities for future research. In essence, we aim to demystify the corpus of existing literature on source code DA for deep learning, and foster further exploration in this sphere. Complementing this, we present a continually updated GitHub repository that hosts a list of update-to-date papers on DA for source code modeling, accessible at \url{https://github.com/terryyz/DataAug4Code}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CASPER: Contrastive Approach for Smart Ponzi Scheme Detecter with More Negative Samples

    cs.CR 2025-07 reject novelty 4.0 of 10

    CASPER claims a triplet-view contrastive learning method with an equal-angle similarity vector improves smart Ponzi scheme detection over SourceP, especially with only 25% labels.

Pith tools