Pith. sign in

REVIEW 1 cited by

BERT_SE: A Pre-trained Language Representation Model for Software Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.00699 v1 pith:PBGSYQMD submitted 2021-12-01 cs.SE

classification cs.SE
keywords softwarebertmodelapplicationembeddingengineeringpre-trainedrequirements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The application of Natural Language Processing (NLP) has achieved a high level of relevance in several areas. In the field of software engineering (SE), NLP applications are based on the classification of similar texts (e.g. software requirements), applied in tasks of estimating software effort, selection of human resources, etc. Classifying software requirements has been a complex task, considering the informality and complexity inherent in the texts produced during the software development process. The pre-trained embedding models are shown as a viable alternative when considering the low volume of textual data labeled in the area of software engineering, as well as the lack of quality of these data. Although there is much research around the application of word embedding in several areas, to date, there is no knowledge of studies that have explored its application in the creation of a specific model for the domain of the SE area. Thus, this article presents the proposal for a contextualized embedding model, called BERT_SE, which allows the recognition of specific and relevant terms in the context of SE. The assessment of BERT_SE was performed using the software requirements classification task, demonstrating that this model has an average improvement rate of 13% concerning the BERT_base model, made available by the authors of BERT. The code and pre-trained models are available at https://github.com/elianedb.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Language Models Potential for Requirement Engineering Applications: Insights into Current Strengths and Limitations

    cs.SE 2024-12 conditional novelty 5.0 of 10

    ChatGPT and Gemini generally underperform task-specific models on requirements engineering benchmarks, except ChatGPT achieves a new top F1 score on the REQuestA question answering dataset.

Pith tools