REVIEW 5 cited by
Comparing BERT against traditional machine learning text classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The BERT model has arisen as a popular state-of-the-art machine learning model in the recent years that is able to cope with multiple NLP tasks such as supervised text classification without human supervision. Its flexibility to cope with any type of corpus delivering great results has make this approach very popular not only in academia but also in the industry. Although, there are lots of different approaches that have been used throughout the years with success. In this work, we first present BERT and include a little review on classical NLP approaches. Then, we empirically test with a suite of experiments dealing different scenarios the behaviour of BERT against the traditional TF-IDF vocabulary fed to machine learning algorithms. Our purpose of this work is to add empirical evidence to support or refuse the use of BERT as a default on NLP tasks. Experiments show the superiority of BERT and its independence of features of the NLP problem such as the language of the text adding empirical evidence to use BERT as a default technique to be used in NLP problems.
Forward citations
Cited by 5 Pith papers
-
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
Ethical refusals in LLM responses sharply reduce user win rates in Chatbot Arena compared to technical refusals and normal answers, though detailed refusals and clearly harmful prompts reduce the penalty.
-
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
Across five offensiveness and hate speech datasets, LLM alignment with human annotators is inconsistent for gender and ethnicity, with the only rounded-average consistency being lower alignment with Black than White a...
-
Measuring and Evaluating the Performance of Generative AI Models for Scam Detection
A new benchmark of 2,742 real scam messages shows top LLMs reach about 64-65% micro-F1 and generalize to an unseen 59,991-sample proprietary set better than a fine-tuned BERT.
-
An Unsupervised Anomaly Detection in Electricity Consumption Using Reinforcement Learning and Time Series Forest Based Framework
A DQN agent that picks among six anomaly detectors using time series forest rewards achieves high F1 on two electricity datasets, but the evaluation is in-sample.
-
To Ensemble or Not: Assessing Majority Voting Strategies for Phishing Detection with Large Language Models
Majority-voting ensembles of LLMs for phishing URL detection improve on the best single model only when ensemble members have comparable performance.
Discussion (0). Continue with ORCID to comment.