Pith. sign in

REVIEW 1 cited by

ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18125 v2 pith:FTLAEP3C submitted 2024-06-26 cs.CL cs.AIcs.CYcs.LG

ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models

classification cs.CL cs.AIcs.CYcs.LG
keywords classificationresumeaccuracymodelschallengesdatasetdatasetslanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The increasing reliance on online recruitment platforms coupled with the adoption of AI technologies has highlighted the critical need for efficient resume classification methods. However, challenges such as small datasets, lack of standardized resume templates, and privacy concerns hinder the accuracy and effectiveness of existing classification models. In this work, we address these challenges by presenting a comprehensive approach to resume classification. We curated a large-scale dataset of 13,389 resumes from diverse sources and employed Large Language Models (LLMs) such as BERT and Gemma1.1 2B for classification. Our results demonstrate significant improvements over traditional machine learning approaches, with our best model achieving a top-1 accuracy of 92\% and a top-5 accuracy of 97.5\%. These findings underscore the importance of dataset quality and advanced model architectures in enhancing the accuracy and robustness of resume classification systems, thus advancing the field of online recruitment practices.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reading Between the Lines: Classifying Resume Seniority with Large Language Models

    cs.CL 2025-09 conditional novelty 4.0

    Fine-tuned RoBERTa reached 90.6% accuracy on resume seniority classification using a new hybrid dataset, outperforming zero-shot GPT-4 and a TF-IDF baseline, though evaluation details are incomplete.