Pith. sign in

REVIEW 2 cited by

Large Language Models in the Workplace: A Case Study on Prompt Engineering for Job Type Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07142 v3 pith:QCYUAA7D submitted 2023-03-13 cs.CL

classification cs.CL
keywords promptmodelsclassificationengineeringgpt-3languagemodelperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This case study investigates the task of job classification in a real-world setting, where the goal is to determine whether an English-language job posting is appropriate for a graduate or entry-level position. We explore multiple approaches to text classification, including supervised approaches such as traditional models like Support Vector Machines (SVMs) and state-of-the-art deep learning methods such as DeBERTa. We compare them with Large Language Models (LLMs) used in both few-shot and zero-shot classification settings. To accomplish this task, we employ prompt engineering, a technique that involves designing prompts to guide the LLMs towards the desired output. Specifically, we evaluate the performance of two commercially available state-of-the-art GPT-3.5-based language models, text-davinci-003 and gpt-3.5-turbo. We also conduct a detailed analysis of the impact of different aspects of prompt engineering on the model's performance. Our results show that, with a well-designed prompt, a zero-shot gpt-3.5-turbo classifier outperforms all other models, achieving a 6% increase in Precision@95% Recall compared to the best supervised approach. Furthermore, we observe that the wording of the prompt is a critical factor in eliciting the appropriate "reasoning" in the model, and that seemingly minor aspects of the prompt significantly affect the model's performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

    cs.AI 2026-01 conditional novelty 6.0 of 10

    In simulated emergency-department encounters, 20 LLMs acquiesced to patient pressure for unindicated CT scans, antibiotics, or opioids at rates varying from 0% to 100%, with no clear link to model capability or recency.

  2. Prompt Engineering Guidelines for Using Large Language Models in Requirements Engineering

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A literature review and three expert interviews yield a proposed mapping of prompt engineering guideline themes onto five requirements engineering activities, with no empirical validation of the mapping.

Pith tools