Pith. sign in

REVIEW 1 cited by

Active PETs: Active Data Annotation Prioritisation for Few-Shot Claim Verification with Pattern Exploiting Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.08749 v2 pith:YO4JTYQA submitted 2022-08-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords dataactivefew-shotmodelsannotationclaimlabelledlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To mitigate the impact of the scarcity of labelled data on fact-checking systems, we focus on few-shot claim verification. Despite recent work on few-shot classification by proposing advanced language models, there is a dearth of research in data annotation prioritisation that improves the selection of the few shots to be labelled for optimal model performance. We propose Active PETs, a novel weighted approach that utilises an ensemble of Pattern Exploiting Training (PET) models based on various language models, to actively select unlabelled data as candidates for annotation. Using Active PETs for few-shot data selection shows consistent improvement over the baseline methods, on two technical fact-checking datasets and using six different pretrained language models. We show further improvement with Active PETs-o, which further integrates an oversampling strategy. Our approach enables effective selection of instances to be labelled where unlabelled data is abundant but resources for labelling are limited, leading to consistently improved few-shot claim verification performance. Our code is available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ALPET: Active Few-shot Learning for Citation Worthiness Detection in Low-Resource Wikipedia Languages

    cs.CL 2025-02 conditional novelty 6.0 of 10

    ALPET, an active-learning plus PET pipeline, detects citation-worthy sentences in Catalan, Basque and Albanian while needing roughly 58-72% fewer labeled examples than its CCW baseline.

Pith tools