REVIEW 4 major objections 5 minor 100 references
ALPET: Active Few-shot Learning for Citation Worthiness Detection in Low-Resource Wikipedia Languages
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read With just 50 labelled examples per class, ALPET reaches an average F1 of 55% on citation worthiness detection across Catalan, Basque, and Albanian Wikipedia, and does so with far fewer labels than the CCW baseline.
desk verdict Useful empirical study of AL+PET for citation-worthiness detection in three low-resource Wikipedias, but the headline claim of beating the existing CCW baseline is not backed by the comparison actually run. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the combination of pool-based active learning with Pattern-Exploiting Training (PET). Active learning provides the acquisition functions, including maximum and minimum average distance, greedy and lightweight coresets, anchor subsampling, prediction entropy, least confidence, breaking ties, CAL, and ALPS, that decide which sentences get labelled; PET then turns each selected sentence into a fill-in-the-blank pattern such as "Aquesta frase va [mask] citació" and maps the masked token's predicted word to the citation-worthy or not-citation-worthy label through a verbalizer. Training happens in six rounds with ten incremental sample sizes from 50 to 500 examples per class, which is the design that lets the paper measure data efficiency and identify the plateau at 300 samples. The same AL query strategies are applied to both ALPET and the adapted CCW baseline so that the comparison is meant to isolate the effect of PET with mBERT versus plain mBERT classification on the same actively selected data.
What would settle it
Run the original, unmodified CCW model on the same ca/eu/sq citation-needed datasets with 50, 100, and 200 labelled examples per class; if it matches or beats ALPET's 50-shot average F1 of 55% at any of these budgets, the paper's central data-efficiency claim fails. A second check: repeat ALPET with labels produced by real human annotators instead of the simulated oracle, and see whether the 50-shot advantage and the 300-sample plateau survive.
Extended reading notes
Core claim
The paper's central claim is that an active-learning-driven few-shot pipeline can detect citation-worthy sentences in low-resource Wikipedias with far fewer labels than the existing contextual model. Concretely, ALPET selects batches of sentences from a large unlabeled pool using diversity, uncertainty, and hybrid query strategies; removes near-duplicates; balances the classes; and trains a Pattern-Exploiting-Training classifier on multilingual BERT with language-specific cloze patterns and verbalizers. Averaged over all query strategies and the three languages, ALPET scores 55% macro F1 at 50 labelled examples per class, while the CCW baseline reaches the same score only at around 200 examples; reductions in labelled-data requirement average 70% for Catalan, 58% for Basque, and 72% for Albanian. The paper further claims that F1 gains shrink to about 1-2% beyond 300 labelled samples, and that only K-Means-clustering-based strategies such as ALPS and LightweightCoreset reliably beat random sampling. The intended upshot is that the main bottleneck for this task, labelled-data scarcity, can be meaningfully reduced rather than simply accepted.
Load-bearing premise
The comparison rests on the assumption that the CCW baseline the authors built, the published CCW model adapted to include an active-learning data-selection step, faithfully represents the original CCW system, so that ALPET's advantage is measured against the right rival.
Editorial extensions
If this is right
- If ALPET is right, a citation-worthiness detector for a new low-resource language can be bootstrapped with about 50 labelled sentences per class and reach F1 around 55%, making manual annotation budgets feasible for small Wikipedia communities.
- Adding labels beyond roughly 300 per class yields less than 1-2% F1 improvement, so annotation effort is better spent on other languages or tasks once that budget is reached.
- Random sampling is a respectable default selection strategy; sophisticated query strategies need to justify their extra computation, since only a few strategies such as ALPS, LightweightCoreset, euclidean-min, and euclidean-cycle beat it in some configurations.
- The same active-plus-few-shot recipe could transfer to neighbouring fact-checking tasks such as claim detection and rumour detection under low-resource constraints, as the authors suggest.
Reading between the lines
- A direct comparison against the original, unmodified CCW implementation would settle whether the reported savings come from Pattern-Exploiting Training, from active data selection, or from both, since the paper's adapted baseline bundles the two changes together.
- Because the oracle is simulated from existing labels, real human-in-the-loop annotation may shift which query strategies win; uncertainty-based strategies could lose value if human labels are noisy or inconsistent, while diversity-based selection might gain value.
- The authors' observation that AL strategies help more when the unlabeled pool is large suggests a testable extension: on a bigger multilingual Wikipedia corpus, K-Means-based selection such as ALPS should show a larger and more consistent margin over random sampling.
- The plateau at 300 samples hints that further gains for low-resource CWD may come less from more labels than from better patterns, verbalizers, or pretrained representations, which the paper did not optimize.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ALPET, a framework that combines active learning (AL) query strategies with Pattern-Exploiting Training (PET) on multilingual BERT for citation worthiness detection (CWD) in three low-resource Wikipedia languages (Catalan, Basque, Albanian). The authors report that ALPET reaches an average macro-F1 of 55% with only 50 labeled examples per class, while their adapted CCW baseline needs about 200 examples to reach the same level, corresponding to claimed labeled-data reductions of 58--72% depending on the language. They also report a performance plateau after 300 labeled examples and, notably, that most AL query strategies do not consistently outperform random sampling. The main technical concerns are that the baseline is an adapted version of CCW rather than the published model, that PET patterns/verbalizers were selected without a documented validation protocol, and that reported F1 differences lack error bars or significance testing.
Significance. If the claims hold, ALPET would be a practically useful data-efficient method for CWD in low-resource languages, and the paper would be among the first to integrate AL with PET for this task. The study's systematic comparison of many AL strategies and its candid negative result--that random sampling remains competitive--are valuable empirical findings for the AL community. Strengths include the use of three public Wikipedia datasets, average results over multiple rounds and training iterations, and an explicit limitation section acknowledging the simulated-oracle setting. However, the headline claim of outperforming the existing CCW baseline is not currently supported because the implemented baseline discards the contextual features that define CCW. The paper's contribution is therefore better characterized as an evaluation of PET+AL against mBERT+AL, with the claimed advantage over the published CCW still unverified.
major comments (4)
- [Section 5.3 (Baseline Model), with Section 5.1 (Datasets)] The comparison against CCW is not a comparison against the published CCW model. Section 5.3 states that the authors 'adapted CCW by adding an AL step for data selection,' while Section 5.1 says they 'focused exclusively on two components: the text of sentences ... and their labels,' omitting the adjacent-sentence context and topic categories that define CCW in reference [4]. The abstract and conclusion claim that ALPET 'outperforms the existing CCW baseline,' but the experiments actually compare PET+mBERT on AL-selected data with mBERT on AL-selected data. This is a load-bearing mismatch: the claimed state-of-the-art improvement and the 58--72% data-efficiency numbers are relative to a baseline that is not the existing CCW model. Please either re-run the original CCW model with its contextual features in the AL setting, or revise the central claims and abstract to state explicitly that the comparison is against an mBERT classifier trained on AL-selected sentences.
- [Section 5.2 (PET with Active Samples), Table 2] The PET pattern and verbalizer selection is not described with a held-out protocol. The text says 'we experimented with a couple of patterns and we choose the best performing ones to report the final results on.' If patterns are chosen based on test-set performance, the reported F1 scores are optimistically biased. Please specify exactly how many patterns were tried per language, whether selection was made on the development set, and whether the reported numbers account for this selection. Without this information, the relative gains of ALPET over the baseline could be partly an artifact of pattern tuning.
- [Section 6 (Experiment Results), Figures 3--5] All central F1 results are reported as point averages without variance or significance tests, even though the experimental setup includes six rounds and three training iterations. For example, Section 6.1 describes differences of '2%-3%' between models, and Section 6.2 asserts a plateau after 300 labeled examples based on changes of about 0.1--1.9 percentage points. These claims are not supported without error bars, confidence intervals, or paired significance tests across the rounds. Please report mean plus standard deviation (or equivalent) for the F1 curves and use an appropriate test for the plateau and model comparisons, or clearly label such differences as suggestive rather than conclusive.
- [Section 2 (Research Objectives and Hypotheses) and Section 6.3 (H3 evaluation)] Hypothesis H3 is stated with a specific quantitative range--'58-72% fewer labelled examples'--before the experiments are described, and Section 6.3 reports reductions of 70% (CA), 58% (EU), and 72% (SQ), exactly matching that range. The paper does not state whether H3 was pre-registered, derived from pilot experiments, or formulated after seeing the results. Additionally, the reduction calculation in Section 6.3 uses ALPET's own 50-shot F1 of 55% as the benchmark for CCW, which anchors the metric to the method being evaluated. Please disclose the origin of the numeric hypothesis and compute data-efficiency reductions relative to an independent performance target, with uncertainty intervals.
minor comments (5)
- [Section 5.2 and Introduction] Please unify the terminology: the Introduction uses 'Pattern Exploit Training' while Section 4.4 correctly uses 'Pattern-Exploiting Training'; also fix the typo 'maually pre-define patterns' in Section 5.2.
- [Table 1] The formatting of Table 1 is confusing: the row 'Train Sets r1-r6 500 each r1-r6 500 each r1-r6' is not a clear description of the six rounds. Please restructure the table so that it is obvious that each of the six rounds contains 500 citation and 500 non-citation training sentences per language.
- [Section 6.4 (Effectiveness of Active Learning Query Strategies)] The sign convention in the random-sampling comparison is counterintuitive: the text says an AL strategy is better if the difference is negative, but the heatmap discussion later describes red as better performance for AL. Please clarify the convention or reverse it so that positive values consistently indicate improvement over random sampling.
- [Section 4.2.1 (Duplicate and similarity removal)] The cosine similarity threshold of 0.8 is justified only by qualitative observation ('almost identical except for minor variations'). Please report how many sentences were removed per language and, ideally, include a small sensitivity analysis to show that the main results are not sensitive to this threshold.
- [Reproducibility] The paper does not mention whether code and configuration files will be released. Given the large number of AL strategies, PET patterns, and hyperparameter choices, releasing the experimental code would substantially strengthen reproducibility and the ability of readers to verify the reductions reported in Table 4.
Circularity Check
No significant circularity: ALPET is an empirical comparison with no mathematical derivation; the adapted-CCW baseline raises validity questions, not circularity.
full rationale
This paper is an empirical evaluation, not a derivation from first principles. ALPET's claimed advantage is measured through experiments on three Wikipedia datasets, and the reported F1 scores and label-efficiency reductions are the output of those measurements rather than analytical consequences of the method's definitions. The only self-referential element is the use of CCW [4], the authors' own prior model, as the baseline; Section 5.3 states that this baseline was adapted by adding an AL step and by restricting inputs to sentence text and labels, omitting the adjacent-sentence context that defines the published CCW. This is a legitimate concern about whether the comparison supports the claim of outperforming the existing CCW model, but it is a threat to external validity or fair benchmarking, not circularity: ALPET's results are not constructed to equal the baseline or to reproduce its own inputs. Similarly, while H3's stated 58-72% reduction range coincidentally matches the later reported means, this is a hypothesis-after-results concern rather than a mathematical identity or a fitted parameter renamed as a prediction. There is no equation in which a predicted quantity is defined in terms of the outcome, no fitted parameter that is then reported as a discovery, and no load-bearing argument whose only support is a self-citation. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- Cosine similarity threshold for duplicate removal =
0.8
- PET pattern/verbalizer selection =
5 patterns per language, best-performing reported
assumptions (3)
- domain assumption The pre-existing labels in the CCW datasets are treated as a perfect oracle for simulating active learning.
- domain assumption The CCW baseline modified with an AL step is a faithful proxy for the published CCW model.
- domain assumption mBERT sentence embeddings capture enough semantic structure in these three languages for distance-based and K-Means AL strategies to be meaningful.
Cite this review
Pith. "Pith review of ALPET: Active Few-shot Learning for Citation Worthiness Detection in Low-Resource Wikipedia Languages." pith.science (2026). https://pith.science/paper/7J2PPLM4
@misc{pith2026250203292,
author = {Pith},
title = {Pith review of: ALPET: Active Few-shot Learning for Citation Worthiness Detection in Low-Resource Wikipedia Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/7J2PPLM4}},
note = {Machine review of arXiv:2502.03292}
}
read the original abstract
Citation Worthiness Detection (CWD) consists in determining which sentences, within an article or collection, should be backed up with a citation to validate the information it provides. This study, introduces ALPET, a framework combining Active Learning (AL) and Pattern-Exploiting Training (PET), to enhance CWD for languages with limited data resources. Applied to Catalan, Basque, and Albanian Wikipedia datasets, ALPET outperforms the existing CCW baseline while reducing the amount of labeled data in some cases above 80\%. ALPET's performance plateaus after 300 labeled samples, showing it suitability for low-resource scenarios where large, labeled datasets are not common. While specific active learning query strategies, like those employing K-Means clustering, can offer advantages, their effectiveness is not universal and often yields marginal gains over random sampling, particularly with smaller datasets. This suggests that random sampling, despite its simplicity, remains a strong baseline for CWD in constraint resource environments. Overall, ALPET's ability to achieve high performance with fewer labeled samples makes it a promising tool for enhancing the verifiability of online content in low-resource language settings.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[4]
Halitaj, A
A. Halitaj, A. Zubiaga, Providing citations to support fact-checking: Contextualizing detection of sentences needing citation on small wikipedias, Natural Language Processing Journal (2024) 100093
2024
-
[1]
Automated Fact Checking: Task formulations, methods and future directions
J. Thorne, A. Vlachos, Automated fact checking: Task formulations, methods and future directions, arXiv preprint arXiv:1806.07687
-
[2]
Z. Guo, M. Schlichtkrull, A. Vlachos, A survey on automated fact-checking, Transactions of the Association for Computational Linguistics 10 (2022) 178–206
2022
-
[3]
X. Zeng, A. S. Abumansour, A. Zubiaga, Automated fact-checking: A survey, Language and Linguistics Compass 15 (10) (2021) e12438
2021
- [5]
- [6]
-
[7]
X-FACT: A New Benchmark Dataset for Multilingual Fact Checking
A. Gupta, V . Srikumar, X-fact: A new benchmark dataset for multilingual fact checking, arXiv preprint arXiv:2106.09248
-
[8]
Konstantinovskiy, O
L. Konstantinovskiy, O. Price, M. Babakar, A. Zubiaga, Toward automated factchecking: Developing an annotation schema and benchmark for consistent automated claim detection, Digital threats: research and practice 2 (2) (2021) 1–16
2021
Show all 100 references
-
[9]
G. K. Shahi, D. Nandini, Fakecovid–a multilingual cross-domain fact check news dataset for covid-19, arXiv preprint arXiv:2006.11343
2006 arXiv
-
[10]
M. Redi, B. Fetahu, J. Morgan, D. Taraborelli, Citation needed: A taxonomy and algorithmic assessment of wikipedia’s verifiability, in: The World Wide Web Conference, 2019, pp. 1567–1578
2019
-
[11]
Y . Wang, Q. Yao, J. T. Kwok, L. M. Ni, Generalizing from a few examples: A survey on few-shot learning, ACM computing surveys (csur) 53 (3) (2020) 1–34
2020
-
[12]
Bonab, H
H. Bonab, H. Zamani, E. Learned-Miller, J. Allan, Citation worthiness of sentences in scientific reports, in: The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, 2018, pp. 1061–1064
2018
-
[13]
Sathe, S
A. Sathe, S. Ather, T. M. Le, N. Perry, J. Park, Automated fact-checking of claims from wikipedia, in: Proceedings of the Twelfth Language Resources and Evaluation Conference, 2020, pp. 6874–6882
2020
-
[14]
Baigutanova, J
A. Baigutanova, J. Myung, D. Saez-Trumper, A.-J. Chou, M. Redi, C. Jung, M. Cha, Longitudinal assessment of reference quality on wikipedia, in: Proceedings of the ACM Web Conference 2023, 2023, pp. 2831–2839
2023
-
[15]
Wright, I
D. Wright, I. Augenstein, Claim check-worthiness detection as positive unlabelled learning, arXiv preprint arXiv:2003.02736
2003 arXiv
-
[16]
Settles, Active learning literature survey
B. Settles, Active learning literature survey
-
[17]
Angluin, Queries and concept learning, Machine learning 2 (1988) 319–342
D. Angluin, Queries and concept learning, Machine learning 2 (1988) 319–342
1988
-
[18]
D. Cohn, L. Atlas, R. Ladner, Improving generalization with active learning, Machine learning 15 (1994) 201–221
1994
-
[19]
D. D. Lewis, W. A. Gale, A sequential algorithm for training text classifiers, 1994. URL https://api.semanticscholar.org/CorpusID:260481767
1994
-
[20]
Brinker, Incorporating diversity in active learning with support vector machines, in: Proceedings of the 20th international conference on machine learning (ICML-03), 2003, pp
K. Brinker, Incorporating diversity in active learning with support vector machines, in: Proceedings of the 20th international conference on machine learning (ICML-03), 2003, pp. 59–66
2003
-
[21]
Reichart, K
R. Reichart, K. Tomanek, U. Hahn, A. Rappoport, Multi-task active learning for linguistic annotations, in: Proceedings of ACL-08: HLT, 2008, pp. 861–869
2008
-
[22]
Acharya, R
A. Acharya, R. J. Mooney, J. Ghosh, Active multitask learning using both latent and supervised shared topics, in: Proceedings of the 2014 SIAM International Conference on Data Mining, SIAM, 2014, pp. 190–198
2014
-
[23]
Harpale, Y
A. Harpale, Y . Yang, Active learning for multi-task adaptive filtering
-
[24]
E. B. Baum, K. Lang, Query learning can work poorly when a human oracle is used, in: International joint conference on neural networks, V ol. 8, Beijing China, 1992, p. 8
1992
-
[25]
R. D. King, K. E. Whelan, F. M. Jones, P. G. Reiser, C. H. Bryant, S. H. Muggleton, D. B. Kell, S. G. Oliver, Functional genomic hypothesis generation and experimentation by a robot scientist, Nature 427 (6971) (2004) 247–252
2004
-
[26]
R. D. King, J. Rowland, S. G. Oliver, M. Young, W. Aubrey, E. Byrne, M. Liakata, M. Markham, P. Pir, L. N. Soldatova, et al., The automation of science, Science 324 (5923) (2009) 85–89
2009
-
[27]
Schumann, I
R. Schumann, I. Rehbein, Active learning via membership query synthesis for semi-supervised sentence classification, in: Proceedings of the 23rd conference on computational natural language learning (CoNLL), 2019, pp. 472–481
2019
-
[28]
Quteineh, S
H. Quteineh, S. Samothrakis, R. Sutcli ffe, Textual data augmentation for e fficient active learning on tiny datasets, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 7400–7410
2020
-
[29]
Zhang, E
Z. Zhang, E. Strubell, E. Hovy, A survey of active learning for natural language processing, arXiv preprint arXiv:2210.10109
-
[30]
Cacciarelli, M
D. Cacciarelli, M. Kulahci, Active learning for data streams: a survey, Machine Learning 113 (1) (2024) 185–239
2024
-
[31]
Dagan, S
I. Dagan, S. P. Engelson, Committee-based sampling for training probabilistic classifiers, in: Machine Learning Proceedings 1995, Elsevier, 1995, pp. 150–157
1995
-
[32]
Van Tran, T
C. Van Tran, T. T. Nguyen, D. T. Hoang, D. Hwang, N. T. Nguyen, Active learning-based approach for named entity recognition on short text streams, in: Multimedia and Network Information Systems: Proceedings of the 10th International Conference MISSI 2016, Springer, 2017, pp. 3...
2016
-
[33]
Smailovi ´c, M
J. Smailovi ´c, M. Gr ˇcar, N. Lavraˇc, M. ˇZnidarˇsiˇc, Stream-based active learning for sentiment analysis in the financial domain, Information sciences 285 (2014) 181–203
2014
-
[34]
Kranjc, J
J. Kranjc, J. Smailovi ´c, V . Podpeˇcan, M. Grˇcar, M. ˇZnidarˇsiˇc, N. Lavraˇc, Active learning for sentiment analysis on data streams: Methodol- ogy and workflow implementation in the clowdflows platform, Information Processing & Management 51 (2) (2015) 187–203
2015
-
[35]
Settles, Synthesis Lectures on Artificial Intelligence and Machine Learning, Morgan & Claypool Publishers, 2012
B. Settles, Synthesis Lectures on Artificial Intelligence and Machine Learning, Morgan & Claypool Publishers, 2012
2012
-
[36]
S. Tong, D. Koller, Support vector machine active learning with applications to text classification, Journal of machine learning research 2 (Nov) (2001) 45–66
2001
-
[37]
Zhang, M
Y . Zhang, M. Lease, B. Wallace, Active discriminative text representation learning, in: Proceedings of the AAAI conference on artificial intelligence, V ol. 31, 2017
2017
-
[38]
Siddhant, Z
A. Siddhant, Z. C. Lipton, Deep bayesian active learning for natural language processing: Results of a large-scale empirical study, arXiv preprint arXiv:1808.05697
-
[39]
Zhang, S
Y . Zhang, S. Feng, C. Tan, Active example selection for in-context learning, arXiv preprint arXiv:2211.04486
-
[40]
G. Tur, D. Hakkani-T ¨ur, R. E. Schapire, Combining active and semi-supervised learning for spoken language understanding, Speech Com- munication 45 (2) (2005) 171–186
2005
-
[41]
Radmard, Y
P. Radmard, Y . Fathullah, A. Lipani, Subsequence based deep active learning for named entity recognition, in: Proceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (...
2021
-
[42]
Tsvigun, A
A. Tsvigun, A. Shelmanov, G. Kuzmin, L. Sanochkin, D. Larionov, G. Gusev, M. Avetisian, L. Zhukov, Towards computationally feasible deep active learning, arXiv preprint arXiv:2205.03598
-
[43]
B. F. Dossou, A. L. Tonja, O. Yousuf, S. Osei, A. Oppong, I. Shode, O. O. Awoyomi, C. C. Emezue, Afrolm: A self-active learning-based multilingual pretrained language model for 23 african languages, arXiv preprint arXiv:2211.03263
-
[44]
Stratos, M
K. Stratos, M. Collins, Simple semi-supervised pos tagging, in: Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing, 2015, pp. 79–87
2015
-
[45]
Ringger, P
E. Ringger, P. McClanahan, R. Haertel, G. Busby, M. Carmen, J. Carroll, K. Seppi, D. Lonsdale, Active learning for part-of-speech tagging: Accelerating corpus annotation, in: Proceedings of the Linguistic Annotation Workshop, 2007, pp. 101–108
2007
-
[46]
Mendonc ¸a, A
V . Mendonc ¸a, A. Sardinha, L. Coheur, A. L. Santos, Query strategies, assemble! active learning with expert advice for low-resource natural language processing, in: 2020 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), IEEE, 2020, pp. 1–8
2020
-
[47]
J. Zhu, H. Wang, T. Yao, B. K. Tsou, Active learning with sampling by uncertainty and density for word sense disambiguation and text classification, in: 22nd International Conference on Computational Linguistics, Coling 2008, 2008, pp. 1137–1144
2008
-
[48]
Dligach, M
D. Dligach, M. Palmer, Good seed makes a good crop: accelerating active learning using language modeling, in: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 2011, pp. 6–10
2011
-
[49]
Alagi ´c, J
D. Alagi ´c, J. ˇSnajder, Experiments on active learning for croatian word sense disambiguation, in: The 5th Workshop on Balto-Slavic Natural Language Processing, 2015, pp. 49–58
2015
-
[50]
Y . Zhao, R. H. Zhang, S. Zhou, Z. Zhang, Active learning approaches to enhancing neural machine translation, in: Findings of the Associa- tion for Computational Linguistics: EMNLP 2020, 2020, pp. 1796–1806
2020
-
[51]
E. A. Chimoto, B. A. Bassett, Comet-qe and active learning for low-resource machine translation, arXiv preprint arXiv:2210.15696
-
[52]
Grießhaber, J
D. Grießhaber, J. Maucher, N. T. Vu, Fine-tuning bert for low-resource natural language understanding via active learning, arXiv preprint arXiv:2012.02462
2012 arXiv
-
[53]
K. Qian, Y . Sang, F. F. Bayat, A. Belyi, X. Chu, Y . Govind, S. Khorshidi, R. Khot, K. Luna, A. Nikfarjam, et al., Ape: Active learning-based tooling for finding informative few-shot examples for llm-based entity matching, arXiv preprint arXiv:2408.04637
-
[54]
Mamooler, R
S. Mamooler, R. Lebret, S. Massonnet, K. Aberer, An e fficient active learning pipeline for legal text classification, arXiv preprint arXiv:2211.08112
-
[55]
Imamura, Y
M. Imamura, Y . Takayama, N. Kaji, M. Toyoda, M. Kitsuregawa, A combination of active learning and semi-supervised learning starting with positive and unlabeled examples for word sense disambiguation: an empirical study on japanese web search query, in: Proceedings of the ACL-...
2009
-
[56]
Tsvigun, I
A. Tsvigun, I. Lysenko, D. Sedashov, I. Lazichny, E. Damirov, V . Karlov, A. Belousov, L. Sanochkin, M. Panov, A. Panchenko, et al., Active learning for abstractive text summarization, arXiv preprint arXiv:2301.03252
-
[57]
Brantley, A
K. Brantley, A. Sharaf, H. Daum ´e III, Active imitation learning with noisy guidance, arXiv preprint arXiv:2005.12801
2005 arXiv
-
[58]
K. Qian, P. C. Raman, Y . Li, L. Popa, Learning structured representations of entity names using active learning and weak supervision, arXiv preprint arXiv:2011.00105
2011 arXiv
-
[59]
Zhang, Y
R. Zhang, Y . Yu, P. Shetty, L. Song, C. Zhang, Prboost: Prompt-based rule discovery and boosting for interactive weakly-supervised learning, arXiv preprint arXiv:2203.09735
-
[60]
Z. L. Zhu, V . Yadav, Z. Afzal, G. Tsatsaronis, Few-shot initializing of active learner via meta-learning, in: Findings of the Association for Computational Linguistics: EMNLP 2022, 2022, pp. 1117–1133
2022
-
[61]
M ¨uller, G
T. M ¨uller, G. P´erez-Torr´o, A. Basile, M. Franco-Salvador, Active few-shot learning with fasl, in: International Conference on Applications of Natural Language to Information Systems, Springer, 2022, pp. 98–110
2022
-
[62]
X. Zeng, A. Zubiaga, Active pets: active data annotation prioritisation for few-shot claim verification with pattern exploiting training, arXiv preprint arXiv:2208.08749
-
[63]
Bayer, C
M. Bayer, C. Reuter, Activellm: Large language model-based active learning for textual few-shot scenarios, arXiv preprint arXiv:2405.10808
-
[64]
Y . Gal, Z. Ghahramani, Dropout as a bayesian approximation: Representing model uncertainty in deep learning, in: international conference on machine learning, PMLR, 2016, pp. 1050–1059
2016
-
[65]
Citovsky, G
G. Citovsky, G. DeSalvo, C. Gentile, L. Karydas, A. Rajagopalan, A. Rostamizadeh, S. Kumar, Batch active learning at scale, Advances in Neural Information Processing Systems 34 (2021) 11933–11944
2021
-
[66]
Beatty, E
G. Beatty, E. Kochis, M. Bloodgood, Impact of batch size on stopping active learning for text classification, in: 2018 IEEE 12th International 23 1–24 24 Conference on Semantic Computing (ICSC), IEEE, 2018, pp. 306–307
2018
-
[67]
Ananthakrishnan, R
S. Ananthakrishnan, R. Prasad, D. Stallard, P. Natarajan, A semi-supervised batch-mode active learning strategy for improved statistical machine translation, in: Proceedings of the Fourteenth Conference on Computational Natural Language Learning, 2010, pp. 126–134
2010
-
[68]
T. Shi, A. Benton, I. Malioutov, O. Irsoy, Diversity-aware batch active learning for dependency parsing, arXiv preprint arXiv:2104.13936
-
[69]
Farinneya, M
P. Farinneya, M. M. A. Pour, S. Hamidian, M. Diab, Active learning for rumor identification on social media, in: Findings of the association for computational linguistics: EMNLP 2021, 2021, pp. 4556–4565
2021
-
[70]
D. Shen, J. Zhang, J. Su, G. Zhou, C. L. Tan, Multi-criteria-based active learning for named entity recognition, in: Proceedings of the 42nd annual meeting of the Association for Computational Linguistics (ACL-04), 2004, pp. 589–596
2004
-
[71]
S. C. Hoi, R. Jin, M. R. Lyu, Large-scale text categorization by batch mode active learning, in: Proceedings of the 15th international conference on World Wide Web, 2006, pp. 633–642
2006
-
[72]
J. T. Ash, C. Zhang, A. Krishnamurthy, J. Langford, A. Agarwal, Deep batch active learning by diverse, uncertain gradient lower bounds, arXiv preprint arXiv:1906.03671
1906 arXiv
-
[73]
Ikhwantri, S
F. Ikhwantri, S. Louvan, K. Kurniawan, B. Abisena, V . Rachman, A. F. Wicaksono, R. Mahendra, Multi-task active learning for neural semantic role labeling on low resource conversational corpus, arXiv preprint arXiv:1806.01523
-
[74]
Rotman, R
G. Rotman, R. Reichart, Multi-task active learning for pre-trained transformer-based models, Transactions of the Association for Computa- tional Linguistics 10 (2022) 1209–1228
2022
-
[75]
B. Zhou, X. Cai, Y . Zhang, W. Guo, X. Yuan, Mtaal: multi-task adversarial active learning for medical named entity recognition and normalization, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 35, 2021, pp. 14586–14593
2021
-
[76]
H. Zhu, W. Ye, S. Luo, X. Zhang, A multitask active learning framework for natural language understanding, in: Proceedings of the 28th International Conference on Computational Linguistics, 2020, pp. 4900–4914
2020
-
[77]
Devlin, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805
J. Devlin, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805
-
[78]
Radford, K
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al., Improving language understanding by generative pre-training
-
[79]
Ra ffel, N
C. Ra ffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, P. J. Liu, Exploring the limits of transfer learning with a unified text-to-text transformer, Journal of machine learning research 21 (140) (2020) 1–67
2020
-
[80]
L. E. Dor, A. Halfon, A. Gera, E. Shnarch, L. Dankin, L. Choshen, M. Danilevsky, R. Aharonov, Y . Katz, N. Slonim, Active learning for bert: an empirical study, in: Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), 2020, pp. 7949–7962
2020
-
[81]
B. Yao, I. Jindal, L. Popa, Y . Katsis, S. Ghosh, L. He, Y . Lu, S. Srivastava, J. A. Hendler, D. Wang, Beyond labels: Empowering human with natural language explanations through a novel active-learning architecture, Association for Computational Linguistics, 2023
2023
-
[82]
Y . Lu, B. Yao, S. Zhang, Y . Wang, P. Zhang, T. Lu, T. J.-J. Li, D. Wang, Human still wins over llm: An empirical study of active learning on domain-specific annotation tasks, arXiv preprint arXiv:2311.09825
-
[83]
D. Li, Z. Wang, Y . Chen, R. Jiang, W. Ding, M. Okumura, A survey on deep active learning: Recent advances and new frontiers, IEEE Transactions on Neural Networks and Learning Systems
-
[84]
Zeng, Few-shot claim verification for automated fact checking, Ph.D
X. Zeng, Few-shot claim verification for automated fact checking, Ph.D. thesis, Queen Mary University of London (2024)
2024
-
[85]
Kasai, K
J. Kasai, K. Qian, S. Gurajada, Y . Li, L. Popa, Low-resource deep entity resolution with transfer and active learning, arXiv preprint arXiv:1906.08042
1906 arXiv
-
[86]
Z. Zhou, A. Waibel, Active learning for massively parallel translation of constrained text into low resource languages, arXiv preprint arXiv:2108.07127
-
[87]
¨Ohman, Active learning for named entity recognition with swedish language models (2021)
J. ¨Ohman, Active learning for named entity recognition with swedish language models (2021)
2021
-
[88]
Maekawa, D
S. Maekawa, D. Zhang, H. Kim, S. Rahman, E. Hruschka, Low-resource interactive active labeling for fine-tuning language models, in: Findings of the Association for Computational Linguistics: EMNLP 2022, 2022, pp. 3230–3242
2022
-
[89]
Z. Ye, D. Liu, K. Pavani, S. Dasgupta, Lamm: Language aware active learning for multilingual models, in: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, 2023, pp. 5255–5256
2023
-
[90]
Barnab `o, F
G. Barnab `o, F. Siciliano, C. Castillo, S. Leonardi, P. Nakov, G. Da San Martino, F. Silvestri, Deep active learning for misinformation detection using geometric deep learning, Online Social Networks and Media 33 (2023) 100244
2023
-
[91]
Margatina, G
K. Margatina, G. Vernikos, L. Barrault, N. Aletras, Active learning by acquiring contrastive examples, arXiv preprint arXiv:2109.03764
-
[92]
Sener, S
O. Sener, S. Savarese, Active learning for convolutional neural networks: A core-set approach, arXiv preprint arXiv:1708.00489
-
[93]
Schr ¨oder, L
C. Schr ¨oder, L. M ¨uller, A. Niekler, M. Potthast, Small-text: Active learning for text classification in python, in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Association for Computa- ...
2023
-
[94]
Bachem, M
O. Bachem, M. Lucic, A. Krause, Scalable k-means clustering via lightweight coresets, in: Proceedings of the 24th ACM SIGKDD Interna- tional Conference on Knowledge Discovery & Data Mining, 2018, pp. 1119–1127
2018
-
[95]
Lesci, A
P. Lesci, A. Vlachos, Anchoral: Computationally e fficient active learning for large and imbalanced datasets, arXiv preprint arXiv:2404.05623
-
[96]
Holub, P
A. Holub, P. Perona, M. C. Burl, Entropy-based active learning for object recognition, in: 2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, IEEE, 2008, pp. 1–8
2008
-
[97]
D. D. Lewis, A sequential algorithm for training text classifiers: Corrigendum and additional data, in: Acm Sigir Forum, V ol. 29, ACM New York, NY , USA, 1995, pp. 13–19
1995
-
[98]
T. Luo, K. Kramer, D. B. Goldgof, L. O. Hall, S. Samson, A. Remsen, T. Hopkins, D. Cohn, Active learning to recognize multiple types of plankton., Journal of Machine Learning Research 6 (4)
-
[99]
Yuan, H.-T
M. Yuan, H.-T. Lin, J. Boyd-Graber, Cold-start active learning through self-supervised language modeling, arXiv preprint arXiv:2010.09535
2010 arXiv
-
[100]
Schick, H
T. Schick, H. Sch ¨utze, Exploiting cloze questions for few shot text classification and natural language inference, arXiv preprint arXiv:2001.07676. 24
2001 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.