REVIEW 4 major objections 5 minor 100 references
A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Fine-tuned classifiers built from $40 of LLM pseudo-labels match GPT-4 on wildlife ad filtering, with F1 up to 0.94.
desk verdict A cost-effective idea with a real application, but the evaluation likely leaks the gold labels into model selection, so the headline comparisons need a clean holdout before they can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the LTS (Learn to Sample) sampling loop. First, topic-model clustering partitions the unlabeled ad collection into diverse groups. Second, Thompson sampling treats each cluster as a bandit arm, drawing a Beta-distributed score per cluster and mining the winning cluster, with rewards defined by whether fine-tuning a small classifier on the newly labeled ads improves validation F1. Third, a few-shot LLM prompt — the same kind of instruction, examples, and rationale a domain expert would write for a human annotator — pseudo-labels a small batch of ads from the chosen cluster. Fourth, a compact text or text-image classifier is fine-tuned on the accumulated labeled ads; its predictions guide the next batch's selection, and its validation score updates the bandit's win/loss counts. The loop continues until the classifier meets a performance threshold or the labeling budget is exhausted, converting a labeling-cost problem into a sampling problem.
What would settle it
Run GPT-4 with the paper's few-shot prompts over the human-labeled validation sets for the three tasks and compute label agreement; if agreement falls well below the ~95% accuracy implied by the reported F1 scores, or if disagreement clusters on exactly the edge cases the prompts exclude (fossil teeth, single teeth, faux leather), then the pseudo-label noise is baked into the trained classifiers and the reported F1 gains overstate real-world performance.
Extended reading notes
Core claim
The paper's central claim is that high-quality classifiers for wildlife ad filtering do not require large-scale human annotation or large-scale LLM labeling. Roughly 2,000–4,000 ads, selected not at random but by a clustering and bandit-guided active-learning loop, and labeled by GPT-4 through few-shot prompts, are sufficient to fine-tune compact classifiers: text-only LTS classifiers reach F1 0.816 on animal products and 0.855 on leather, versus GPT-4's 0.827 and 0.850, and 0.872 on shark products versus GPT-4's 0.946, while a text-image LTS model reaches F1 0.940 for leather. The paper further claims that the sampling strategy is the key ingredient: classifiers trained on samples chosen by LTS outperform those trained on random samples, keyword-biased samples, or a coreset active-learning baseline, and that the choice of LLM used to generate pseudo-labels is the dominant factor in final classifier quality.
Load-bearing premise
The pipeline assumes GPT-4's few-shot labels are accurate enough to serve as ground truth for training, yet the paper never measures how those pseudo-labels compare with the roughly 500 human expert labels collected for each task.
Editorial extensions
If this is right
- A one-time sampling cost of about $40 yields a reusable classifier for a research question, whereas labeling the full collection with GPT-4 would cost between $320 and $14,000 for the three tasks studied.
- Classifier quality is dominated by the pseudo-label source: with GPT-4 labels, LTS-trained text models reach F1 between 0.82 and 0.87, but switching to an 8-billion-parameter open model drops F1 to between 0.40 and 0.62.
- LTS beats random sampling, keyword-based sampling, and a coreset active-learning baseline on the leather and animal-product tasks, while remaining far cheaper: the coreset baseline timed out after 72 hours on the 700k-ad animal-products collection, where LTS finished in under 5 hours.
- Text-plus-image classifiers help for small leather goods (F1 0.94) but not for shark products, where images did not improve over text-only classification.
- The same data-curation recipe is intended to transfer to other online data triage tasks, including identifying advertisements linked to human trafficking and illegal gun sales.
Reading between the lines
- Beyond the paper: the pipeline never measures GPT-4 pseudo-label agreement with the 500 human labels per task, so a systematic LLM bias (e.g., on fossil teeth or faux leather) would be baked into the trained classifiers and inflate apparent performance.
- Beyond the paper: a direct comparison between an LTS-trained classifier and a classifier trained on the same number of human labels would quantify the real cost of pseudo-label noise.
- Beyond the paper: the same bandit-guided active-learning loop should transfer to other rare-class text triage tasks, such as detecting ads for human trafficking or illegal weapons.
- Beyond the paper: deploying the cheap classifier as a daily monitor, verifying only its positive predictions with the LLM, would turn a one-time $40 labeling cost into an ongoing trafficking surveillance system.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LTS (Learn to Sample), a pipeline that combines clustering, Thompson sampling, and active learning to select a small, diverse set of ads from online marketplaces, uses GPT-4 few-shot prompting to pseudo-label those ads, and then fine-tunes small BERT-based classifiers on the resulting labeled data. The method is evaluated on three wildlife-trafficking classification tasks: shark products, small leather products, and animal products, with approximately 500 human expert labels per task used for evaluation. The paper reports F1 scores up to about 0.94, claims that LTS-derived classifiers are comparable to or better than zero-shot GPT-4 at a fraction of the cost (roughly $40 in API labeling costs), and presents two domain use cases in criminology and environmental science. The central claim is that resource-constrained researchers can build specialized wildlife-trafficking classifiers without large-scale manual labeling.
Significance. If the evaluation protocol is sound, the paper makes a useful applied contribution: it offers a practical, low-cost recipe for bootstrapping specialized classifiers with LLM pseudo-labels, and it is one of the few studies in this domain to compare against random sampling, keyword-based sampling, a standard active-learning baseline, and several LLM baselines on three real tasks with human gold labels. The open-source code, real market data, and concrete cost figures are genuine strengths. The significance is proportionate: the method is not theoretically surprising, but it could enable broader monitoring of online wildlife trafficking by groups without large labeling budgets. However, the main claim of 'outperforming LLMs' is currently overstated relative to Table 5, and the evaluation may be compromised by overlap between the validation labels used inside LTS and the test labels used for final reporting.
major comments (4)
- [§3.4 (Retrain Model) and §4.1 (Experimental Setup)] The same set of approximately 500 human-labeled ads appears to be used both as the validation gold data that drives the Thompson-sampling rewards, base-model selection, and stopping rule in §3.4, and as the evaluation data that produces the F1 scores in Tables 2 and 5. The paper describes no separate held-out test split and no nested resampling procedure. As a result, LTS is rewarded during development for agreeing with the exact labels on which its performance is later reported, while the RS, KBS, ACTL, and zero-shot GPT-4 baselines never see those gold labels during development. This selection-on-validation risk can inflate LTS's apparent advantage and must be addressed, for example by holding out an untouched test set for final reporting or by reporting nested cross-validation estimates.
- [Abstract and §4.3] The abstract claims that the classifiers 'outperform LLMs at a lower cost', but Table 5 shows that LTS-text is below GPT-4 on Animal Products (0.816 vs 0.827) and Shark Products (0.873 vs 0.946), and only 0.005 above GPT-4 on Leather Products (0.855 vs 0.850). The data support a weaker claim: comparable performance to GPT-4 on two tasks and slightly better on one, at a much lower labeling and inference cost. The abstract and Section 6 should be revised to avoid overstating the accuracy comparison.
- [Tables 2, 3, 5, and 6] All reported precision, recall, F1, and accuracy numbers are single-run point estimates with no confidence intervals or standard deviations, despite the fact that LTS involves stochastic components (Thompson sampling, clustering, fine-tuning) and that several reported differences between methods are small, e.g., 0.005 on Leather Products in Table 5 and 0.052 on Sharks in Table 3. Without repeated runs, bootstrap intervals, or per-seed variability, the comparisons against GPT-4 and ACTL are not statistically supported. Please report variability across seeds or iterations.
- [§3.4 (LLM Pseudo-Labeling) and Table 6] The paper never measures the agreement between the GPT-4 pseudo-labels and the human expert labels, even though the 500 human labels exist for each task. Table 6 shows that pseudo-label quality is the main performance driver, with F1 dropping from 0.816 to 0.40 when smaller LLMs generate the labels. Systematic GPT-4 errors on difficult cases such as fossil teeth or synthetic leather would therefore propagate directly into the trained classifiers. Reporting pseudo-label accuracy or agreement per task, ideally stratified by these known tricky cases, would provide direct evidence for the central assumption that LLM pseudo-labels are reliable enough to train on.
minor comments (5)
- [Figure 3, Animal Products prompt] The fifth example is duplicated: the text reads '5. Advertisement: 5. Advertisement: 1/10X Wholesale ...', and the numbering then repeats '5. Advertisement: <advertisement_title>'. This should be cleaned up, as the prompt is copied verbatim into the pseudo-labeling pipeline.
- [§3.4, Clustering] The sentence 'For our current implementation, use used a topic modeling method [79]' contains a typo ('use used') and should read 'we used'.
- [Tables 2 and 3] The ACTL timeout is reported inconsistently: the Table 2 footnote says 'After 78 hours only 3,000 training examples', while Table 3 reports '72*' for the same Animal Products run. Please reconcile these numbers.
- [Figures 5 and 6] Both figures contain unresolved 'Figure ??' cross-references in the text of Section 5.1; the figures should be numbered and cited properly.
- [§4.1, validation data] The paper does not describe how the approximately 500 gold labels per task were produced, such as the number of annotators, their expertise, or inter-annotator agreement. A brief description of the annotation protocol would help readers judge label noise.
Circularity Check
LTS's reported F1 scores are computed on the same validation gold data that drives its Thompson-sampling rewards and model accept/revert decisions.
-
fitted input called prediction
[Section 3.4 (Retrain Model) and Section 4.1 (Data Collection and Validation Data); Tables 2 and 5]
""After training, we evaluate the model’s performance against the validation gold data. The results are used to decide how to update the Thompson sampling parameters, specifically the counts of wins and losses." ... "To evaluate the derived models, we used approximately 500 ads manually labeled by domain experts for each research question.""
The only gold-labeled data described in the paper is this 'validation gold data' of about 500 expert-labeled ads per task. LTS uses it every iteration to compute the Thompson-sampling reward and to decide whether to keep or revert the base model, then the paper reports F1 scores on the same data in Tables 2 and 5. No separate held-out test split or nested resampling is described, so the headline F1 is the quantity LTS actively optimizes through cluster selection and model acceptance, not an independent out-of-sample prediction. Random/KBS/ACTL and zero-shot GPT-4 never see the gold labels during development, so the comparison is biased in LTS's favor by construction.
full rationale
The LLM pseudo-labeling/training loop itself is not circular: GPT-4 generates training labels, BERT is fine-tuned on those labels, and the final evaluation uses human expert labels, so the classifier is not trained directly on the test labels. The cost analysis (2,000-4,000 GPT-4 calls) and the use-case analyses are also independent of any circularity. However, the paper's only gold labels serve both as LTS's validation/reward signal and as the evaluation set for the reported F1 scores. Because LTS's cluster selection and base-model decisions are rewarded for improving F1 on that exact set, the Table 2 and Table 5 numbers are optimistically biased and are not true held-out predictions. This is a partial circularity rather than a complete reduction: the diversity clustering and cost-control components are not derived from the validation labels, and the pseudo-label quality result (Table 6) remains informative. A separate held-out test set or nested resampling would remove this issue; absent that, the central comparative claim is partly forced by the evaluation design.
Assumptions & free parameters
free parameters (7)
- Cluster count K =
not reported
- Samples per iteration N =
200
- Thompson sampling decay factor delta =
0.99
- Initial Beta parameters for Thompson sampling =
not reported
- Maximum positive samples per cluster =
not reported
- Baseline F1 and stopping threshold =
0.5 initially, then previous model F1; target not stated
- BERT and MLP grid-search hyperparameters =
selected by grid search
assumptions (6)
- standard math Thompson sampling with Beta posteriors converges to good cluster selection for stationary reward distributions.
- domain assumption GPT-4 few-shot labels from ad titles are sufficiently accurate to serve as pseudo-ground truth for training.
- domain assumption The approximately 500 manually labeled ads per task are correct and representative of each full ad collection.
- domain assumption Topic modeling over ad titles produces clusters whose diversity is useful for selecting representative ads.
- domain assumption The ACHE scoped crawl with seed species and product keywords approximates the population of wildlife ads on target marketplaces.
- ad hoc to paper Hand-written few-shot prompt examples in Figure 3 correctly encode each research question's selection criteria.
Cite this review
Pith. "Pith review of A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces." pith.science (2026). https://pith.science/paper/4SNFTWRT
@misc{pith2026250421211,
author = {Pith},
title = {Pith review of: A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/4SNFTWRT}},
note = {Machine review of arXiv:2504.21211}
}
read the original abstract
Wildlife trafficking remains a critical global issue, significantly impacting biodiversity, ecological stability, and public health. Despite efforts to combat this illicit trade, the rise of e-commerce platforms has made it easier to sell wildlife products, putting new pressure on wild populations of endangered and threatened species. The use of these platforms also opens a new opportunity: as criminals sell wildlife products online, they leave digital traces of their activity that can provide insights into trafficking activities as well as how they can be disrupted. The challenge lies in finding these traces. Online marketplaces publish ads for a plethora of products, and identifying ads for wildlife-related products is like finding a needle in a haystack. Learning classifiers can automate ad identification, but creating them requires costly, time-consuming data labeling that hinders support for diverse ads and research questions. This paper addresses a critical challenge in the data science pipeline for wildlife trafficking analytics: generating quality labeled data for classifiers that select relevant data. While large language models (LLMs) can directly label advertisements, doing so at scale is prohibitively expensive. We propose a cost-effective strategy that leverages LLMs to generate pseudo labels for a small sample of the data and uses these labels to create specialized classification models. Our novel method automatically gathers diverse and representative samples to be labeled while minimizing the labeling costs. Our experimental evaluation shows that our classifiers achieve up to 95% F1 score, outperforming LLMs at a lower cost. We present real use cases that demonstrate the effectiveness of our approach in enabling analyses of different aspects of wildlife trafficking.
Figures
Reference graph
Works this paper leans on
-
[1]
Umang Aggarwal, Adrian Popescu, and Céline Hudelot. 2021. Minority class oriented active learning for imbalanced datasets. In International Conference on Pattern Recognition (ICPR). IEEE, online, 9920–9927
2021
-
[2]
Juliana Barbosa, Sunandan Chakraborty, and Juliana Freire. 2024. A Flexible and Scalable Approach for Collecting Wildlife Advertisements on the Web. arXiv:2407.18898 [cs.IR] https://arxiv.org/abs/2407.18898
work page Pith review arXiv 2024
-
[3]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901
2020
-
[4]
Trang Bui and Lam Quan. 2021. Traffickers v. Covid-19. https://www. investigative.earth/traffickers-v-covid-19 . Accessed on 07/31/2024
2021
-
[5]
12 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces SIGMOD ’25, June 22-27, 2025, Berlin, Germany
Ana Sofia Cardoso, Sofiya Bryukhova, Francesco Renna, Luís Reino, Chi Xu, Zixiang Xiao, Ricardo Correia, Enrico Di Minin, Joana Ribeiro, and Ana Sofia Vaz. 12 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces SIGMOD ’25, June 22-27, 2025, Berlin, Germany
2025
-
[6]
Dominique Chabot, Seth Stapleton, and Charles M Francis. 2022. Using Web images to train a deep neural network to detect sparsely distributed wildlife in large volumes of remotely sensed imagery: A case study of polar bears on sea ice. Ecological Informatics 68 (2022), 101547
2022
-
[7]
Chengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li, and Yuyu Luo. 2022. Selec- tive data acquisition in the wild for model charging. Proceedings of the VLDB Endowment (pVLDB) 15, 7 (2022), 1466–1478
2022
-
[8]
Hasna Chamlal, Hajar Kamel, and Tayeb Ouaderhman. 2024. A hybrid multi- criteria meta-learner based classifier for imbalanced data. Knowledge-based systems 285 (2024), 111367
2024
Show all 100 references
-
[9]
Zixiang Chang. 2022. A survey of modern crawler methods. In Proceedings of the 6th International Conference on Control Engineering and Artificial Intelligence . IEEE, Online, 21–28
2022
-
[10]
Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. Advances in neural information processing systems 24 (2011), 2249 – 2257
2011
-
[11]
Sandra Charity and Juliana Machado Ferreira. 2020. Wildlife trafficking in Brazil. TRAFFIC International, Cambridge, United Kingdom 140 (2020)
2020
-
[12]
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer
-
[13]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations?. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.)...
2023
-
[14]
Cody Coleman, Edward Chou, Julian Katz-Samuels, Sean Culatana, Peter Bailis, Alexander C Berg, Robert Nowak, Roshan Sumbaly, Matei Zaharia, and I Zeki Yalniz. 2022. Similarity search for efficient active learning and search of rare concepts. In Proceedings of the AAAI Conferen...
2022
-
[15]
DARPA. 2023. Memex program. https://www.darpa.mil/program/memex. Accessed: 2025-01-12
2023
-
[16]
Nilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi, Duen Horng Chau, and Xu Chu. 2020. Goggles: Automatic image labeling with affinity coding. InProceedings of the ACM SIGMOD International Conference on Management of Data. ACM, 1717– 1732
2020
-
[17]
Elodie Demeau, Miguel Eduardo Vargas Monroy, and Jeffrey Karolan. 2019. Wildlife trafficking on the internet: a virtual market similar to drug traffick- ing? Revista Criminalidad 61, 2 (2019), 101–112
2019
-
[18]
Michael Desmond, Zahra Ashktorab, Qian Pan, Casey Dugan, and James M. Johnson. 2024. EvaluLLM: LLM assisted evaluation of generative outputs. In Companion Proceedings of the International Conference on Intelligent User Interfaces. ACM, 30–32
2024
-
[19]
Enrico Di Minin, Christoph Fink, Henrikki Tenkanen, and Tuomo Hiippala. 2018. Machine learning for tracking illegal wildlife trade on social media. Nature ecology & evolution 2, 3 (2018), 406–407
2018
-
[20]
Isabella Dominguez, Marjan Hindriks, Jordi Janssen, and Daan van Uhm. 2024. Online Illegal Trade in Reptiles in the Netherlands. In Criminal Justice, Wildlife Conservation and Animal Rights in the Anthropocene . Bristol University Press, Online, 52–69
2024
-
[21]
Nicholas K Dulvy, Nathan Pacoureau, Cassandra L Rigby, Riley A Pollom, Rima W Jabado, David A Ebert, Brittany Finucci, Caroline M Pollock, Jessica Cheok, Danielle H Derrick, et al. 2021. Overfishing drives over one-third of all sharks and rays toward a global extinction crisis...
2021
-
[22]
eBay Inc. 2023. eBay: Buy, Sell, and Save on Brands You Love. https://www. ebay.com. Accessed: 2024-10-14
2023
-
[23]
Environmental Investigation Agency. 2020. While you’ve been in lockdown, so have wildlife criminals – and many of them have been ‘working from home’. https://eia-international.org/news/while-youve-been-in-lockdown- so-have-wildlife-criminals-and-many-of-them-have-been-working-...
2020
-
[24]
Hugging Face. 2024. HuggingFace. https://huggingface.co/
2024
-
[25]
Ju Fan, Jianhong Tu, Guoliang Li, Peng Wang, Xiaoyong Du, Xiaofeng Jia, Song Gao, and Nan Tang. 2024. Unicorn: A Unified Multi-Tasking Matching Model. SIGMOD Rec. 53, 1 (May 2024), 44–53. doi: 10.1145/3665252.3665263
2024
-
[26]
Chenhao Fang, Xiaohan Li, Zezhong Fan, Jianpeng Xu, Kaushiki Nag, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. 2024. LLM-Ensemble: Optimal Large Language Model Ensemble Method for E-commerce Product Attribute Value Extraction. In Proceedings of the International ACM SIGIR...
2024
-
[27]
Lalita Gomez and Chris R Shepherd. 2019. Bearly on the radar–an analysis of seizures of bears in Indonesia. European Journal of Wildlife Research 65, 6 (2019), 89
2019
-
[28]
Anjana Gosain and Saanchi Sardana. 2017. Handling class imbalance problem using oversampling techniques: A review. In International conference on advances in computing, communications and informatics (ICACCI) . IEEE, 79–85
2017
-
[29]
Timothy C Haas and Sam M Ferreira. 2015. Federated databases and action- able intelligence: using social network analysis to disrupt transnational wildlife trafficking criminal networks. Security Informatics 4, 1 (2015), 1–14
2015
-
[30]
Lauren Harrington, David Macdonald, and Neil D’Cruze. 2019. Popularity of pet otters on YouTube: evidence of an emerging trade threat. Nature Conservation 36 (2019), 17–45
2019
-
[31]
Simone Haysom. 2019. In Search of Cyber-Enabled Disruption
2019
-
[32]
Geon Heo, Yuji Roh, Seonghyeon Hwang, Dayun Lee, and Steven Euijong Whang
-
[33]
Julio Hernandez-Castro and David L Roberts. 2015. Automatic detection of potentially illegal online sales of elephant ivory via data mining. PeerJ Computer Science 1 (2015), e10
2015
-
[34]
L Hingley. 2020. Conservation implications of land-based trophy shark fishing . Ph. D. Dissertation. Masters thesis. The Univeristy of Western Australia
2020
-
[35]
Sara Bronwen Hunter, Fiona Mathews, and Julie Weeds. 2023. Using hierarchical text classification to investigate the utility of machine learning in automating online analyses of wildlife exploitation. Ecological Informatics 75 (2023), 102076
2023
-
[36]
Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, and Dan Suciu. 2024. Chorus: Foundation Models for Unified Data Discovery and Exploration. Proceedings of the VLDB Endowment (pVLDB)17, 8 (2024), 2104–2114
2024
-
[37]
Burcu B Keskin, Emily C Griffin, Jonathan O Prell, Bistra Dilkina, Aaron Ferber, John MacDonald, Rowan Hilend, Stanley Griffis, and Meredith L Gore. 2022. Quantitative investigation of wildlife trafficking supply chains: A review. Omega 115, 102780 (2022), 102780
2022
-
[38]
Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan A, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts
-
[39]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimiza- tion. In Proceedings of the International Conference on Learning Representations (ICLR)
2014
-
[40]
Ritwik Kulkarni and Enrico Di Minin. 2023. Towards automatic detection of wildlife trade using machine vision models. Biological Conservation 279 (2023), 109924
2023
-
[41]
Larry Greenemeier. 2015. Human Traffickers Caught on Hidden Inter- net. https://www.scientificamerican.com/article/human-traffickers- caught-on-hidden-internet . Accessed: 2025-01-12
2015
-
[42]
Pietro Lesci and Andreas Vlachos. 2024. AnchorAL: Computationally Efficient Ac- tive Learning for Large and Imbalanced Datasets. InProceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume ...
2024
-
[43]
Wei-Chao Lin, Chih-Fong Tsai, Ya-Han Hu, and Jing-Shang Jhang. 2017. Clustering-based undersampling in class-imbalanced data. Information Sciences 409 (2017), 17–26
2017
-
[44]
Proceedings of the Conference on Empirical Methods in Natural Language Processing
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In "Proceedings of the Conference on Empirical Methods in Natural Language Processing". "2511–2522"
2023
-
[45]
Rowan O Martin, Cristiana Senni, and Neil C D’Cruze. 2018. Trade in wild- sourced African grey parrots: Insights via social media. Global Ecology and Conservation 15 (2018), e00429
2018
-
[46]
De-Yao Meng, Tao Li, Hao-Xuan Li, Mei Zhang, Kun Tan, Zhi-Pang Huang, Na Li, Rong-Hai Wu, Xiao-Wei Li, Ben-Hui Chen, et al. 2023. A method for auto- matic identification and separation of wildlife images using ensemble learning. Ecological Informatics 77 (2023), 102262
2023
-
[47]
Nguyen Minh and Madelon Willemsen. 2016. A rapid assessment of e-commerce wildlife trade in Viet Nam. TRAFFIC Bulletin 28, 2 (2016), 53
2016
-
[48]
Robert Munro Monarch. 2021. Human-in-the-Loop Machine Learning: Active learning and annotation for human-centered AI . Simon and Schuster
2021
-
[49]
Annika Mozer and Stefan Prost. 2023. An introduction to illegal wildlife trade and its effects on biodiversity and society. Forensic Science International: Animals and Environments 3 (2023), 100064
2023
-
[50]
Sravani Nalluri, S Jeevan Rishi Kumar, Manik Soni, Soheb Moin, and K Nikhil
-
[51]
Avanika Narayan, Ines Chami, Laurel Orr, and Christopher Ré. 2022. Can Foun- dation Models Wrangle Your Data? Proceedings of the VLDB Endowment (pVLDB) 16, 4 (2022), 738–746
2022
-
[52]
Gohar A Petrossian, Stephen F Pires, and Daan P van Uhm. 2016. An overview of seized illegal wildlife entering the United States. Global Crime 17, 2 (2016), 13 SIGMOD ’25, June 22-27, 2025, Berlin, Germany Juliana Barbosa et al. 181–201
2016
-
[53]
Jennifer M Pytka, Alec BM Moore, and Adel Heenan. 2023. Internet trade of a previously unknown wildlife product from a critically endangered marine fish. Conservation Science and Practice 5, 3 (2023), e12896
2023
-
[54]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[55]
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al . 2018. Improving language understanding by generative pre-training
2018
-
[56]
Anant Raj and Francis Bach. 2022. Convergence of Uncertainty Sampling for Ac- tive Learning. In Proceedings of the International Conference on Machine Learning , Vol. 162. 18310–18331
2022
-
[57]
Alexander J Ratner, Stephen H Bach, Henry R Ehrenberg, and Chris Ré. 2017. Snorkel: Fast training set generation for information extraction. In Proceedings of the ACM International Conference on Management of Data . ACM, 1683–1686
2017
-
[58]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . 3982–3992
2019
-
[59]
Salim Rezvani and Xizhao Wang. 2023. A broad review on class imbalance learning techniques. Applied Soft Computing 143 (2023), 110415
2023
-
[60]
Matthias C Rillig, Marlene Ågerstrand, Mohan Bi, Kenneth A Gould, and Uli Sauerland. 2023. Risks and benefits of large language models for the environment. Environmental Science & Technology 57, 9 (2023), 3464–3466
2023
-
[61]
David L Roberts, Katya Mun, and EJ Milner-Gulland. 2022. A systematic survey of online trade: trade in Saiga antelope horn on Russian-language websites. Oryx 56, 3 (2022), 352–359
2022
-
[62]
Arunabha M Roy, Jayabrata Bhaduri, Teerath Kumar, and Kislay Raj. 2023. WilDect-YOLO: An efficient and robust computer vision-based accurate object localization model for automated endangered wildlife detection. Ecological Infor- matics 75 (2023), 101919
2023
-
[63]
Amir Reza Salehi and Majid Khedmati. 2024. A cluster-based SMOTE both- sampling (CSBBoost) ensemble algorithm for classifying imbalanced data. Scien- tific Reports 14, 1 (2024), 5152
2024
-
[64]
Brett R Scheffers, Brunno F Oliveira, Ieuan Lamb, and David P Edwards. 2019. Global wildlife trade across the tree of life. Science 366, 6461 (2019), 71–76
2019
-
[65]
Christopher Schröder, Lydia Müller, Andreas Niekler, and Martin Potthast. 2023. Small-Text: Active Learning for Text Classification in Python. In Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations. 84–95
2023
-
[66]
CITES Secretariat. 2022. World Wildlife Trade Report 2022 . Technical Report. Secretariat of the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES), Geneva, Switzerland
2022
-
[67]
Ozan Sener and Silvio Savarese. 2017. Active learning for convolutional neural networks: A core-set approach
2017
-
[68]
Burr Settles. 2009. Active learning literature survey . Technical Report. University of Wisconsin-Madison Department of Computer Sciences
2009
-
[69]
Penthai Siriwat and Vincent Nijman. 2018. Illegal pet trade on social media as an emerging impediment to the conservation of Asian otters species. Journal of Asia-Pacific Biodiversity 11, 4 (2018), 469–475
2018
-
[70]
Beautiful Soup. 2023. Beautiful Soup Documentation . Beautiful Soup. https: //www.crummy.com/software/BeautifulSoup/bs4/doc/
2023
-
[71]
Sarah Stoner. 2014. Tigers: exploring the threat from illegal online trade.TRAFFIC Bulletin 26, 1 (2014), 26–30
2014
-
[72]
Oliver C Stringham, Stephanie Moncayo, Katherine GW Hill, Adam Toomes, Lewis Mitchell, Joshua V Ross, and Phillip Cassey. 2021. Text classification to streamline online wildlife trade analyses. Plos one 16, 7 (2021), e0254007
2021
-
[73]
Oliver C Stringham, Stephanie Moncayo, Eilish Thomas, Sarah Heinrich, Adam Toomes, Jacob Maher, Katherine GW Hill, Lewis Mitchell, Joshua V Ross, Chris R Shepherd, et al. 2021. Dataset of seized wildlife and their intended uses. Data in Brief 39 (2021), 107531
2021
-
[74]
Alaa Tharwat and Wolfram Schenck. 2023. A survey on active learning: State-of- the-art, practical challenges and research directions. Mathematics 11, 4 (2023), 820
2023
-
[75]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models
2023
-
[76]
Devis Tuia, Benjamin Kellenberger, Sara Beery, Blair R Costelloe, Silvia Zuffi, Ben- jamin Risse, Alexander Mathis, Mackenzie W Mathis, Frank van Langevelde, Tilo Burghardt, et al. 2022. Perspectives in machine learning for wildlife conservation. Nature communications 13, 1 (2...
2022
-
[77]
Daan P van Uhm, Stephen F Pires, Monique Sosnowski, and Gohar A Petrossian
-
[78]
Paroma Varma and Christopher Ré. 2018. Snuba: Automating weak supervision to label training data. Proceedings of the VLDB Endowment (pVLDB) 12 (2018), 223
2018
-
[79]
Ike Vayansky and Sathish AP Kumar. 2020. A review of topic modeling methods. Information Systems 94 (2020), 101582
2020
-
[80]
Sofia Venturini and David L Roberts. 2020. Disguising elephant ivory as other materials in the Online Trade. Tropical Conservation Science 13 (2020), 1940082920974604
2020
-
[81]
VIDA-NYU. [n. d.]. ACHE Crawler. https://github.com/VIDA-NYU/ache. Ac- cessed: 2025-04-09
2025
-
[82]
Ohad Volk and Gonen Singer. 2024. An adaptive cost-sensitive learning approach in neural networks to minimize local training–test class distributions mismatch. Intelligent Systems with Applications 21 (2024), 200316
2024
-
[83]
Meng Wang and Xian-Sheng Hua. 2011. Active learning in multimedia annotation and retrieval: A survey. ACM Transactions on Intelligent Systems and Technology (TIST) 2, 2 (2011), 1–21
2011
-
[84]
Tingting Wang, Shixun Huang, Zhifeng Bao, J Shane Culpepper, Volkan Dedeoglu, and Reza Arablouei. 2024. Optimizing Data Acquisition to Enhance Machine Learning Performance. Proceedings of the VLDB Endowment (pVLDB) 17, 6 (2024), 1310–1323
2024
-
[85]
Steven Euijong Whang, Yuji Roh, Hwanjun Song, and Jae-Gil Lee. 2023. Data collection and quality challenges in deep learning: A data-centric ai perspective. The VLDB Journal 32, 4 (2023), 791–813
2023
-
[86]
White, Q
J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer- Smith, and D. C. Schmidt. 2023. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. arXiv:arXiv:2302.11382 [cs.CL]
2023 arXiv
-
[87]
Wildlife Conservation Society. n.d.. About WCS. https://www.wcs.org/about. Accessed: 2024-10-08
2024
-
[88]
Qing Xu, Mingxiang Cai, and Tim K Mackey. 2020. The illegal wildlife dig- ital market: an analysis of Chinese wildlife marketing and sale on Facebook. Environmental conservation 47, 3 (2020), 206–212
2020
-
[89]
Qing Xu, Jiawei Li, Mingxiang Cai, and Tim K Mackey. 2019. Use of machine learning to detect wildlife product promotion and sales on Twitter. Frontiers in big Data 2 (2019), 28
2019
-
[90]
Guobin Yang, Chenhong Sui, Fuhao Jiang, Yunhao Pan, Ankang Zang, and Jian Hu
-
[91]
Li Yang, Qifan Wang, Zac Yu, Anand Kulkarni, Sumit Sanghai, Bin Shu, Jon Elsas, and Bhargav Kanagal. 2022. MAVE: A Product Dataset for Multi-source Attribute Value Extraction. In Proceedings of the ACM International Conference on Web Search and Data Mining (WSDM) . 1256–1265
2022
-
[92]
Yi Yang, Zhigang Ma, Feiping Nie, Xiaojun Chang, and Alexander G Haupt- mann. 2015. Multi-class active learning by uncertainty sampling with diversity maximization. International Journal of Computer Vision 113 (2015), 113–127
2015
-
[93]
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2023. Flask: Fine- grained language model evaluation based on alignment skill sets. arXiv preprint arXiv:2307.10928. 14 ,
2023 arXiv
-
[2002]
Journal of artificial intelligence research 16 (2002), 321–357
SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16 (2002), 321–357
2002
-
[2019]
In Quantitative Studies in Green and Conservation Criminology
A comparison of seizures of illegal wildlife between the US and the EU: Implications for prevention. In Quantitative Studies in Green and Conservation Criminology. Routledge, Online, 127–145
-
[2020]
Proceedings of the VLDB Endowment (pVLDB) 14, 1 (2020), 28 – 36
Inspector gadget: A data programming-based labeling system for industrial images. Proceedings of the VLDB Endowment (pVLDB) 14, 1 (2020), 28 – 36
2020
-
[2021]
In Proceedings of Inter- national Conference on Advances in Computer Engineering and Communication Systems: ICACECS
A survey on identification of illegal wildlife trade. In Proceedings of Inter- national Conference on Advances in Computer Engineering and Communication Systems: ICACECS. 127–135
-
[2022]
In 2022 Inter- national Conference on Automation, Robotics and Computer Engineering (ICARCE)
Lightweight Conv-Swin Transformer for Wildlife Detection. In 2022 Inter- national Conference on Automation, Robotics and Computer Engineering (ICARCE) . IEEE, 1–5
2022
-
[2023]
Biological Conservation 279 (2023), 109905
Detecting wildlife trafficking in images from online platforms: A test case using deep learning with pangolin images. Biological Conservation 279 (2023), 109905
2023
-
[2024]
In Proceedings of the International Conference on Learning Representa- tions (ICLR)
DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines. In Proceedings of the International Conference on Learning Representa- tions (ICLR). –. https://openreview.net/forum?id=sY5N0zY5Od
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.