REVIEW 2 major objections 1 minor 299 references
Task Decomposition for Efficient Annotation
T0 review · 2 major / 1 minor · reviewed 2026-06-25 · grok-4.3
Pith's one-line read Decomposing structured annotation tasks by isolating center identification reduces aggregate inferential load.
desk verdict The decomposition guidelines and allocation procedure are the practical bits worth a look, but the formal model of inferential load has no shown link to actual annotator effort. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Formal model of inferential load defined by degrees of freedom in the space of valid annotations, with centers from centering theory serving as the salient anchor entities that sub-tasks must realize.
What would settle it
A controlled experiment on the same corpus that measures total annotation time or downstream error rate for end-to-end versus center-first decomposed workflows and finds no reduction in load for the decomposed version.
Extended reading notes
Core claim
By modeling inferential load through degrees of freedom in the space of valid annotations, the paper establishes that annotation decompositions which isolate and advance the identification of centers—salient anchor entities realized by sub-tasks—constrain output space complexity and reduce the aggregate inferential load. This holds across heterogeneous annotators that include both models and humans with varying expertise, and it is supported by guidelines and allocation procedures illustrated with prior cost-efficiency examples.
Load-bearing premise
The formal count of degrees of freedom in valid annotations accurately captures the practical difficulty experienced by human or model annotators.
Editorial extensions
If this is right
- Decompositions that isolate center identification constrain output space complexity.
- Allocating sub-tasks across heterogeneous annotators maximizes quality under a fixed budget.
- Guidelines for decomposition produce measurable cost-efficiency gains as shown in prior examples.
- Modern annotation projects can redesign workflows to match distinct challenges to annotator strengths.
Reading between the lines
- The same decomposition logic could be tested on other structured prediction tasks such as semantic parsing or information extraction.
- Dynamic assignment of sub-tasks might further improve efficiency if annotator performance on each sub-task can be estimated in advance.
- Empirical measurement of actual annotation time before and after decomposition would provide a direct test of the degrees-of-freedom model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that decomposing structured annotation tasks into sub-tasks reduces aggregate inferential load on heterogeneous annotators (humans and models). Drawing from centering theory, it introduces a formal model of inferential load defined via degrees of freedom in the space of valid annotations; identifying 'centers' (salient anchor entities) is argued to constrain output-space complexity, with decompositions that isolate center identification thereby lowering total load. The manuscript supplies guidelines for such decompositions, illustrates them with cost-efficiency examples drawn from the authors' prior work, and outlines a procedure for allocating sub-tasks across annotators to maximize quality under a fixed budget.
Significance. If the formal model is shown to track real annotator effort and the allocation procedure yields measurable gains, the approach could meaningfully improve efficiency and quality control for large-scale structured annotation in NLP, especially when mixing model and human annotators with differing expertise.
major comments (2)
- [formal model of inferential load] The section introducing the formal model of inferential load: the definition of load as degrees of freedom in the valid-annotation space is presented without any derivation, empirical correlation, or validation against observable annotator metrics (time, error rate, cognitive load). This assumption is load-bearing for the central claim that center-isolating decompositions reduce practical difficulty.
- [guidelines and examples] The section on guidelines and examples: cost-efficiency improvements are supported solely by references to prior work rather than new experiments that apply the proposed decomposition procedure and measure the predicted load reduction.
minor comments (1)
- [abstract] The abstract refers to 'our prior work' without citations; adding specific references would clarify the empirical grounding of the examples.
Simulated Author's Rebuttal
We thank the referee for their constructive comments. We address each major point below, clarifying the theoretical nature of the contribution while agreeing where revisions can strengthen the manuscript.
read point-by-point responses
-
Referee: [formal model of inferential load] The section introducing the formal model of inferential load: the definition of load as degrees of freedom in the valid-annotation space is presented without any derivation, empirical correlation, or validation against observable annotator metrics (time, error rate, cognitive load). This assumption is load-bearing for the central claim that center-isolating decompositions reduce practical difficulty.
Authors: The degrees-of-freedom formulation is introduced as a direct formalization of centering theory's notion of salience constraining discourse entities, rather than as an empirically fitted metric. No derivation from first principles or correlation to time/error rates is provided because the paper's focus is the resulting decomposition guidelines, not metric validation. We will revise the model section to include an explicit step-by-step derivation showing how annotation constraints map to degrees of freedom. revision: yes
-
Referee: [guidelines and examples] The section on guidelines and examples: cost-efficiency improvements are supported solely by references to prior work rather than new experiments that apply the proposed decomposition procedure and measure the predicted load reduction.
Authors: The guidelines are illustrated with cost-efficiency outcomes from our earlier annotation projects to show concrete applicability; the manuscript is a methodological proposal, not an empirical study. New experiments applying the procedure and measuring load would be a natural next step but fall outside the current scope. We will add a short subsection outlining how such validation experiments could be designed. revision: partial
Circularity Check
Formal model equates inferential load to degrees of freedom, making reduction claims tautological by definition
-
self definitional
[Abstract]
"we introduce a formal model of inferential load based on the degrees of freedom in the space of valid annotations. Using this model, we show that identifying these centers (i.e. salient anchor entities realized by annotation sub-tasks) constrains the output space complexity, and decompositions which isolate and advance center identification reduce the aggregate inferential load."
Inferential load is defined as degrees of freedom in the valid annotation space. The result that center-identifying decompositions reduce aggregate load follows directly from the definition (constraining the space reduces degrees of freedom), without additional steps or validation against actual annotator effort.
full rationale
The paper introduces a formal model defining inferential load explicitly as degrees of freedom in the annotation space, then uses that model to 'show' that center identification reduces load by constraining the space. This reduction holds by construction of the definition itself rather than through independent derivation or external evidence. Support for guidelines also draws from the authors' prior work, adding a self-citation element, but the core formal step is self-definitional. No equations or further derivations are visible to alter this assessment.
Assumptions & free parameters
assumptions (1)
- domain assumption Centering theory supplies a useful notion of salient anchor entities for constraining annotation spaces
invented entities (1)
-
formal model of inferential load based on degrees of freedom
Cite this review
Pith. "Pith review of Task Decomposition for Efficient Annotation." pith.science (2026). https://pith.science/paper/2GD2H5VV
@misc{pith2026260624734,
author = {Pith},
title = {Pith review of: Task Decomposition for Efficient Annotation},
year = {2026},
howpublished = {\url{https://pith.science/paper/2GD2H5VV}},
note = {Machine review of arXiv:2606.24734}
}
read the original abstract
High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, although cheaper to generate, requires expensive validation and potentially significant supervision to ensure that the annotation quality is strong enough to be useful downstream. In traditional annotation workflows, annotation of each complete example is performed end-to-end by a single annotator. However, structured annotation is complex, and each aspect of the task represents a unique challenge with an associated inferential load for a given annotator. Modern annotation projects can incorporate heterogeneous groups of annotators, including both models and human annotators with varying domain and linguistic expertise. It remains unclear, however, how to redesign annotation tasks in this setting, where efforts are discriminately allocated across heterogeneous annotators with respect to distinct annotation challenges. We propose to decompose annotation tasks into sub-tasks in order to reduce the aggregate inferential load of annotation projects. Inspired by the notion of centers from centering theory, we introduce a formal model of inferential load based on the degrees of freedom in the space of valid annotations. Using this model, we show that identifying these centers (i.e. salient anchor entities realized by annotation sub-tasks) constrains the output space complexity, and decompositions which isolate and advance center identification reduce the aggregate inferential load. We provide guidelines for decomposing complex structured annotation tasks, supported by examples demonstrating improved cost-efficiency from our prior work. Finally, we present a procedure for allocating sub-tasks across annotators to maximize quality under a fixed budget.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Proceedings of the human language technology conference , pages=
The GENIA corpus: An annotated research abstract corpus in molecular biology domain , author=. Proceedings of the human language technology conference , pages=. 2002 , organization=
2002
-
[2]
BMC bioinformatics , volume=
Concept annotation in the CRAFT corpus , author=. BMC bioinformatics , volume=. 2012 , publisher=
2012
-
[3]
Proceedings of the Third Workshop on Building and Evaluating Resources for Biomedical Text Mining , pages=
Developing Specifications for Light Annotation Tasks in the Biomedical Domain , author=. Proceedings of the Third Workshop on Building and Evaluating Resources for Biomedical Text Mining , pages=
-
[4]
Does Synthetic Data Generation of LLMs Help Clinical Text Mining?
Does Synthetic Data Generation of LLMs Help Clinical Text Mining? , author=. arXiv preprint arXiv:2303.04360 , year=
-
[5]
Annotating Named Entities in T witter Data with Crowdsourcing
Finin, Tim and Murnane, William and Karandikar, Anand and Keller, Nicholas and Martineau, Justin and Dredze, Mark. Annotating Named Entities in T witter Data with Crowdsourcing. Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with A mazon ' s Mechanical Turk. 2010
2010
-
[6]
Journal of Combinatorial Theory , volume=
On Stirling numbers of the second kind , author=. Journal of Combinatorial Theory , volume=. 1969 , publisher=
1969
-
[7]
Re- TASK : Revisiting LLM Tasks from Capability, Skill, and Knowledge Perspectives
Wang, Zhihu and Zhao, Shiwan and Wang, Yu and Huang, Heyuan and Xie, Sitao and Zhang, Yubo and Shi, Jiaxin and Wang, Zhixing and Li, Hongyan and Yan, Junchi. Re- TASK : Revisiting LLM Tasks from Capability, Skill, and Knowledge Perspectives. Findings of the Association for Computational Linguistics: ACL 2025. 2025. doi:10.18653/v1/2025.findings-acl.254
-
[8]
arXiv preprint arXiv:2510.04311 , year=
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems , author=. arXiv preprint arXiv:2510.04311 , year=
Show all 299 references
-
[9]
arXiv preprint arXiv:2312.11511 , year=
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity , author=. arXiv preprint arXiv:2312.11511 , year=
-
[10]
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLM s in Detecting Offensive Language with Annotation Disagreement
Lu, Junyu and Ma, Kai and Wang, Kaichun and Xiao, Kelaiti and Lee, Roy Ka-Wei and Xu, Bo and Yang, Liang and Lin, Hongfei. Is LLM an Overconfident Judge? Unveiling the Capabilities of LLM s in Detecting Offensive Language with Annotation Disagreement. Findings of the Associati...
2025 doi
-
[11]
Pushing the Limits of Low-Resource NER Using LLM Artificial Data Generation
Santoso, Joan and Sutanto, Patrick and Cahyadi, Billy and Setiawan, Esther. Pushing the Limits of Low-Resource NER Using LLM Artificial Data Generation. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.findings-acl.575
2024 doi
-
[12]
BMC bioinformatics , volume=
Various criteria in the evaluation of biomedical named entity recognition , author=. BMC bioinformatics , volume=. 2006 , publisher=
2006
-
[13]
Hire a Linguist!: Learning Endangered Languages in LLM s with In-Context Linguistic Descriptions
Zhang, Kexun and Choi, Yee and Song, Zhenqiao and He, Taiqi and Wang, William Yang and Li, Lei. Hire a Linguist!: Learning Endangered Languages in LLM s with In-Context Linguistic Descriptions. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.1...
2024 doi
-
[14]
Journal of the American Medical Informatics Association , volume=
Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutions , author=. Journal of the American Medical Informatics Association , volume=. 2011 , publisher=
2011
-
[15]
Computer Standards & Interfaces , volume=
Named entity recognition: Fallacies, challenges and opportunities , author=. Computer Standards & Interfaces , volume=. 2013 , publisher=
2013
-
[16]
N u NER : Entity Recognition Encoder Pre-training via LLM -Annotated Data
Bogdanov, Sergei and Constantin, Alexandre and Bernard, Timoth \'e e and Crabb \'e , Benoit and Bernard, Etienne P. N u NER : Entity Recognition Encoder Pre-training via LLM -Annotated Data. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing...
2024 doi
-
[17]
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
Golde, Jonas and Haller, Patrick and Ploner, Max and Barth, Fabio and Jedema, Nicolaas and Akbik, Alan. Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data. Proceedings of the 2025 Conference of the Nation...
2025 doi
-
[18]
Evaluating Sequence Labeling on the basis of Information Theory
Amigo, Enrique and \'A lvarez-Mellado, Elena and Gonzalo, Julio and Carrillo-de-Albornoz, Jorge. Evaluating Sequence Labeling on the basis of Information Theory. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 20...
2025 doi
-
[19]
Database , volume=
BioCreative V CDR task corpus: a resource for chemical disease relation extraction , author=. Database , volume=. 2016 , publisher=
2016
-
[20]
Bioinformatics , volume=
GENIA corpus—a semantically annotated corpus for bio-textmining , author=. Bioinformatics , volume=. 2003 , publisher=
2003
-
[21]
FSUIE : A Novel Fuzzy Span Mechanism for Universal Information Extraction
Peng, Tianshuo and Li, Zuchao and Zhang, Lefei and Du, Bo and Zhao, Hai. FSUIE : A Novel Fuzzy Span Mechanism for Universal Information Extraction. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.186...
2023 doi
-
[22]
and Joshi, Aravind K
Grosz, Barbara J. and Joshi, Aravind K. and Weinstein, Scott. C entering: A Framework for Modeling the Local Coherence of Discourse. Computational Linguistics. 1995
1995
-
[23]
arXiv preprint arXiv:2404.01334 , year=
Augmenting NER datasets with LLMs: towards automated and refined annotation , author=. arXiv preprint arXiv:2404.01334 , year=
-
[24]
Journal of the American Medical Informatics Association , volume=
Utilizing active learning strategies in machine-assisted annotation for clinical named entity recognition: a comprehensive analysis considering annotation costs and target effectiveness , author=. Journal of the American Medical Informatics Association , volume=. 2024 , publisher=
2024
-
[25]
Thinking about GPT -3 In-Context Learning for Biomedical IE ? Think Again
Jimenez Gutierrez, Bernal and McNeal, Nikolas and Washington, Clayton and Chen, You and Li, Lang and Sun, Huan and Su, Yu. Thinking about GPT -3 In-Context Learning for Biomedical IE ? Think Again. Findings of the Association for Computational Linguistics: EMNLP 2022. 2022. do...
2022 doi
-
[26]
Proceedings of the 3rd Machine Learning for Health Symposium , pages =
LLMs Accelerate Annotation for Medical Information Extraction , author =. Proceedings of the 3rd Machine Learning for Health Symposium , pages =. 2023 , editor =
2023
-
[27]
Soviet physics doklady , volume=
Binary codes capable of correcting deletions, insertions, and reversals , author=. Soviet physics doklady , volume=. 1966 , organization=
1966
-
[28]
GPT - NER : Named Entity Recognition via Large Language Models
Wang, Shuhe and Sun, Xiaofei and Li, Xiaoya and Ouyang, Rongbin and Wu, Fei and Zhang, Tianwei and Li, Jiwei and Wang, Guoyin and Guo, Chen. GPT - NER : Named Entity Recognition via Large Language Models. Findings of the Association for Computational Linguistics: NAACL 2025. 2...
2025 doi
-
[29]
In-Context Learning for Text Classification with Many Labels
Milios, Aristides and Reddy, Siva and Bahdanau, Dzmitry. In-Context Learning for Text Classification with Many Labels. Proceedings of the 1st GenBench Workshop on (Benchmarking) Generalisation in NLP. 2023. doi:10.18653/v1/2023.genbench-1.14
2023 doi
-
[30]
In-Context Learning on a Budget: A Case Study in Token Classification
Berger, Uri and Baumel, Tal and Stanovsky, Gabriel. In-Context Learning on a Budget: A Case Study in Token Classification. The Sixth Workshop on Insights from Negative Results in NLP. 2025. doi:10.18653/v1/2025.insights-1.2
2025 doi
-
[31]
Revisiting Relation Extraction in the era of Large Language Models
Wadhwa, Somin and Amir, Silvio and Wallace, Byron. Revisiting Relation Extraction in the era of Large Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.868
2023 doi
-
[32]
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
Peng, Siyao and Sun, Zihang and Loftus, Sebastian and Plank, Barbara. Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations. Proceedings of the Third Workshop on Understanding Implicit and Underspecified Language. 2024
2024
-
[33]
NER etrieve: Dataset for Next Generation Named Entity Recognition and Retrieval
Katz, Uri and Vetzler, Matan and Cohen, Amir and Goldberg, Yoav. NER etrieve: Dataset for Next Generation Named Entity Recognition and Retrieval. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.218
2023 doi
-
[34]
arXiv preprint arXiv:2404.07376 , year=
LLMs in Biomedicine: A study on clinical Named Entity Recognition , author=. arXiv preprint arXiv:2404.07376 , year=
-
[35]
Journal of cheminformatics , volume=
The CHEMDNER corpus of chemicals and drugs and its annotation principles , author=. Journal of cheminformatics , volume=. 2015 , publisher=
2015
-
[36]
Comparative evaluation of boundary-relaxed annotation for Entity Linking performance
Herman Bernardim Andrade, Gabriel and Yada, Shuntaro and Aramaki, Eiji. Comparative evaluation of boundary-relaxed annotation for Entity Linking performance. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. ...
2023 doi
-
[37]
When and how to paraphrase for named entity recognition?
Sharma, Saket and Joshi, Aviral and Zhao, Yiyun and Mukhija, Namrata and Bhathena, Hanoz and Singh, Prateek and Santhanam, Sashank. When and how to paraphrase for named entity recognition?. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics...
2023 doi
-
[38]
LLM s are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition
Bai, Fan and Hassanzadeh, Hamid and Saeedi, Ardavan and Dredze, Mark. LLM s are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2...
2025 doi
-
[39]
Quality and quantity , volume=
Measuring the Reliability of Qualitative Text Analysis Data , author=. Quality and quantity , volume=. 2004 , publisher=
2004
-
[40]
Workshop on Information Retrieval Techniques for Speech Applications , pages=
Segmenting Conversations by Topic, Initiative, and Style , author=. Workshop on Information Retrieval Techniques for Speech Applications , pages=. 2001 , organization=
2001
-
[41]
2017 , publisher=
Discovery of Grounded Theory: Strategies for Qualitative Research , author=. 2017 , publisher=
2017
-
[42]
Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence,
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models , author =. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence,. 2025 , month =. doi:10.24963/ijcai.2025/873 , url =
2025 doi
-
[43]
Relation Extraction in Underexplored Biomedical Domains: A Diversity-optimized Sampling and Synthetic Data Generation Approach
Delmas, Maxime and Wysocka, Magdalena and Freitas, Andr \'e. Relation Extraction in Underexplored Biomedical Domains: A Diversity-optimized Sampling and Synthetic Data Generation Approach. Computational Linguistics. 2024. doi:10.1162/coli_a_00520
2024 doi
-
[44]
arXiv preprint arXiv:1503.02531 , year=
Distilling the Knowledge in a Neural Network , author=. arXiv preprint arXiv:1503.02531 , year=
-
[45]
Z ero-shot L abel-Aware E vent T rigger and A rgument C lassification
Zhang, Hongming and Wang, Haoyu and Roth, Dan. Z ero-shot L abel-Aware E vent T rigger and A rgument C lassification. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2021. doi:10.18653/v1/2021.findings-acl.114
2021 doi
-
[46]
Nature , volume=
Artificial intelligence and illusions of understanding in scientific research , author=. Nature , volume=. 2024 , publisher=
2024
-
[47]
S ci ER : An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents
Zhang, Qi and Chen, Zhijia and Pan, Huitong and Caragea, Cornelia and Latecki, Longin Jan and Dragut, Eduard. S ci ER : An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents. Proceedings of the 2024 Conference on Empirical Methods i...
2024 doi
-
[48]
Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!
Ma, Yubo and Cao, Yixin and Hong, Yong and Sun, Aixin. Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.710
2023 doi
-
[49]
Exploiting Asymmetry for Synthetic Training Data Generation: S ynth IE and the Case of Information Extraction
Josifoski, Martin and Sakota, Marija and Peyrard, Maxime and West, Robert. Exploiting Asymmetry for Synthetic Training Data Generation: S ynth IE and the Case of Information Extraction. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 202...
2023 doi
-
[50]
Language Models are Few-Shot Learners , url =
Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom a...
-
[51]
BMC Medical Informatics and Decision Making , volume=
MedTAG: a portable and customizable annotation tool for biomedical documents , author=. BMC Medical Informatics and Decision Making , volume=. 2021 , publisher=
2021
-
[52]
Oscar Sainz and Iker Garc. Go. The Twelfth International Conference on Learning Representations , year=
-
[53]
International conference on learning representations , volume=
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition , author=. International conference on learning representations , volume=
-
[54]
2010 , howpublished =
Fourth i2b2/VA Shared-Task and Workshop: Challenges in Natural Language Processing for Clinical Data (Relations task) , author =. 2010 , howpublished =
2010
-
[55]
Annotating Mentions Alone Enables Efficient Domain Adaptation for Coreference Resolution
Gandhi, Nupoor and Field, Anjalie and Strubell, Emma. Annotating Mentions Alone Enables Efficient Domain Adaptation for Coreference Resolution. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v...
2023 doi
-
[56]
arXiv preprint arXiv:2510.19410 , year=
ToMMeR--Efficient Entity Mention Detection from Large Language Models , author=. arXiv preprint arXiv:2510.19410 , year=
-
[57]
Empirical Study of Zero-Shot NER with C hat GPT
Xie, Tingyu and Li, Qi and Zhang, Jian and Zhang, Yan and Liu, Zuozhu and Wang, Hongwei. Empirical Study of Zero-Shot NER with C hat GPT. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.493
2023 doi
-
[58]
LLM s are not Zero-Shot Reasoners for Biomedical Information Extraction
Nagar, Aishik and Schlegel, Viktor and Nguyen, Thanh-Tung and Li, Hao and Wu, Yuping and Binici, Kuluhan and Winkler, Stefan. LLM s are not Zero-Shot Reasoners for Biomedical Information Extraction. The Sixth Workshop on Insights from Negative Results in NLP. 2025. doi:10.1865...
2025 doi
-
[59]
arXiv preprint arXiv:2203.03903 , year=
InstructionNER: A Multi-Task Instruction-Based Generative Framework for Few-shot NER , author=. arXiv preprint arXiv:2203.03903 , year=
-
[60]
Large Language Models as Annotators of Named Entities in Climate Change and Biodiversity: A Preliminary Study
Volkanovska, Elena. Large Language Models as Annotators of Named Entities in Climate Change and Biodiversity: A Preliminary Study. Proceedings of the 1st Workshop on Ecology, Environment, and Natural Language Processing (NLP4Ecology2025). 2025
2025
-
[61]
Aho and Jeffrey D
Alfred V. Aho and Jeffrey D. Ullman , title =. 1972
1972
-
[62]
Publications Manual , year = "1983", publisher =
1983
-
[63]
Chandra and Dexter C
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243
1981 doi
-
[64]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of
-
[65]
and Li, Toby Jia-Jun and Jiang, Meng and Metoyer, Ronald A
Szymanski, Annalisa and Ziems, Noah and Eicher-Miller, Heather A. and Li, Toby Jia-Jun and Jiang, Meng and Metoyer, Ronald A. , title =. Proceedings of the 30th International Conference on Intelligent User Interfaces , pages =. 2025 , isbn =. doi:10.1145/3708359.3712091 , abstract =
2025 doi
-
[66]
Dan Gusfield , title =. 1997
1997
-
[67]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[68]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =
-
[69]
The Case for Scalable, Data-Driven Theory: A Paradigm for Scientific Progress in NLP
Michael, Julian. The Case for Scalable, Data-Driven Theory: A Paradigm for Scientific Progress in NLP. Proceedings of the Big Picture Workshop. 2023. doi:10.18653/v1/2023.bigpicture-1.4
2023 doi
-
[70]
1998 , institution=
Proposal for an Interactive Environment for Information Extraction , author=. 1998 , institution=
1998
-
[71]
AAAI , volume=
Reducing labeling effort for structured prediction tasks , author=. AAAI , volume=
-
[72]
arXiv preprint arXiv:2312.17543 , year=
Building Efficient Universal Classifiers with Natural Language Inference , author=. arXiv preprint arXiv:2312.17543 , year=
-
[73]
2012 , publisher=
Natural Language Annotation for Machine Learning: A guide to corpus-building for applications , author=. 2012 , publisher=
2012
-
[74]
LLM s instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Bavaresco, Anna and Bernardi, Raffaella and Bertolazzi, Leonardo and Elliott, Desmond and Fern \'a ndez, Raquel and Gatt, Albert and Ghaleb, Esam and Giulianelli, Mario and Hanna, Michael and Koller, Alexander and Martins, Andre and Mondorf, Philipp and Neplenbroek, Vera and P...
2025 doi
-
[75]
Decomposing Unitization and Typing for Efficient and Consistent Span-Bound Concept Annotation
Gandhi, Nupoor and Bada, Michael and Strubell, Emma. Decomposing Unitization and Typing for Efficient and Consistent Span-Bound Concept Annotation. Findings of the Association for Computational Linguistics: ACL 2026. 2026. doi:10.18653/v1/2023.findings-acl.865
2026 doi
-
[76]
Data-efficient Active Learning for Structured Prediction with Partial Annotation and Self-Training
Zhang, Zhisong and Strubell, Emma and Hovy, Eduard. Data-efficient Active Learning for Structured Prediction with Partial Annotation and Self-Training. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.865
2023 doi
-
[77]
GL i NER 2: Schema-Driven Multi-Task Learning for Structured Information Extraction
Zaratiana, Urchade and Pasternak, Gil and Boyd, Oliver and Hurn-Maloney, George and Lewis, Ash. GL i NER 2: Schema-Driven Multi-Task Learning for Structured Information Extraction. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System D...
2025 doi
-
[78]
Management science , volume=
Mathematical methods of organizing and planning production , author=. Management science , volume=. 1960 , publisher=
1960
-
[79]
Machine learning , volume=
Multitask learning , author=. Machine learning , volume=. 1997 , publisher=
1997
-
[80]
AMIA Annual Symposium Proceedings , volume=
Clinical text annotation--what factors are associated with the cost of time? , author=. AMIA Annual Symposium Proceedings , volume=
-
[81]
Information Extraction from Visually Rich Documents using LLM -based Organization of Documents into Independent Textual Segments
Bhattacharyya, Aniket and Tripathi, Anurag and Das, Ujjal and Karmakar, Archan and Pathak, Amit and Gupta, Maneesh. Information Extraction from Visually Rich Documents using LLM -based Organization of Documents into Independent Textual Segments. Proceedings of the 63rd Annual ...
2025 doi
-
[82]
ICLR , year=
Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding , author=. ICLR , year=
-
[83]
Exploring Multimodal Relation Extraction of Hierarchical Tabular Data with Multi-task Learning
Zhang, Xinyu and Song, Aibo and Qiu, Jingyi and Jin, Jiahui and Zhang, Tianbo and Fang, Xiaolin. Exploring Multimodal Relation Extraction of Hierarchical Tabular Data with Multi-task Learning. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguist...
2025 doi
-
[84]
Towards Event Extraction with Massive Types: LLM -based Collaborative Annotation and Partitioning Extraction
Liu, Wenxuan and Li, Zixuan and Bai, Long and Zuo, Yuxin and Xu, Daozhu and Jin, Xiaolong and Guo, Jiafeng and Cheng, Xueqi. Towards Event Extraction with Massive Types: LLM -based Collaborative Annotation and Partitioning Extraction. Proceedings of the 2025 Conference on Empi...
2025 doi
-
[85]
Advances in Neural Information Processing Systems , volume=
Revisiting Out-of-distribution Robustness in NLP: Benchmarks, analysis, and LLMs evaluations , author=. Advances in Neural Information Processing Systems , volume=
-
[86]
Aarohi Srivastava and Abhinav Rastogi and Abhishek Rao and Abu Awal Md Shoeb and Abubakar Abid and Adam Fisch and Adam R. Brown and Adam Santoro and Aditya Gupta and Adrià Garriga-Alonso and Agnieszka Kluska and Aitor Lewkowycz and Akshat Agarwal and Alethea Power and Alex Ray...
-
[87]
ICLR , year=
Learning the difference that makes a difference with counterfactually-augmented data , author=. ICLR , year=
-
[88]
arXiv preprint arXiv:1608.07836 , year=
What to do about non-standard (or non-canonical) language in NLP , author=. arXiv preprint arXiv:1608.07836 , year=
-
[89]
NeurIPS , year=
Shifts: A Dataset of Real Distributional Shift Across Multiple Large-Scale Tasks , author=. NeurIPS , year=
-
[90]
Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society , pages=
Measuring and mitigating unintended bias in text classification , author=. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society , pages=
2018
-
[91]
Journal of computational and applied mathematics , volume=
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis , author=. Journal of computational and applied mathematics , volume=. 1987 , publisher=
1987
-
[92]
Exploring The Landscape of Distributional Robustness for Question Answering Models
Awadalla, Anas and Wortsman, Mitchell and Ilharco, Gabriel and Min, Sewon and Magnusson, Ian and Hajishirzi, Hannaneh and Schmidt, Ludwig. Exploring The Landscape of Distributional Robustness for Question Answering Models. Findings of the Association for Computational Linguist...
2022 doi
-
[93]
ICLR , year=
Prompting GPT-3 To Be Reliable , author=. ICLR , year=
-
[94]
International Conference on Machine Learning , pages=
Large Language Models Struggle to Learn Long-Tail Knowledge , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[95]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of. 2007 , url=
2007
-
[96]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =. 2005 , url=
2005
-
[97]
and Tukey, John W
Cooley, James W. and Tukey, John W. , journal=. An algorithm for the machine calculation of complex. 1965 , url=
1965
-
[98]
Clarendon Press , year=
Centering theory in discourse , author=. Clarendon Press , year=
-
[99]
2012 , publisher=
Natural language annotation for machine learning , author=. 2012 , publisher=
2012
-
[100]
Proceedings of the human language technology conference of the NAACL, Companion Volume: Short Papers , pages=
Ontonotes: the 90\ author=. Proceedings of the human language technology conference of the NAACL, Companion Volume: Short Papers , pages=
-
[101]
World Politics , volume=
Government Responses to Climate Change , author=. World Politics , volume=. 2025 , publisher=
2025
-
[102]
Frontiers in Climate , volume=
Interactions within climate policyscapes: a network analysis of the electricity generation space in the United Kingdom, 1956--2022 , author=. Frontiers in Climate , volume=. 2024 , publisher=
1956
-
[103]
Policy Design and Practice , pages=
Micro-dimensions in practice: how do Argentina and Brazil shape digital transformation through instrumental choices? , author=. Policy Design and Practice , pages=. 2025 , publisher=
2025
-
[104]
Policy Design and Practice , volume=
The institutional design of data governance in Brazil: entropy, restrictiveness and institutional grammar , author=. Policy Design and Practice , volume=. 2025 , publisher=
2025
-
[105]
1995 , publisher=
Learning From Strangers: The Art and Method of Qualitative Interview Studies , author=. 1995 , publisher=
1995
-
[106]
1990 , publisher=
Basics of Qualitative Research , author=. 1990 , publisher=
1990
-
[107]
Forum qualitative sozialforschung/forum: qualitative social research , volume=
Remodeling Grounded Theory , author=. Forum qualitative sozialforschung/forum: qualitative social research , volume=
-
[108]
Cam Journal , volume=
Codebook Development for Team-Based Qualitative Analysis , author=. Cam Journal , volume=. 1998 , publisher=
1998
-
[109]
1994 , publisher=
Qualitative Data Analysis: An Expanded Sourcebook , author=. 1994 , publisher=
1994
-
[110]
2011 , journal =
Paradigmatic Controversies, Contradictions, and Emerging Confluences, Revisited , author =. 2011 , journal =
2011
-
[111]
Encyclopedia of Social Measurement , publisher =
Non-Probability Sampling , editor =. Encyclopedia of Social Measurement , publisher =. 2005 , isbn =. doi:https://doi.org/10.1016/B0-12-369398-5/00382-0 , url =
2005 doi
-
[112]
SAGE research methods foundations , year=
Snowball sampling , author=. SAGE research methods foundations , year=
-
[113]
2024 , eprint=
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery , author=. 2024 , eprint=
2024
-
[114]
2024 , eprint=
Is Semantic Chunking Worth the Computational Cost? , author=. 2024 , eprint=
2024
-
[115]
Advanced Science , volume=
Data-Driven Materials Science: Status, Challenges, and Perspectives , author=. Advanced Science , volume=. 2019 , publisher=
2019
-
[116]
Applied Physics Reviews , volume=
Data-driven materials research enabled by natural language processing and information extraction , author=. Applied Physics Reviews , volume=. 2020 , publisher=
2020
-
[117]
Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval , pages=
Query expansion using local and global document analysis , author=. Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval , pages=
-
[118]
Journal of the ACM (JACM) , volume=
Local feedback in full-text retrieval systems , author=. Journal of the ACM (JACM) , volume=. 1977 , publisher=
1977
-
[119]
arXiv preprint arXiv:2401.10825 , year=
Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study , author=. arXiv preprint arXiv:2401.10825 , year=
-
[120]
ACM Computing Surveys , volume=
Named Entity Recognition and Classification in Historical Documents: A Survey , author=. ACM Computing Surveys , volume=. 2023 , publisher=
2023
-
[121]
Dialogue & Discourse , volume=
Opinion Piece: Can we Fix the Scope for Coreference? Problems and Solutions for Benchmarks beyond OntoNotes , author=. Dialogue & Discourse , volume=
-
[122]
IEEE transactions on knowledge and data engineering , volume=
A Survey on Deep Learning for Named Entity Recognition , author=. IEEE transactions on knowledge and data engineering , volume=. 2020 , publisher=
2020
-
[123]
arXiv preprint arXiv:1811.05468 , year=
Few-shot Learning for Named Entity Recognition in Medical Text , author=. arXiv preprint arXiv:1811.05468 , year=
-
[124]
arXiv preprint arXiv:2305.15444 , year=
PromptNER: Prompting For Named Entity Recognition , author=. arXiv preprint arXiv:2305.15444 , year=
-
[125]
arXiv preprint arXiv:2012.14978 , year=
Few-Shot Named Entity Recognition: A Comprehensive Study , author=. arXiv preprint arXiv:2012.14978 , year=
2012
-
[126]
Cross-document coreference: An approach to capturing coreference without context
Wright-Bettner, Kristin and Palmer, Martha and Savova, Guergana and de Groen, Piet and Miller, Timothy. Cross-document coreference: An approach to capturing coreference without context. Proceedings of the Tenth International Workshop on Health Text Mining and Information Analy...
2019 doi
-
[127]
Cross-document Coreference Resolution over Predicted Mentions
Cattan, Arie and Eirew, Alon and Stanovsky, Gabriel and Joshi, Mandar and Dagan, Ido. Cross-document Coreference Resolution over Predicted Mentions. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. 2021. doi:10.18653/v1/2021.findings-acl.453
2021 doi
-
[128]
Source code for biology and medicine , volume=
Layout-aware text extraction from full-text PDF of scientific articles , author=. Source code for biology and medicine , volume=. 2012 , publisher=
2012
-
[129]
arXiv preprint arXiv:2408.15836 , year=
Knowledge Navigator: LLM-guided Browsing Framework for Exploratory Search in Scientific Literature , author=. arXiv preprint arXiv:2408.15836 , year=
-
[130]
2017 ACM/IEEE joint conference on digital libraries (JCDL) , pages=
A Benchmark and Evaluation for Text Extraction from PDF , author=. 2017 ACM/IEEE joint conference on digital libraries (JCDL) , pages=. 2017 , organization=
2017
-
[131]
arXiv preprint arXiv:2401.14490 , year=
LongHealth: A Question Answering Benchmark with Long Clinical Documents , author=. arXiv preprint arXiv:2401.14490 , year=
-
[132]
Nature Communications , volume=
Structured information extraction from scientific text with large language models , author=. Nature Communications , volume=. 2024 , publisher=
2024
-
[133]
OPENER-Open-NER in domains without annotated resources , author=. Proc. IberSPEECH 2024 , pages=
2024
-
[134]
Harvard Data Science Review , volume=
Why the Data Revolution Needs Qualitative Thinking , author=. Harvard Data Science Review , volume=
-
[135]
International Conference on Information , pages=
A Benchmark of PDF Information Extraction Tools using a Multi-Task and Multi-Domain Evaluation Framework for Academic Documents , author=. International Conference on Information , pages=. 2023 , organization=
2023
-
[136]
arXiv preprint arXiv:2402.12424 , year=
Tables as Images? Exploring the Strengths and Limitations of LLMs on Multimodal Representations of Tabular Data , author=. arXiv preprint arXiv:2402.12424 , year=
-
[137]
arXiv preprint arXiv:2309.08172 , year=
LASER: LLM Agent with State-Space Exploration for Web Navigation , author=. arXiv preprint arXiv:2309.08172 , year=
-
[138]
Proceedings of the 30th ACM International Conference on Multimedia , pages =
Huang, Yupan and Lv, Tengchao and Cui, Lei and Lu, Yutong and Wei, Furu , title =. Proceedings of the 30th ACM International Conference on Multimedia , pages =. 2022 , isbn =. doi:10.1145/3503161.3548112 , abstract =
2022 doi
-
[139]
arXiv preprint arXiv:2409.12191 , year=
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution , author=. arXiv preprint arXiv:2409.12191 , year=
-
[140]
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=
Liu, Haotian and Li, Chunyuan and Li, Yuheng and Li, Bo and Zhang, Yuanhan and Shen, Sheng and Lee, Yong Jae , month=. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=
-
[141]
Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=
Docvqa: A dataset for vqa on document images , author=. Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=
-
[142]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Singh, Amanpreet and Natarajan, Vivek and Shah, Meet and Jiang, Yu and Chen, Xinlei and Batra, Dhruv and Parikh, Devi and Rohrbach, Marcus , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
-
[143]
arXiv preprint arXiv:2410.05613 , year=
Stereotype or Personalization? User Identity Biases Chatbot Recommendations , author=. arXiv preprint arXiv:2410.05613 , year=
-
[144]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
ChatGPT Doesn’t Trust Chargers Fans: Guardrail Sensitivity in Context , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
2024
-
[145]
IEEE Transactions on Visualization and Computer Graphics , year=
CiteRivers: Visual Analytics of Citation Patterns , author=. IEEE Transactions on Visualization and Computer Graphics , year=
-
[146]
Lo, Kyle and Chang, Joseph Chee and Head, Andrew and Bragg, Jonathan and Zhang, Amy X. and Trier, Cassidy and Anastasiades, Chloe and August, Tal and Authur, Russell and Bragg, Danielle and Bransom, Erin and Cachola, Isabel and Candra, Stefan and Chandrasekhar, Yoganand and Ch...
2024 doi
-
[147]
and Hearst, Marti A
Head, Andrew and Lo, Kyle and Kang, Dongyeop and Fok, Raymond and Skjonsberg, Sam and Weld, Daniel S. and Hearst, Marti A. , title =. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , articleno =. 2021 , isbn =. doi:10.1145/3411764.3445648 , abstract =
2021 doi
-
[148]
2024 , eprint=
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval , author=. 2024 , eprint=
2024
-
[149]
arXiv preprint arXiv:1911.03868 , year=
Knowledge guided text retrieval and reading for open domain question answering , author=. arXiv preprint arXiv:1911.03868 , year=
1911
-
[150]
Advances in Neural Information Processing Systems , volume=
Learning to navigate wikipedia by taking random walks , author=. Advances in Neural Information Processing Systems , volume=
-
[151]
arXiv preprint arXiv:1809.09600 , year=
HotpotQA: A dataset for diverse, explainable multi-hop question answering , author=. arXiv preprint arXiv:1809.09600 , year=
-
[152]
arXiv preprint arXiv:2305.14337 , year=
Anchor Prediction: Automatic Refinement of Internet Links , author=. arXiv preprint arXiv:2305.14337 , year=
-
[153]
N ews S ense: Reference-free Verification via Cross-document Comparison
Milbauer, Jeremiah and Ding, Ziqi and Wu, Zhijin and Wu, Tongshuang. N ews S ense: Reference-free Verification via Cross-document Comparison. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 2023. doi:10.18653/v1/20...
2023 doi
-
[154]
Matteo Cargnelutti and Jack Cushman , title =
-
[155]
2024 , month =
David Wilkins , title =. 2024 , month =
2024
-
[156]
Techcrunch.com , date =
Kyle Wiggers , title =. Techcrunch.com , date =
-
[157]
reuters.com , date =
Sara Merken , title =. reuters.com , date =
-
[158]
forbes.com , date =
Ray Ravaglia , title =. forbes.com , date =
-
[159]
2019 , publisher=
Climate action planning: a guide to creating low-carbon, resilient communities , author=. 2019 , publisher=
2019
-
[160]
California Climate Action Plan Database , author =
-
[161]
Departmental papers (ASC) , year=
Computing Krippendorff’s alpha-reliability , author=. Departmental papers (ASC) , year=
-
[162]
Educational and psychological measurement , volume=
A coefficient of agreement for nominal scales , author=. Educational and psychological measurement , volume=. 1960 , publisher=
1960
-
[163]
Scientometrics , volume=
Information extraction from scientific articles: a survey , author=. Scientometrics , volume=. 2018 , publisher=
2018
-
[164]
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, Nils and Gurevych, Iryna. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. 2019
2019
-
[165]
arXiv preprint arXiv:2303.16416 , year=
Zero-shot clinical entity recognition using chatgpt , author=. arXiv preprint arXiv:2303.16416 , year=
-
[166]
arXiv preprint arXiv:2307.00186 , year=
How far is language model from 100\ author=. arXiv preprint arXiv:2307.00186 , year=
-
[167]
How different is different? Systematically identifying distribution shifts and their impacts in NER datasets , author=
-
[168]
arXiv preprint arXiv:2402.10744 , year=
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models , author=. arXiv preprint arXiv:2402.10744 , year=
-
[169]
arXiv preprint arXiv:2305.14450 , year=
Is information extraction solved by chatgpt? an analysis of performance, evaluation criteria, robustness and errors , author=. arXiv preprint arXiv:2305.14450 , year=
-
[171]
2008--2024 , archivePrefix =
GROBID , howpublished =. 2008--2024 , archivePrefix =
2008
-
[172]
Research and Advanced Technology for Digital Libraries: 13th European Conference, ECDL 2009, Corfu, Greece, September 27-October 2, 2009
GROBID: Combining automatic bibliographic data recognition and term extraction for scholarship publications , author=. Research and Advanced Technology for Digital Libraries: 13th European Conference, ECDL 2009, Corfu, Greece, September 27-October 2, 2009. Proceedings 13 , pag...
2009
-
[173]
Journal of the Medical Library Association , volume=
Modeling public health interventions for improved access to the gray literature , author=. Journal of the Medical Library Association , volume=. 2005 , publisher=
2005
-
[174]
, author=
Measuring nominal scale agreement among many raters. , author=. Psychological bulletin , volume=. 1971 , publisher=
1971
-
[175]
Ecological Indicators , volume=
Climate adaptation indicators and metrics: State of local policy practice , author=. Ecological Indicators , volume=. 2022 , publisher=
2022
-
[176]
JMIR Medical Informatics , volume=
Is Boundary Annotation Necessary? Evaluating Boundary-Free Approaches to Improve Clinical Named Entity Annotation Efficiency: Case Study , author=. JMIR Medical Informatics , volume=. 2024 , publisher=
2024
-
[177]
Journal of the American Medical Informatics Association , volume=
2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text , author=. Journal of the American Medical Informatics Association , volume=. 2011 , publisher=
2010
-
[178]
X - AMR Annotation Tool
Ahmed, Shafiuddin Rehan and Cai, Jon and Palmer, Martha and Martin, James H. X - AMR Annotation Tool. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations. 2024
2024
-
[179]
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
Sundar, Kawshik Manikantan and Toshniwal, Shubham and Tapaswi, Makarand and Gandhi, Vineet. Major Entity Identification: A Generalizable Alternative to Coreference Resolution. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10....
2024 doi
-
[180]
COLM , year=
Are Large Language Models Robust Coreference Resolvers? , author=. COLM , year=
-
[181]
Nature climate change , volume=
A systematic global stocktake of evidence on human adaptation to climate change , author=. Nature climate change , volume=. 2021 , publisher=
2021
-
[182]
and Lin, Kevin and Hewitt, John and Paranjape, Ashwin and Bevilacqua, Michele and Petroni, Fabio and Liang, Percy
Liu, Nelson F. and Lin, Kevin and Hewitt, John and Paranjape, Ashwin and Bevilacqua, Michele and Petroni, Fabio and Liang, Percy. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00638
2024 doi
-
[183]
The State of Relation Extraction Data Quality: Is Bigger Always Better?
Cai, Erica and O ' Connor, Brendan. The State of Relation Extraction Data Quality: Is Bigger Always Better?. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.findings-acl.470
2024 doi
-
[184]
arXiv preprint arXiv:2101.05779 , year=
Structured prediction as translation between augmented natural languages , author=. arXiv preprint arXiv:2101.05779 , year=
-
[185]
arXiv preprint arXiv:2307.09288 , year=
Llama 2: Open foundation and fine-tuned chat models , author=. arXiv preprint arXiv:2307.09288 , year=
-
[186]
Advances in Neural Information Processing Systems , volume=
Unlimiformer: Long-range transformers with unlimited length input , author=. Advances in Neural Information Processing Systems , volume=
-
[187]
arXiv preprint arXiv:2308.16137 , year=
Lm-infinite: Simple on-the-fly length generalization for large language models , author=. arXiv preprint arXiv:2308.16137 , year=
-
[188]
arXiv preprint arXiv:2308.13418 , year=
Nougat: Neural optical understanding for academic documents , author=. arXiv preprint arXiv:2308.13418 , year=
-
[189]
2017 ACM/IEEE joint conference on digital libraries (JCDL) , pages=
A benchmark and evaluation for text extraction from PDF , author=. 2017 ACM/IEEE joint conference on digital libraries (JCDL) , pages=. 2017 , organization=
2017
-
[190]
arXiv preprint arXiv:2403.12924 , year=
Supporting Energy Policy Research with Large Language Models , author=. arXiv preprint arXiv:2403.12924 , year=
-
[191]
Proceedings of the eighth conference on computational natural language learning (CoNLL-2004) at HLT-NAACL 2004 , pages=
A linear programming formulation for global inference in natural language tasks , author=. Proceedings of the eighth conference on computational natural language learning (CoNLL-2004) at HLT-NAACL 2004 , pages=
2004
-
[192]
Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2010, Barcelona, Spain, September 20-24, 2010, Proceedings, Part III 21 , pages=
Modeling relations and their mentions without labeled text , author=. Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2010, Barcelona, Spain, September 20-24, 2010, Proceedings, Part III 21 , pages=. 2010 , organization=
2010
-
[193]
arXiv preprint arXiv:2404.02822 , year=
Identifying Climate Targets in National Laws and Policies using Machine Learning , author=. arXiv preprint arXiv:2404.02822 , year=
-
[194]
International Conference on Information , pages=
A benchmark of pdf information extraction tools using a multi-task and multi-domain evaluation framework for academic documents , author=. International Conference on Information , pages=. 2023 , organization=
2023
-
[195]
Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations , title =
Jan-Christoph Klie and Michael Bugert and Beto Boullosa and Richard Eckart de Castilho and Iryna Gurevych , eventdate =. Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations , title =. 2018 , month =
2018
-
[196]
Journal of the Young Librarians Association , volume=
Grey literature: A valuable untapped stockpile of information , author=. Journal of the Young Librarians Association , volume=
-
[197]
The handbook of research synthesis and meta-analysis , volume=
Grey literature , author=. The handbook of research synthesis and meta-analysis , volume=
-
[198]
Australian Academic & Research Libraries , volume=
Collecting the evidence: Improving access to grey literature and data for public policy and practice , author=. Australian Academic & Research Libraries , volume=. 2015 , publisher=
2015
-
[199]
Journal of Planning Education and Research , pages=
What is in a plan? Using natural language processing to read 461 California city general plans , author=. Journal of Planning Education and Research , pages=. 2021 , publisher=
2021
-
[200]
npj Urban Sustainability , volume=
A computational approach to analyzing climate strategies of cities pledging net zero , author=. npj Urban Sustainability , volume=. 2022 , publisher=
2022
-
[201]
Regional Environmental Change , volume=
Machine learning for research on climate change adaptation policy integration: an exploratory UK case study , author=. Regional Environmental Change , volume=. 2020 , publisher=
2020
-
[202]
ACE , year=
NYU's English ACE 2005 system description , author=. ACE , year=
2005
-
[203]
Guideline Learning for In-Context Information Extraction
Pang, Chaoxu and Cao, Yixuan and Ding, Qiang and Luo, Ping. Guideline Learning for In-Context Information Extraction. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.950
2023 doi
-
[204]
Don't Take Their Word for It (February 03, 2025) , year=
Large Language Models for Legal Interpretation? Don't Take Their Word for It , author=. Don't Take Their Word for It (February 03, 2025) , year=
2025
-
[205]
Computational Linguistics , volume=
A corpus-based evaluation of centering and pronoun resolution , author=. Computational Linguistics , volume=. 2001 , publisher=
2001
-
[206]
Proceedings of the 5th International Conference on Language Resources and Evaluation (LREC) , year=
Multi-level annotation of linguistic data with MMAX2 , author=. Proceedings of the 5th International Conference on Language Resources and Evaluation (LREC) , year=
-
[207]
Centering in Discourse , pages=
Centering: A Parametric Theory and its Instantiations , author=. Centering in Discourse , pages=. 2004 , publisher=
2004
-
[208]
Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL) , year=
A Reliable Human Evaluation Metric for Grammatical Error Correction , author=. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL) , year=
-
[209]
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) , year=
TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task , author=. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) , year=
-
[210]
Journal of the American Statistical Association , volume=
Hierarchical grouping to optimize an objective function , author=. Journal of the American Statistical Association , volume=. 1963 , publisher=
1963
-
[211]
arXiv preprint arXiv:2502.16377 , year=
Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines , author=. arXiv preprint arXiv:2502.16377 , year=
-
[212]
Journal of Public Administration Research and Theory , volume =
Use of Boilerplate Language in Regulatory Documents: Evidence from Environmental Impact Statements , author =. Journal of Public Administration Research and Theory , volume =. 2022 , month =. doi:10.1093/jopart/muab048 , url =
2022 doi
-
[213]
Contribution of Working Groups I, II and III to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change , author =
Climate Change 2023: Synthesis Report. Contribution of Working Groups I, II and III to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change , author =. 2023 , publisher =
2023
-
[214]
Findings of the Association for Computational Linguistics: ACL 2023 , pages=
Is Summary Useful or Not? An Extrinsic Human Evaluation of Text Summaries on Downstream Tasks , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=. 2023 , url=
2023
-
[215]
, year =
Lucy, Li and Dodge, Jesse and Bamman, David and Keith, Katherine A. , year =. Words as. doi:10.48550/ARXIV.2212.09676 , abstract =
- [216]
-
[217]
Centering in Discourse II , pages=
Centering: A parametric theory and its instantiations , author=. Centering in Discourse II , pages=. 2004 , publisher=
2004
-
[218]
BMC bioinformatics , volume=
Corpus annotation for mining biomedical events from literature , author=. BMC bioinformatics , volume=. 2008 , publisher=
2008
-
[219]
arXiv preprint arXiv:2407.12826 , year=
Assessing the Effectiveness of GPT-4o in Climate Change Evidence Synthesis and Systematic Assessments: Preliminary Insights , author=. arXiv preprint arXiv:2407.12826 , year=
-
[220]
Journal of the American Planning Association , volume=
Explaining progress in climate adaptation planning across 156 US municipalities , author=. Journal of the American Planning Association , volume=. 2015 , publisher=
2015
-
[221]
Journal of the American planning association , volume=
Innovation and climate action planning: Perspectives from municipal plans , author=. Journal of the American planning association , volume=. 2010 , publisher=
2010
-
[222]
Proceedings of the VLDB Endowment , volume=
Nlprov: Natural language provenance , author=. Proceedings of the VLDB Endowment , volume=. 2016 , publisher=
2016
-
[223]
Datasheets for datasets , year =
Gebru, Timnit and Morgenstern, Jamie and Vecchione, Briana and Vaughan, Jennifer Wortman and Wallach, Hanna and III, Hal Daum\'. Datasheets for datasets , year =. Commun. ACM , month =. doi:10.1145/3458723 , abstract =
-
[224]
Under submission , year=
Collage: Decomposable Rapid Prototyping for Information Extraction on Scientific PDFs , author=. Under submission , year=
-
[225]
arXiv:2004.05150 , year=
Longformer: The Long-Document Transformer , author=. arXiv:2004.05150 , year=
2004 arXiv
-
[226]
Challenges in End-to-End Policy Extraction from Climate Action Plans
Gandhi, Nupoor and Corringham, Tom and Strubell, Emma. Challenges in End-to-End Policy Extraction from Climate Action Plans. Proceedings of the 1st Workshop on Natural Language Processing Meets Climate Change (ClimateNLP 2024). 2024
2024
-
[227]
Regulation & Governance , volume=
Regulatory policy outputs and impacts: Exploring a complex relationship , author=. Regulation & Governance , volume=. 2012 , publisher=
2012
-
[228]
IEEE Intelligent Systems , volume=
Information extraction , author=. IEEE Intelligent Systems , volume=. 2015 , publisher=
2015
-
[229]
Corringham, T. W. and Spokoyny, D. and Xiao, E. and Cha, C. and Lemarchand, C. and Syal, M. and Olson, E. and Gershunov, A. , title =. ICML Workshop, Tackling Climate Change with Machine Learning , year =
-
[230]
arXiv preprint arXiv:2401.09646 , year=
Climategpt: Towards ai synthesizing interdisciplinary research on climate change , author=. arXiv preprint arXiv:2401.09646 , year=
-
[231]
arXiv preprint arXiv:2307.01972 , year=
Open-domain hierarchical event schema induction by incremental prompting and verification , author=. arXiv preprint arXiv:2307.01972 , year=
-
[232]
arXiv preprint arXiv:2302.10205 , year=
Zero-shot information extraction via chatting with chatgpt , author=. arXiv preprint arXiv:2302.10205 , year=
-
[233]
Cross-Document Event Coreference Resolution: Instruct Humans or Instruct GPT ?
Zhao, Jin and Xue, Nianwen and Min, Bonan. Cross-Document Event Coreference Resolution: Instruct Humans or Instruct GPT ?. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. doi:10.18653/v1/2023.conll-1.38
2023 doi
-
[234]
Neurological Research and practice , volume=
How to use and assess qualitative research methods , author=. Neurological Research and practice , volume=. 2020 , publisher=
2020
-
[235]
2025 , eprint=
Beyond Text: Characterizing Domain Expert Needs in Document Research , author=. 2025 , eprint=
2025
-
[236]
Named Entity Recognition Under Domain Shift via Metric Learning for Life Sciences
Liu, Hongyi and Wang, Qingyun and Karisani, Payam and Ji, Heng. Named Entity Recognition Under Domain Shift via Metric Learning for Life Sciences. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language ...
2024 doi
-
[237]
Interpretable Multi-dataset Evaluation for Named Entity Recognition
Fu, Jinlan and Liu, Pengfei and Neubig, Graham. Interpretable Multi-dataset Evaluation for Named Entity Recognition. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. doi:10.18653/v1/2020.emnlp-main.489
2020 doi
-
[238]
COPEN : Probing Conceptual Knowledge in Pre-trained Language Models
Peng, Hao and Wang, Xiaozhi and Hu, Shengding and Jin, Hailong and Hou, Lei and Li, Juanzi and Liu, Zhiyuan and Liu, Qun. COPEN : Probing Conceptual Knowledge in Pre-trained Language Models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing...
2022 doi
-
[239]
Proceedings of the 26th Conference on Artificial Intelligence (AAAI) , year=
Fine-grained entity recognition , author=. Proceedings of the 26th Conference on Artificial Intelligence (AAAI) , year=
-
[240]
and Dreier, Sarah , journal =
Tanweer, Anissa and Gade, Emily Kalah and Krafft, P.M. and Dreier, Sarah , journal =. Why the. 2021 , month =
2021
-
[241]
Large Language Models Enable Few-Shot Clustering
Viswanathan, Vijay and Gashteovski, Kiril and Gashteovski, Kiril and Lawrence, Carolin and Wu, Tongshuang and Neubig, Graham. Large Language Models Enable Few-Shot Clustering. Transactions of the Association for Computational Linguistics. 2024
2024
-
[242]
Text Classification Using Label Names Only: A Language Model Self-Training Approach
Meng, Yu and Zhang, Yunyi and Huang, Jiaxin and Xiong, Chenyan and Ji, Heng and Zhang, Chao and Han, Jiawei. Text Classification Using Label Names Only: A Language Model Self-Training Approach. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process...
2020 doi
-
[243]
Scientific Data , volume=
Towards understanding policy design through text-as-data approaches: The policy design annotations (POLIANNA) dataset , author=. Scientific Data , volume=. 2023 , publisher=
2023
-
[244]
Tackling climate change with machine learning workshop at ICML , year=
Neuralnere: Neural named entity relationship extraction for end-to-end climate change knowledge graph construction , author=. Tackling climate change with machine learning workshop at ICML , year=
-
[245]
Towards Fine-grained Classification of Climate Change related Social Media Text
Vaid, Roopal and Pant, Kartikey and Shrivastava, Manish. Towards Fine-grained Classification of Climate Change related Social Media Text. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop. 2022. doi:10.18653/v1/2...
2022 doi
-
[246]
arXiv preprint arXiv:2012.00483 , year=
ClimaText: A dataset for climate change topic detection , author=. arXiv preprint arXiv:2012.00483 , year=
2012
-
[247]
An NLP Benchmark Dataset for Assessing Corporate Climate Policy Engagement , url =
Morio, Gaku and Manning, Christopher D , booktitle =. An NLP Benchmark Dataset for Assessing Corporate Climate Policy Engagement , url =
-
[248]
and Spokoyny, D
Laud, T. and Spokoyny, D. and Corringham, T. W. and Berg-Kirkpatrick, T. , title =. EMNLP 2022 NLP4PI Workshop , year =
2022
-
[249]
Policy Studies Journal , volume=
Toward a comparative measure of climate policy output , author=. Policy Studies Journal , volume=. 2015 , publisher=
2015
-
[250]
C o A nnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
Li, Minzhi and Shi, Taiwei and Ziems, Caleb and Kan, Min-Yen and Chen, Nancy and Liu, Zhengyuan and Yang, Diyi. C o A nnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation. Proceedings of the 2023 Conference on Empirical Meth...
2023 doi
-
[251]
and Cai, J
Spokoyny, D. and Cai, J. and Corringham, T. W. and Berg-Kirkpatrick, T. , title =. Proceedings of the 1st Workshop on Natural Language Processing Meets Climate Change (ClimateNLP 2024) , year =
2024
-
[252]
and Laud, T
Spokoyny, D. and Laud, T. and Corringham, T. W. and Berg-Kirkpatrick, T. , title =. arXiv [Cs.CL] , year =
-
[253]
Bali Principles of Climate Justice , year =
-
[254]
Environmental Justice Principles , year =
-
[255]
This is an example of sample bibitem article title , journal =
Surname, FirstName , year =. This is an example of sample bibitem article title , journal =
-
[256]
This is an example of sample bibitem article title , booktitle =
Surname, FirstName , year =. This is an example of sample bibitem article title , booktitle =
-
[257]
Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[258]
Virtual Citation Proximity ( VCP ): Empowering Document Recommender Systems by Learning a Hypothetical In-Text Citation-Proximity Metric for Uncited Documents
Molloy, Paul and Beel, Joeran and Aizawa, Akiko. Virtual Citation Proximity ( VCP ): Empowering Document Recommender Systems by Learning a Hypothetical In-Text Citation-Proximity Metric for Uncited Documents. Proceedings of the 8th International Workshop on Mining Scientific P...
2020
-
[259]
Citations Beyond Self Citations: Identifying Authors, Affiliations, and Nationalities in Scientific Papers
Matsubara, Yoshitomo and Singh, Sameer. Citations Beyond Self Citations: Identifying Authors, Affiliations, and Nationalities in Scientific Papers. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[260]
S mart C ite C on: Implicit Citation Context Extraction from Academic Literature Using Supervised Learning
Guo, Chenrui and Cui, Haoran and Zhang, Li and Wang, Jiamin and Lu, Wei and Wu, Jian. S mart C ite C on: Implicit Citation Context Extraction from Academic Literature Using Supervised Learning. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[261]
Synthetic vs
Grennan, Mark and Beel, Joeran. Synthetic vs. Real Reference Strings for Citation Parsing, and the Importance of Re-training and Out-Of-Sample Data for Meaningful Evaluations: Experiments with GROBID , GIANT and CORA. Proceedings of the 8th International Workshop on Mining Sci...
2020
-
[262]
Term-Recency for TF - IDF , BM 25 and USE Term Weighting
Marwah, Divyanshu and Beel, Joeran. Term-Recency for TF - IDF , BM 25 and USE Term Weighting. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[263]
The Normalized Impact Index for Keywords in Scholarly Papers to Detect Subtle Research Topics
Ikeda, Daisuke and Taniguchi, Yuta and Koga, Kazunori. The Normalized Impact Index for Keywords in Scholarly Papers to Detect Subtle Research Topics. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[264]
Representing and Reconstructing P hy SH : Which Embedding Competent?
Chen, Xiaoli and Zhang, Zhixiong. Representing and Reconstructing P hy SH : Which Embedding Competent?. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[265]
Combining Representations For Effective Citation Classification
de Andrade, Claudio Mois \'e s Valiense and Gon c alves, Marcos Andr \'e. Combining Representations For Effective Citation Classification. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[266]
Scubed at 3 C task A - A simple baseline for citation context purpose classification
Mishra, Shubhanshu and Mishra, Sudhanshu. Scubed at 3 C task A - A simple baseline for citation context purpose classification. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[267]
Scubed at 3 C task B - A simple baseline for citation context influence classification
Mishra, Shubhanshu and Mishra, Sudhanshu. Scubed at 3 C task B - A simple baseline for citation context influence classification. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[268]
A mrita \_ CEN \_ NLP @ WOSP 3 C Citation Context Classification Task
B, Premjith and KP, Soman. A mrita \_ CEN \_ NLP @ WOSP 3 C Citation Context Classification Task. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[269]
Overview of the 2020 WOSP 3 C Citation Context Classification Task
Kunnath, Suchetha Nambanoor and Pride, David and Gyawali, Bikash and Knoth, Petr. Overview of the 2020 WOSP 3 C Citation Context Classification Task. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[270]
Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020
2020
-
[271]
May I Ask Who ' s Calling? Named Entity Recognition on Call Center Transcripts for Privacy Law Compliance
Kaplan, Micaela. May I Ask Who ' s Calling? Named Entity Recognition on Call Center Transcripts for Privacy Law Compliance. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.1
2020 doi
-
[272]
`` Did you really mean what you said? '' : Sarcasm Detection in H indi- E nglish Code-Mixed Data using Bilingual Word Embeddings
Aggarwal, Akshita and Wadhawan, Anshul and Chaudhary, Anshima and Maurya, Kavita. `` Did you really mean what you said? '' : Sarcasm Detection in H indi- E nglish Code-Mixed Data using Bilingual Word Embeddings. Proceedings of the Sixth Workshop on Noisy User-generated Text (W...
2020 doi
-
[273]
Noisy Text Data: Achilles ' Heel of BERT
Kumar, Ankit and Makhija, Piyush and Gupta, Anuj. Noisy Text Data: Achilles ' Heel of BERT. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.3
2020 doi
-
[274]
Determining Question-Answer Plausibility in Crowdsourced Datasets Using Multi-Task Learning
Gardner, Rachel and Varma, Maya and Zhu, Clare and Krishna, Ranjay. Determining Question-Answer Plausibility in Crowdsourced Datasets Using Multi-Task Learning. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.4
2020 doi
-
[275]
Combining BERT with Static Word Embeddings for Categorizing Social Media
Alghanmi, Israa and Espinosa Anke, Luis and Schockaert, Steven. Combining BERT with Static Word Embeddings for Categorizing Social Media. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.5
2020 doi
-
[276]
Enhanced Sentence Alignment Network for Efficient Short Text Matching
Hu, Zhe and Fu, Zuohui and Peng, Cheng and Wang, Weiwei. Enhanced Sentence Alignment Network for Efficient Short Text Matching. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.6
2020 doi
-
[277]
PHINC : A Parallel H inglish Social Media Code-Mixed Corpus for Machine Translation
Srivastava, Vivek and Singh, Mayank. PHINC : A Parallel H inglish Social Media Code-Mixed Corpus for Machine Translation. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.7
2020 doi
-
[278]
Cross-lingual sentiment classification in low-resource B engali language
Sazzed, Salim. Cross-lingual sentiment classification in low-resource B engali language. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.8
2020 doi
-
[279]
The Non-native Speaker Aspect: I ndian E nglish in Social Media
Sarkar, Rupak and Mahinder, Sayantan and KhudaBukhsh, Ashiqur. The Non-native Speaker Aspect: I ndian E nglish in Social Media. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.9
2020 doi
-
[280]
Sentence Boundary Detection on Line Breaks in J apanese
Hayashibe, Yuta and Mitsuzawa, Kensuke. Sentence Boundary Detection on Line Breaks in J apanese. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.10
2020 doi
-
[281]
Non-ingredient Detection in User-generated Recipes using the Sequence Tagging Approach
Yamaguchi, Yasuhiro and Inuzuka, Shintaro and Hiramatsu, Makoto and Harashima, Jun. Non-ingredient Detection in User-generated Recipes using the Sequence Tagging Approach. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.11
2020 doi
-
[282]
Generating Fact Checking Summaries for Web Claims
Mishra, Rahul and Gupta, Dhruv and Leippold, Markus. Generating Fact Checking Summaries for Web Claims. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.12
2020 doi
-
[283]
Intelligent Analyses on Storytelling for Impact Measurement
Kicken, Koen and De Maesschalck, Tessa and Vanrumste, Bart and De Keyser, Tom and Shim, Hee Reen. Intelligent Analyses on Storytelling for Impact Measurement. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.13
2020 doi
-
[284]
An Empirical Analysis of Human-Bot Interaction on R eddit
Ma, Ming-Cheng and Lalor, John P. An Empirical Analysis of Human-Bot Interaction on R eddit. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.14
2020 doi
-
[285]
Detecting Trending Terms in Cybersecurity Forum Discussions
Hughes, Jack and Aycock, Seth and Caines, Andrew and Buttery, Paula and Hutchings, Alice. Detecting Trending Terms in Cybersecurity Forum Discussions. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.15
2020 doi
-
[286]
Service registration chatbot: collecting and comparing dialogues from AMT workers and service ' s users
Molteni, Luca and Singh, Mittul and Leinonen, Juho and Leino, Katri and Kurimo, Mikko and Della Valle, Emanuele. Service registration chatbot: collecting and comparing dialogues from AMT workers and service ' s users. Proceedings of the Sixth Workshop on Noisy User-generated T...
2020 doi
-
[287]
Automated Assessment of Noisy Crowdsourced Free-text Answers for H indi in Low Resource Setting
Agarwal, Dolly and Gupta, Somya and Baghel, Nishant. Automated Assessment of Noisy Crowdsourced Free-text Answers for H indi in Low Resource Setting. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.17
2020 doi
-
[288]
Punctuation Restoration using Transformer Models for High-and Low-Resource Languages
Alam, Tanvirul and Khan, Akib and Alam, Firoj. Punctuation Restoration using Transformer Models for High-and Low-Resource Languages. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.18
2020 doi
-
[289]
Truecasing G erman user-generated conversational text
Grishina, Yulia and Gueudre, Thomas and Winkler, Ralf. Truecasing G erman user-generated conversational text. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.19
2020 doi
-
[290]
Fine-Tuning MT systems for Robustness to Second-Language Speaker Variations
Alam, Md Mahfuz Ibn and Anastasopoulos, Antonios. Fine-Tuning MT systems for Robustness to Second-Language Speaker Variations. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.20
2020 doi
-
[291]
Impact of ASR on A lzheimer ' s Disease Detection: All Errors are Equal, but Deletions are More Equal than Others
Balagopalan, Aparna and Shkaruta, Ksenia and Novikova, Jekaterina. Impact of ASR on A lzheimer ' s Disease Detection: All Errors are Equal, but Deletions are More Equal than Others. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653...
2020 doi
-
[292]
Detecting Entailment in Code-Mixed H indi- E nglish Conversations
Chakravarthy, Sharanya and Umapathy, Anjana and Black, Alan W. Detecting Entailment in Code-Mixed H indi- E nglish Conversations. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.22
2020 doi
-
[293]
Detecting Objectifying Language in Online Professor Reviews
Waller, Angie and Gorman, Kyle. Detecting Objectifying Language in Online Professor Reviews. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.23
2020 doi
-
[294]
Annotation Efficient Language Identification from Weak Labels
Palakodety, Shriphani and KhudaBukhsh, Ashiqur. Annotation Efficient Language Identification from Weak Labels. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.24
2020 doi
-
[295]
Fantastic Features and Where to Find Them: Detecting Cognitive Impairment with a Subsequence Classification Guided Approach
Eyre, Ben and Balagopalan, Aparna and Novikova, Jekaterina. Fantastic Features and Where to Find Them: Detecting Cognitive Impairment with a Subsequence Classification Guided Approach. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18...
2020 doi
-
[296]
Quantifying the Evaluation of Heuristic Methods for Textual Data Augmentation
Kashefi, Omid and Hwa, Rebecca. Quantifying the Evaluation of Heuristic Methods for Textual Data Augmentation. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.26
2020 doi
-
[297]
An Empirical Survey of Unsupervised Text Representation Methods on T witter Data
Wang, Lili and Gao, Chongyang and Wei, Jason and Ma, Weicheng and Liu, Ruibo and Vosoughi, Soroush. An Empirical Survey of Unsupervised Text Representation Methods on T witter Data. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653...
2020 doi
-
[298]
and Dredze, Mark
Sech, Justin and DeLucia, Alexandra and Buczak, Anna L. and Dredze, Mark. Civil Unrest on T witter ( CUT ): A Dataset of Tweets to Support Research on Civil Unrest. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.28
2020 doi
-
[299]
Tweeki: Linking Named Entities on T witter to a Knowledge Graph
Harandizadeh, Bahareh and Singh, Sameer. Tweeki: Linking Named Entities on T witter to a Knowledge Graph. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.29
2020 doi
-
[300]
Representation learning of writing style
Hay, Julien and Doan, Bich-Lien and Popineau, Fabrice and Ait Elhara, Ouassim. Representation learning of writing style. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.30
2020 doi
Reviewed June 25, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.