REVIEW 4 major objections 4 minor 94 references
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper argues that NLU diagnostics benchmarks have no shared naming convention for their linguistic categories, making model error analysis incomparable across benchmarks, and proposes building a global hierarchy of linguistic…
desk verdict Useful taxonomy comparison of five diagnostics benchmarks, but the field-wide 'no naming convention' claim outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying objects are the diagnostics datasets themselves and their two-level taxonomies: macro-categories such as Lexical Semantics, Predicate-Argument Structure, Logic, Quantification, and Knowledge and Common Sense, each subdivided into micro-categories. The comparative method is to tabulate which phenomena each benchmark covers, at which level of the hierarchy, with how many sentence pairs, and with what distribution of entailment classes (entailment, neutral/unknown, contradiction). The taxonomies are also the evidence: placing the same phenomenon at different levels in different datasets is what demonstrates the absence of a shared naming convention.
What would settle it
The claim that no naming convention exists would be settled by searching for an NLU diagnostics benchmark outside the five examined that publishes a fixed hierarchy of macro- and micro-categories with a defined phenomenon list; finding one would directly weaken the claim. A complementary test would be to attempt a faithful mapping of the five taxonomies onto a single hierarchy; if the mapping succeeds without residue, the problem is terminological, whereas if phenomena refuse to align, the taxonomies encode genuinely different analyses.
Extended reading notes
Core claim
The paper's central claim is that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena across NLU diagnostics benchmarks. Rather than a single finding, the claim is established by a comparative analysis showing systematic incommensurability: FraCaS keeps Quantifiers, Plurals, and Adjectives as separate macro-categories; GLUE and ALUE group the same material under Lexical Semantics or Predicate-Argument Structure; CLUE has only coarse-grained categories with no fine-grained structure; FraCaS lacks any Logic category while the other diagnostics include one; and world knowledge is an explicit category in GLUE and ALUE but deliberately implicit in FraCaS and the specialized TE dataset. The authors also report distributional statistics, including that the majority of samples in each dataset fall under the entailment class. From this comparison they conclude that an evaluation standard for diagnostics benchmarks is missing and that a global hierarchy of linguistic phenomena, built under the supervision of linguistics experts, would allow more insightful comparisons of model results across benchmarks.
Load-bearing premise
The survey's conclusion that no naming convention exists rests on treating the five examined datasets as representative of all NLU diagnostics benchmarks, and the paper does not state a systematic search strategy or inclusion criteria.
Editorial extensions
If this is right
- Without a shared naming convention, error analyses on different benchmarks cannot be aligned, so a reported weakness in "monotonicity" or "anaphora" may refer to different constructions depending on the benchmark.
- A standardized hierarchy would let researchers compare a model's failure profile across languages, since the survey covers English, Arabic, and Chinese diagnostics.
- Benchmark builders could use the hierarchy to check coverage, avoiding gaps such as CLUE's absence of fine-grained categories or FraCaS's lack of a Logic category.
- The reported class distributions (most samples are entailment) show that any evaluation standard would need to address class balance and minimum sample size per phenomenon.
Reading between the lines
- The missing convention is a reproducibility problem: two papers reporting a model's weakness on "anaphora" may be testing different linguistic constructions, so the survey's complaint matters beyond aesthetics.
- A testable extension would be to construct the proposed mapping table aligning micro-categories across the five taxonomies; success would show the issue is terminological, failure would show the taxonomies are genuinely incommensurable.
- The survey's own statistics suggest a standard would need to prescribe not just category names but also class balance and per-phenomenon sample counts, since most diagnostics samples fall in the entailment class and sizes vary widely.
- With a shared hierarchy, the field could ask whether model error profiles are language-specific or universal, since the same phenomena (monotonicity, quantifiers, anaphora) appear in English, Arabic, and Chinese diagnostics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys NLU benchmarks that contain diagnostics datasets, with a focus on English, Arabic, and multilingual benchmarks. It compares the macro- and micro-level linguistic phenomena covered by five diagnostics resources—FraCaS, a specialized TE dataset, GLUE, ALUE, and CLUE—and provides statistics on sample counts and class distributions. The authors' main claim is that there is no naming convention for macro and micro categories, nor a standard set of linguistic phenomena, across NLU diagnostics benchmarks. Based on this observed gap, they pose an open research question about why no evaluation standard exists for diagnostics benchmarks and suggest that a global hierarchy of linguistic phenomena be built under the supervision of linguistics experts.
Significance. If the central observation were established, the paper would provide a useful and timely call for standardization in an area where comparisons across model evaluations are genuinely difficult. The survey collates information that is otherwise scattered: Table 4 and Figures 7–8 give a compact numerical and visual summary of how five diagnostic resources differ in their category hierarchies, and the discussion of individual phenomena (e.g., Ellipsis, Anaphora, Monotonicity) highlights real terminological inconsistencies. The paper is a survey rather than a formal derivation, so it does not offer machine-checked proofs or code; its value lies in synthesis and in articulating a gap. However, the central claim is currently overgeneralized relative to the evidence, and the paper would need to narrow or systematically support the universal negative before the contribution can be endorsed as stated.
major comments (4)
- [§3, Tables 3–4; Abstract/Conclusion] The central claim, repeated in the abstract and conclusion, is a universal negative: 'there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena.' The evidence base is a qualitative comparison of only five datasets, with no systematic search strategy, inclusion/exclusion criteria, or screening protocol reported. Table 3 lists more than twenty NLU benchmarks but does not explain why only FraCaS, the TE specialized dataset, GLUE, ALUE, and CLUE count as diagnostics with analyzable taxonomies; diagnostic-style suites such as HANS and SuperGLUE's AX-b are neither analyzed nor explicitly excluded. If any excluded benchmark uses a shared or standardized taxonomy, the unqualified claim is false. The claim should be narrowed to 'among the five surveyed datasets' or supported by a systematic literature review with transparent criteria.
- [§2.3] The text states that GLUE is 'the first Diagnostics dataset' immediately after presenting FraCaS (1996) as the first linguistic-phenomenon hierarchy and the 2010 specialized TE methodology as a later proposal that discusses linguistic phenomena. This is internally inconsistent and obscures what the authors mean by 'diagnostics dataset.' Since Table 4 treats FraCaS and the TE specialized dataset as diagnostics, the paper should either clarify why GLUE is called the first, or revise the historical claim and define the term 'diagnostics dataset' precisely.
- [Table 4, §2.3] The specialized TE row in Table 4 totals 205 entries (32+18+44+67+44), while Section 2.3 says the Bentivogli et al. methodology was applied to a sample of 90 pairs from RTE-5. The reader cannot tell whether the table counts annotated phenomena rather than pairs, or whether each pair can contribute to multiple categories. The table should state its counting unit and overlap rule; as written, the statistics for TE Specialized cannot be reliably interpreted.
- [§4, Monotonicity] The discussion says 'Monotonicity was consistently included as a key phenomenon through all diagnostics datasets for decades.' However, the specialized TE dataset, as described in Table 4 and Figure 4, does not list Monotonicity among its categories. Unless the authors can point to monotonicity examples within one of its macro-categories (e.g., Reasoning), this statement is an overgeneralization that undermines confidence in the comparative analysis. Please provide the evidence or revise the claim.
minor comments (4)
- [Throughout] The spelling 'FraCas' and 'FraCaS' is used inconsistently; please choose one form and use it consistently.
- [Figure 8] Figure 8 is described as showing the class distribution in the studied diagnostics datasets, but it displays only four of the five datasets and omits CLUE; either add CLUE or adjust the caption and the surrounding text.
- [References] References [4] and [9] duplicate the same RTE-5 entry, and references [7] and [13] duplicate the same Fourth PASCAL RTE Challenge entry; these should be merged or distinguished clearly.
- [§3 and §4] The paper uses 'SoTA' without expanding the abbreviation at first use; please define it in the introduction or in a list of abbreviations.
Circularity Check
No circularity found: the survey contains no derivations or fitted predictions; its central claim is an inductive observation about five external datasets, and the sole self-citation (ArNLI) is a non-load-bearing list entry.
full rationale
This paper is a literature survey and comparison; it contains no equations, no fitted parameters, and no predictions, so there is no derivation chain that could reduce to its own inputs. The central claim — that 'there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered' — is presented as an observed gap derived from a qualitative comparison of five diagnostics datasets (FraCaS, TE specialized, GLUE, ALUE, CLUE) in Section 3 and Tables 3-4. That is an inductive empirical generalization, not a circular construction: it is not defined into existence, since 'diagnostics dataset' is defined independently as 'a specialized evaluation dataset that is used by humans to pinpoint specific areas where models struggle,' and the macro/micro category labels are taken from the external datasets themselves. The only self-citation with overlapping authorship is [35] (ArNLI), which appears in Section 2.1 in a plain list of available Arabic NLI datasets ('some available datasets are: ArbTEDS corpus [34], ArNLI dataset [35], ArEntil Dataset [36]'); it supports none of the paper's comparative claims, statistics, or conclusions, so it is not load-bearing. No uniqueness theorem, ansatz, or renamed result is imported from the authors' prior work. The paper even flags its own limitation ('we have just made initial statistics on macro/micro categories counts... but it still lacks important evaluation criteria'), showing the statistics are offered as descriptive, not as validated predictions. The skeptic's concern — that the negative universal claim rests on only five hand-selected datasets with no systematic search or inclusion criteria — is a coverage and correctness critique, not a circularity one, and per the review rules it cannot raise the circularity score without an exhibited reduction. Consequently the circularity burden is nil and the score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The five studied diagnostics datasets are representative of all NLU diagnostics benchmarks in English, Arabic, and multilingual settings.
- domain assumption The macro and micro category labels and their counts, as reported by the original benchmark papers, are accurate and can be compared across datasets.
Cite this review
Pith. "Pith review of Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?." pith.science (2026). https://pith.science/paper/3R6CHEEB
@misc{pith2026250720419,
author = {Pith},
title = {Pith review of: Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?},
year = {2026},
howpublished = {\url{https://pith.science/paper/3R6CHEEB}},
note = {Machine review of arXiv:2507.20419}
}
read the original abstract
Natural Language Understanding (NLU) is a basic task in Natural Language Processing (NLP). The evaluation of NLU capabilities has become a trending research topic that attracts researchers in the last few years, resulting in the development of numerous benchmarks. These benchmarks include various tasks and datasets in order to evaluate the results of pretrained models via public leaderboards. Notably, several benchmarks contain diagnostics datasets designed for investigation and fine-grained error analysis across a wide range of linguistic phenomena. This survey provides a comprehensive review of available English, Arabic, and Multilingual NLU benchmarks, with a particular emphasis on their diagnostics datasets and the linguistic phenomena they covered. We present a detailed comparison and analysis of these benchmarks, highlighting their strengths and limitations in evaluating NLU tasks and providing in-depth error analysis. When highlighting the gaps in the state-of-the-art, we noted that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered. Consequently, we formulated a research question regarding the evaluation metrics of the evaluation diagnostics benchmarks: "Why do not we have an evaluation standard for the NLU evaluation diagnostics benchmarks?" similar to ISO standard in industry. We conducted a deep analysis and comparisons of the covered linguistic phenomena in order to support experts in building a global hierarchy for linguistic phenomena in future. We think that having evaluation metrics for diagnostics evaluation could be valuable to gain more insights when comparing the results of the studied models on different diagnostics benchmarks.
Reference graph
Works this paper leans on
-
[33]
SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,
F. Iandola, A. Shaw, R. Krishna, and K. Keutzer, “SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,” in Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, N. S. Moosavi, A. Fan, V. Shwartz, G. Glavaš, S. Joty, A. Wang, and T. Wolf, Eds., Online: Association for Computational Linguistics,...
-
[1]
Principles of Evaluation in Natural Language Processing,
P. Paroubek, S. Chaudiron, and L. Hirschman, “Principles of Evaluation in Natural Language Processing,” in Traitement Automatique des Langues, Volume 48, Numéro 1 : Principes de l’évaluation en Traitement Automatique des Langues [Principles of Evaluation in Natural Language Processing], P. Paroubek, S. Chaudiron, and L. Hirschman, Eds., France: ATALA (Ass...
2007
-
[2]
Natural Language Inference ,
Y. A. Wilks, “Natural Language Inference ,” Aug. 1973
1973
-
[3]
PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,
I. Dagan and O. Glickman, “PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,” 2004. [Online]. Available: https://api.semanticscholar.org/CorpusID:17200692
2004
-
[5]
Probabilistic textual entailment: Generic applied modeling of language variability,
O. Glickman and I. Dagan, “Probabilistic textual entailment: Generic applied modeling of language variability,” in Proceedings of the Workshop on Learning Methods for Text Understanding and Mining, 2004
2004
-
[6]
The Seventh PASCAL Recognizing Textual Entailment Challenge,
L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo, “The Seventh PASCAL Recognizing Textual Entailment Challenge,” Theory and Applications of Categories, 2011, [Online]. Available: https://api.semanticscholar.org/CorpusID:5791809
2011
-
[8]
The Sixth PASCAL Recognizing Textual Entailment Challenge,
L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo, “The Sixth PASCAL Recognizing Textual Entailment Challenge,” in Text Analysis Conference, 2009. [Online]. Available: https://api.semanticscholar.org/CorpusID:858065
2009
-
[9]
The Fifth PASCAL Recognizing Textual Entailment Challenge,
L. Bentivogli, B. Magnini, I. Dagan, H. T. Dang, and D. Giampiccolo, “The Fifth PASCAL Recognizing Textual Entailment Challenge,” in Proceedings of the Second Text Analysis Conference, TAC 2009, Gaithersburg, Maryland, USA, November 16-17, 2009, NIST, 2009. [Online]. Available: https://tac.nist.gov/publications/2009/additional.papers/RTE5_overview.proceedings.pdf
2009
Show all 94 references
-
[10]
The Third PASCAL Recognizing Textual Entailment Challenge,
D. Giampiccolo, B. Magnini, I. Dagan, and B. Dolan, “The Third PASCAL Recognizing Textual Entailment Challenge,” in Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, S. Sekine, K. Inui, I. Dagan, B. Dolan, D. Giampiccolo, and B. Magnini, Eds., Prag...
2007
-
[11]
The Second PASCAL Recognising Textual Entailment Challenge,
R. Bar-Haim et al., “The Second PASCAL Recognising Textual Entailment Challenge,” 2006. [Online]. Available: https://api.semanticscholar.org/CorpusID:13385138
2006
-
[12]
The PASCAL Recognising Textual Entailment Challenge,
O. and M. B. Dagan Ido and Glickman, “The PASCAL Recognising Textual Entailment Challenge,” in Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, I. and M. B. and d’Alché-B. F. Quiñonero-Candela Joa...
2006
-
[13]
The Fourth PASCAL Recognizing Textual Entailment Challenge,
D. Giampiccolo, H. T. Dang, B. Magnini, I. Dagan, E. Cabrio, and W. B. Dolan, “The Fourth PASCAL Recognizing Textual Entailment Challenge,” in Text Analysis Conference, 2008. [Online]. Available: https://api.semanticscholar.org/CorpusID:12381965
2008
-
[14]
Recognizing Textual Entailment: Models and Applications,
I. Dagan, D. Roth, M. Sammons, and F. M. Zanzotto, “Recognizing Textual Entailment: Models and Applications,” Synthesis Lectures on Human Language Technologies, vol. 6, pp. 1–220, Jul. 2013, doi: 10.2200/S00509ED1V01Y201305HLT023
2013 doi
-
[15]
The Winograd Schema Challenge,
H. J. Levesque, E. Davis, and L. Morgenstern, “The Winograd Schema Challenge,” in AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, 2011. [Online]. Available: https://api.semanticscholar.org/CorpusID:15710851
2011
-
[16]
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,
N. and B. S. Williams Adina and Nangia, “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1...
2018
-
[17]
A large annotated corpus for learning natural language inference,
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning, “A large annotated corpus for learning natural language inference,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, L. Màrquez, C. Callison-Burch, and J. Su, Eds., Lisbon, Portugal...
2015 doi
-
[18]
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Brussels, Belgium...
2018 doi
-
[19]
PaLM: Scaling Language Modeling with Pathways,
A. Chowdhery et al., “PaLM: Scaling Language Modeling with Pathways,” ArXiv, vol. abs/2204.02311, 2022, [Online]. Available: https://api.semanticscholar.org/CorpusID:247951931
2022 arXiv
-
[20]
Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE,
Q. Zhong et al., “Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE,” ArXiv, vol. abs/2212.01853, 2022, [Online]. Available: https://api.semanticscholar.org/CorpusID:254246784
2022 arXiv
-
[21]
RoBERTa: A Robustly Optimized BERT Pretraining Approach,
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” 2019
2019
-
[22]
Semantics-aware BERT for Language Understanding,
Z. Zhang et al., “Semantics-aware BERT for Language Understanding,” in AAAI Conference on Artificial Intelligence, 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:202539891
2019
-
[23]
XLNet: Generalized Autoregressive Pretraining for Language Understanding,
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V Le, “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” in Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alché-Buc, E. Fox, and R....
2019
-
[24]
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model,
S. Soltan et al., “AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model,” ArXiv, vol. abs/2208.01448, 2022, [Online]. Available: https://api.semanticscholar.org/CorpusID:251253416
2022 arXiv
-
[25]
BloombergGPT: A Large Language Model for Finance,
S. Wu et al., “BloombergGPT: A Large Language Model for Finance,” ArXiv, vol. abs/2303.17564, 2023, [Online]. Available: https://api.semanticscholar.org/CorpusID:257833842
2023 arXiv
- [26]
-
[27]
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,
L. Xue et al., “ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,” Trans Assoc Comput Linguist, vol. 10, pp. 291–306, 2022, doi: 10.1162/tacl_a_00461
2022 doi
-
[28]
Rethinking embedding coupling in pre-trained language models,
H. W. Chung, T. Févry, H. Tsai, M. Johnson, and S. Ruder, “Rethinking embedding coupling in pre-trained language models,” ArXiv, vol. abs/2010.12821, 2020, [Online]. Available: https://api.semanticscholar.org/CorpusID:225067567
2010 arXiv
-
[29]
mGPT: Few-Shot Learners Go Multilingual,
O. Shliazhko, A. Fenogenova, M. Tikhonova, A. Kozlova, V. Mikhailov, and T. Shavrina, “mGPT: Few-Shot Learners Go Multilingual,” Trans Assoc Comput Linguist, vol. 12, pp. 58–79, 2024, doi: 10.1162/tacl_a_00633
2024 doi
-
[30]
DeBERTa: Decoding-enhanced BERT with Disentangled Attention,
P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with Disentangled Attention,” 2021
2021
-
[31]
Available: https://aclanthology.org/2007.tal-1.1
[Online]. Available: https://aclanthology.org/2007.tal-1.1
2007
-
[32]
SpanBERT: Improving Pre-training by Representing and Predicting Spans,
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “SpanBERT: Improving Pre-training by Representing and Predicting Spans,” Trans Assoc Comput Linguist, vol. 8, pp. 64–77, 2020, doi: 10.1162/tacl_a_00300
2020 doi
-
[34]
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,
V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” ArXiv, vol. abs/1910.01108, 2019, [Online]. Available: https://api.semanticscholar.org/CorpusID:203626972
1910 arXiv
-
[35]
A Dataset for Arabic Textual Entailment,
M. Alabbas, “A Dataset for Arabic Textual Entailment,” in Proceedings of the Student Research Workshop associated with RANLP 2013, I. Temnikova, I. Nikolova, and N. Konstantinova, Eds., Hissar, Bulgaria: INCOMA Ltd. Shoumen, BULGARIA, Sep. 2013, pp. 7–13. [Online]. Available: ...
2013
-
[36]
ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,
K. Al Jallad and N. Ghneim, “ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,” Computer Science, vol. 24, no. 2, Mar. 2023, doi: 10.7494/csci.2023.24.2.4378
2023 doi
-
[37]
ArEntail: manually-curated Arabic natural language inference dataset from news headlines,
R. Obeidat, Y. Al-Harahsheh, M. Al-Ayyoub, and M. Gharaibeh, “ArEntail: manually-curated Arabic natural language inference dataset from news headlines,” Lang Resour Eval, 2024, doi: 10.1007/s10579-024-09731-1
2024 doi
-
[38]
Baselines and Test Data for Cross-Lingual Inference,
Ž. Agić and N. Schluter, “Baselines and Test Data for Cross-Lingual Inference,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), N. Calzolari, K. Choukri, C. Cieri, T. Declerck, S. Goggi, K. Hasida, H. Isahara, B. Maegaa...
2018
-
[39]
XNLI: Evaluating Cross-lingual Sentence Representations,
A. Conneau et al., “XNLI: Evaluating Cross-lingual Sentence Representations,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds., Brussels, Belgium: Association for Computational ...
2018 doi
-
[40]
Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,
W. Yin, J. Hay, and D. Roth, “Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pro...
2019 doi
-
[41]
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,
M. Lewis et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault...
2020 doi
-
[42]
RoBERTa: A Robustly Optimized BERT Pretraining Approach,
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” ArXiv, vol. abs/1907.11692, 2019, [Online]. Available: https://api.semanticscholar.org/CorpusID:198953378
1907 arXiv
-
[43]
MiniLM: Deep Self-Attention Distillation for Task- Agnostic Compression of Pre-Trained Transformers,
W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “MiniLM: Deep Self-Attention Distillation for Task- Agnostic Compression of Pre-Trained Transformers,” ArXiv, vol. abs/2002.10957, 2020, [Online]. Available: https://api.semanticscholar.org/CorpusID:211296536
2002 arXiv
-
[44]
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient- Disentangled Embedding Sharing,
P. He, J. Gao, and W. Chen, “DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient- Disentangled Embedding Sharing,” ArXiv, vol. abs/2111.09543, 2021, [Online]. Available: https://api.semanticscholar.org/CorpusID:244346093
2021 arXiv
-
[45]
Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,
M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers, “Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,” Political Analysis, vol. 32, no. 1, pp. 84–100, Jan. 2024, doi: 10.1017/pan.2023.20
2024 doi
-
[46]
GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,
M. T. I. Khondaker, A. Waheed, E. M. B. Nagoudi, and M. Abdul-Mageed, “GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore...
2023 doi
-
[47]
CLSE: Corpus of Linguistically Significant Entities,
A. Chuklin, J. Zhao, and M. Kale, “CLSE: Corpus of Linguistically Significant Entities,” in Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM), A. Bosselut, K. Chandu, K. Dhole, V. Gangal, S. Gehrmann, Y. Jernite, J. Novikova, and L. ...
2022 doi
-
[48]
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,
S. Gehrmann et al., “The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,” in Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), A. Bosselut, E. Durmus, V. P. Gangal, S. Gehrmann, Y. Jernite, L. Perez-Belt...
2021 doi
-
[49]
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,
S. Gehrmann et al., “GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,” in EMNLP 2022 - 2022 Conference on Empirical Methods in Natural Language Processing, W. Che and E. Shutova, Eds., Association for Computational Linguistics (ACL), Dec. 2022, pp. 266–281. doi: ...
2022 doi
-
[50]
IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,
A. Kumar et al., “IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds., Abu Dhabi, United Arab Emirates: Asso...
2022 doi
-
[51]
MTG: A Benchmark Suite for Multilingual Text Generation,
Y. Chen et al., “MTG: A Benchmark Suite for Multilingual Text Generation,” in Findings of the Association for Computational Linguistics: NAACL 2022, M. Carpuat, M.-C. de Marneffe, and I. V. Meza Ruiz, Eds., Seattle, United States: Association for Computational Linguistics, Jul...
2022 doi
-
[52]
IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,
S. Cahyawijaya et al., “IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W. Yih, Eds., Online and Punta C...
2021 doi
-
[53]
Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,
E. M. B. Nagoudi, A. Elmadany, A. El-Shangiti, and M. Abdul-Mageed, “Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,” in Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore: Association for Compu...
2023
-
[54]
TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,
E. M. B. Nagoudi, A. Elmadany, and M. Abdul-Mageed, “TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,” in Proceedinsg of the 5th Workshop on Open-Source Arabic Corpora and Processing Tools with Shared Tasks on Qur`an QA and Fine-Grained Hate Speech Detection, ...
2022
-
[55]
AraBench: Benchmarking Dialectal Arabic-English Machine Translation,
H. Sajjad, A. Abdelali, N. Durrani, and F. Dalvi, “AraBench: Benchmarking Dialectal Arabic-English Machine Translation,” in Proceedings of the 28th International Conference on Computational Linguistics, D. Scott, N. Bel, and C. Zong, Eds., Barcelona, Spain (Online): Internatio...
2020 doi
-
[56]
BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,
A. Bhattacharjee, T. Hasan, W. U. Ahmad, and R. Shahriyar, “BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,” in Findings of the Association for Computational Linguistics: EACL 2023, A. Vlachos and I. Augenstei...
2023 doi
-
[57]
AraT5: Text-to-Text Transformers for Arabic Language Generation,
E. M. B. Nagoudi, A. Elmadany, and M. Abdul-Mageed, “AraT5: Text-to-Text Transformers for Arabic Language Generation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio...
2022 doi
-
[58]
CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark,
Y. Yao et al., “CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark,” 2022. [Online]. Available: https://arxiv.org/abs/2112.13610
2022 arXiv
-
[59]
PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,
L. Doan, L. T. Nguyen, N. L. Tran, T. Hoang, and D. Q. Nguyen, “PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang...
2021 doi
-
[60]
Benchmarking Multidomain English-Indonesian Machine Translation,
T. W. Guntara, A. F. Aji, and R. E. Prasojo, “Benchmarking Multidomain English-Indonesian Machine Translation,” in Proceedings of the 13th Workshop on Building and Using Comparable Corpora, R. Rapp, P. Zweigenbaum, and S. Sharoff, Eds., Marseille, France: European Language Res...
2020
-
[61]
LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,
J. Guan et al., “LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,” Trans Assoc Comput Linguist, vol. 10, pp. 434–451, 2022, doi: 10.1162/tacl_a_00469
2022 doi
-
[62]
XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,
Y. Liang et al., “XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association f...
2020
-
[63]
GLGE: A New General Language Generation Evaluation Benchmark,
D. Liu et al., “GLGE: A New General Language Generation Evaluation Benchmark,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, C. Zong, F. Xia, W. Li, and R. Navigli, Eds., Online: Association for Computational Linguistics, Aug. 2021, pp. 408–420...
2021 doi
-
[64]
XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,
S. Ruder et al., “XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W. Yih, Eds., Online and Punta Cana, Dominican Republi...
2021 doi
-
[65]
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,
A. Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA: Curran Associates Inc., 2019
2019
-
[66]
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,
J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,” 2020
2020
-
[67]
ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,
M. Abdul-Mageed, A. Elmadany, and E. M. B. Nagoudi, “ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Proces...
2021 doi
-
[68]
ORCA: A Challenging Benchmark for Arabic Language Understanding,
A. Elmadany, E. M. B. Nagoudi, and M. Abdul-Mageed, “ORCA: A Challenging Benchmark for Arabic Language Understanding,” in Annual Meeting of the Association for Computational Linguistics, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:254926875
2022
-
[69]
ALUE: Arabic Language Understanding Evaluation,
H. Seelawi et al., “ALUE: Arabic Language Understanding Evaluation,” in Proceedings of the Sixth Arabic Natural Language Processing Workshop, N. Habash, H. Bouamor, H. Hajj, W. Magdy, W. Zaghouani, F. Bougares, N. Tomeh, I. Abu Farha, and S. Touileb, Eds., Kyiv, Ukraine (Virtu...
2021
-
[70]
LAraBench: Benchmarking Arabic AI with Large Language Models,
A. Abdelali et al., “LAraBench: Benchmarking Arabic AI with Large Language Models,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Y. Graham and M. Purver, Eds., St. Julian’s, Malta: Assoc...
2024
-
[71]
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks,
J. Weston, A. Bordes, S. Chopra, and T. Mikolov, “Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks,” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2- 4, 2016, Conference Track Proceedings, Y. Bengio an...
2016 arXiv
-
[72]
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,
F. Koto et al., “ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Srikumar, Eds., Bangkok, Thailand: Association for Computational Linguistics, Aug. 2...
2024 doi
-
[73]
KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,
J. Ham, Y. J. Choe, K. Park, I. Choi, and H. Soh, “KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association for Computat...
2020 doi
-
[74]
CLUE: A Chinese Language Understanding Evaluation Benchmark,
L. Xu et al., “CLUE: A Chinese Language Understanding Evaluation Benchmark,” in Proceedings of the 28th International Conference on Computational Linguistics, D. Scott, N. Bel, and C. Zong, Eds., Barcelona, Spain (Online): International Committee on Computational Linguistics, ...
2020 doi
-
[75]
JGLUE: Japanese General Language Understanding Evaluation,
K. Kurihara, D. Kawahara, and T. Shibata, “JGLUE: Japanese General Language Understanding Evaluation,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, Marseille, France: European Language Resources Association, Jun. 2022, pp. 2957–2966. [Online]....
2022
-
[76]
FlauBERT: Unsupervised Language Model Pre-training for French,
H. Le et al., “FlauBERT: Unsupervised Language Model Pre-training for French,” in Proceedings of the Twelfth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. M...
2020
-
[77]
KLUE: Korean Language Understanding Evaluation,
S. Park et al., “KLUE: Korean Language Understanding Evaluation,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. [Online]. Available: https://openreview.net/forum?id=q-8h8-LZiUm
2021
-
[78]
UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,
V. Basile, L. Bioglio, A. Bosca, C. Bosco, and V. Patti, “UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), D. Bollegala, R. Hu...
2023 doi
-
[79]
SuperGLEBer: German Language Understanding Evaluation Benchmark,
J. Pfister and A. Hotho, “SuperGLEBer: German Language Understanding Evaluation Benchmark,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gom...
2024 doi
-
[80]
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,
B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natura...
2020 doi
-
[81]
IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,
D. Kakwani et al., “IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association for...
2020
-
[82]
Using the Framework,
R. Cooper et al., “Using the Framework,” Mar. 1996. Accessed: Aug. 02, 2024. [Online]. Available: https://files.ifi.uzh.ch/cl/hess/classes/seminare/interface/framework.pdf
1996
-
[83]
An extended model of natural logic
C. D. Manning and B. MacCartney, “An extended model of natural logic”
-
[84]
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,
P. N.-T. Do, S. Q. Tran, P. G. Hoang, K. Van Nguyen, and N. L.-T. Nguyen, “VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,” in Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gome...
2024 doi
-
[85]
Analysis of identifying linguistic phenomena for recognizing inference in text,
M.-Y. Day and Y.-J. Wang, “Analysis of identifying linguistic phenomena for recognizing inference in text,” Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration, IEEE IRI 2014, pp. 607–612, Mar. 2015, doi: 10.1109/IRI.2014.7051945
2014
-
[86]
EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,
A. Ravichander, A. Naik, C. Rose, and E. Hovy, “EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,” Mar. 2019, pp. 349–361. doi: 10.18653/v1/K19-1033
2019 doi
-
[87]
Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,
S. Sugawara, Y. Kido, H. Yokono, and A. Aizawa, “Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,” Mar. 2017, pp. 806–817. doi: 10.18653/v1/P17-1075
2017 doi
-
[88]
NaturalLI: Natural Logic Inference for Common Sense Reasoning,
G. Angeli and C. D. Manning, “NaturalLI: Natural Logic Inference for Common Sense Reasoning,” in Conference on Empirical Methods in Natural Language Processing, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:2854390
2014
-
[89]
Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,
L. Bentivogli, E. Cabrio, I. Dagan, D. Giampiccolo, M. Leggio, and B. Magnini, “Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,” Mar. 2010
2010
- [90]
-
[91]
Neural Networks and Textual Inference: How did we get here and where do we go now?,
I. Cases and L. Karttunen, “Neural Networks and Textual Inference: How did we get here and where do we go now?,” Sep. 2017
2017
-
[94]
On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,
A. Poliak, Y. Belinkov, J. Glass, and B. Van Durme, “On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: H...
2018 doi
-
[520]
Available: https://aclanthology.org/2024.eacl-long.30/
[Online]. Available: https://aclanthology.org/2024.eacl-long.30/
2024
-
[1422]
doi: 10.18653/v1/2023.findings-emnlp.98
2023 doi
-
[4961]
doi: 10.18653/v1/2020.findings-emnlp.445
2020 doi
-
[6018]
doi: 10.18653/v1/2020.emnlp-main.484
2020 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.