Pith. sign in

REVIEW 4 major objections 4 minor 94 references

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper argues that NLU diagnostics benchmarks have no shared naming convention for their linguistic categories, making model error analysis incomparable across benchmarks, and proposes building a global hierarchy of linguistic…

desk verdict Useful taxonomy comparison of five diagnostics benchmarks, but the field-wide 'no naming convention' claim outruns the evidence. read the letter →

arxiv 2507.20419 v1 pith:3R6CHEEB submitted 2025-07-27 cs.CL cs.AIcs.HCcs.LG

classification cs.CLcs.AIcs.HCcs.LG
keywords NLUdiagnosticsbenchmarkslinguisticphenomenataxonomymacroandmicrocategoriesnaturallanguageinferencetextualentailmentbenchmarkstandardizationerroranalysiscross-lingualevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey examines the diagnostics portions of five natural language understanding benchmarks—FraCaS, a specialized textual entailment dataset, GLUE, ALUE, and CLUE—and argues that the field has no naming convention for coarse-grained (macro) and fine-grained (micro) linguistic categories, nor an agreed list of phenomena that a diagnostics suite should cover. The authors show that the same phenomenon is treated at different levels in different taxonomies: ellipsis is a macro-category in FraCaS but a micro-category within Predicate-Argument Structure in GLUE and ALUE, while CLUE offers no fine-grained categories at all. A sympathetic reader would care because, without a common taxonomy, a model's error profile on one benchmark cannot be compared against its profile on another, defeating the purpose of diagnostics. The paper therefore poses the research question of why no evaluation standard for diagnostics benchmarks exists, and it calls for linguistics experts to build a global hierarchy of linguistic phenomena.

What carries the argument

The carrying objects are the diagnostics datasets themselves and their two-level taxonomies: macro-categories such as Lexical Semantics, Predicate-Argument Structure, Logic, Quantification, and Knowledge and Common Sense, each subdivided into micro-categories. The comparative method is to tabulate which phenomena each benchmark covers, at which level of the hierarchy, with how many sentence pairs, and with what distribution of entailment classes (entailment, neutral/unknown, contradiction). The taxonomies are also the evidence: placing the same phenomenon at different levels in different datasets is what demonstrates the absence of a shared naming convention.

What would settle it

The claim that no naming convention exists would be settled by searching for an NLU diagnostics benchmark outside the five examined that publishes a fixed hierarchy of macro- and micro-categories with a defined phenomenon list; finding one would directly weaken the claim. A complementary test would be to attempt a faithful mapping of the five taxonomies onto a single hierarchy; if the mapping succeeds without residue, the problem is terminological, whereas if phenomena refuse to align, the taxonomies encode genuinely different analyses.

Watch

Extended reading notes

Core claim

The paper's central claim is that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena across NLU diagnostics benchmarks. Rather than a single finding, the claim is established by a comparative analysis showing systematic incommensurability: FraCaS keeps Quantifiers, Plurals, and Adjectives as separate macro-categories; GLUE and ALUE group the same material under Lexical Semantics or Predicate-Argument Structure; CLUE has only coarse-grained categories with no fine-grained structure; FraCaS lacks any Logic category while the other diagnostics include one; and world knowledge is an explicit category in GLUE and ALUE but deliberately implicit in FraCaS and the specialized TE dataset. The authors also report distributional statistics, including that the majority of samples in each dataset fall under the entailment class. From this comparison they conclude that an evaluation standard for diagnostics benchmarks is missing and that a global hierarchy of linguistic phenomena, built under the supervision of linguistics experts, would allow more insightful comparisons of model results across benchmarks.

Load-bearing premise

The survey's conclusion that no naming convention exists rests on treating the five examined datasets as representative of all NLU diagnostics benchmarks, and the paper does not state a systematic search strategy or inclusion criteria.

Editorial extensions

If this is right

  • Without a shared naming convention, error analyses on different benchmarks cannot be aligned, so a reported weakness in "monotonicity" or "anaphora" may refer to different constructions depending on the benchmark.
  • A standardized hierarchy would let researchers compare a model's failure profile across languages, since the survey covers English, Arabic, and Chinese diagnostics.
  • Benchmark builders could use the hierarchy to check coverage, avoiding gaps such as CLUE's absence of fine-grained categories or FraCaS's lack of a Logic category.
  • The reported class distributions (most samples are entailment) show that any evaluation standard would need to address class balance and minimum sample size per phenomenon.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The missing convention is a reproducibility problem: two papers reporting a model's weakness on "anaphora" may be testing different linguistic constructions, so the survey's complaint matters beyond aesthetics.
  • A testable extension would be to construct the proposed mapping table aligning micro-categories across the five taxonomies; success would show the issue is terminological, failure would show the taxonomies are genuinely incommensurable.
  • The survey's own statistics suggest a standard would need to prescribe not just category names but also class balance and per-phenomenon sample counts, since most diagnostics samples fall in the entailment class and sizes vary widely.
  • With a shared hierarchy, the field could ask whether model error profiles are language-specific or universal, since the same phenomena (monotonicity, quantifiers, anaphora) appear in English, Arabic, and Chinese diagnostics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper surveys NLU benchmarks that contain diagnostics datasets, with a focus on English, Arabic, and multilingual benchmarks. It compares the macro- and micro-level linguistic phenomena covered by five diagnostics resources—FraCaS, a specialized TE dataset, GLUE, ALUE, and CLUE—and provides statistics on sample counts and class distributions. The authors' main claim is that there is no naming convention for macro and micro categories, nor a standard set of linguistic phenomena, across NLU diagnostics benchmarks. Based on this observed gap, they pose an open research question about why no evaluation standard exists for diagnostics benchmarks and suggest that a global hierarchy of linguistic phenomena be built under the supervision of linguistics experts.

Significance. If the central observation were established, the paper would provide a useful and timely call for standardization in an area where comparisons across model evaluations are genuinely difficult. The survey collates information that is otherwise scattered: Table 4 and Figures 7–8 give a compact numerical and visual summary of how five diagnostic resources differ in their category hierarchies, and the discussion of individual phenomena (e.g., Ellipsis, Anaphora, Monotonicity) highlights real terminological inconsistencies. The paper is a survey rather than a formal derivation, so it does not offer machine-checked proofs or code; its value lies in synthesis and in articulating a gap. However, the central claim is currently overgeneralized relative to the evidence, and the paper would need to narrow or systematically support the universal negative before the contribution can be endorsed as stated.

major comments (4)
  1. [§3, Tables 3–4; Abstract/Conclusion] The central claim, repeated in the abstract and conclusion, is a universal negative: 'there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena.' The evidence base is a qualitative comparison of only five datasets, with no systematic search strategy, inclusion/exclusion criteria, or screening protocol reported. Table 3 lists more than twenty NLU benchmarks but does not explain why only FraCaS, the TE specialized dataset, GLUE, ALUE, and CLUE count as diagnostics with analyzable taxonomies; diagnostic-style suites such as HANS and SuperGLUE's AX-b are neither analyzed nor explicitly excluded. If any excluded benchmark uses a shared or standardized taxonomy, the unqualified claim is false. The claim should be narrowed to 'among the five surveyed datasets' or supported by a systematic literature review with transparent criteria.
  2. [§2.3] The text states that GLUE is 'the first Diagnostics dataset' immediately after presenting FraCaS (1996) as the first linguistic-phenomenon hierarchy and the 2010 specialized TE methodology as a later proposal that discusses linguistic phenomena. This is internally inconsistent and obscures what the authors mean by 'diagnostics dataset.' Since Table 4 treats FraCaS and the TE specialized dataset as diagnostics, the paper should either clarify why GLUE is called the first, or revise the historical claim and define the term 'diagnostics dataset' precisely.
  3. [Table 4, §2.3] The specialized TE row in Table 4 totals 205 entries (32+18+44+67+44), while Section 2.3 says the Bentivogli et al. methodology was applied to a sample of 90 pairs from RTE-5. The reader cannot tell whether the table counts annotated phenomena rather than pairs, or whether each pair can contribute to multiple categories. The table should state its counting unit and overlap rule; as written, the statistics for TE Specialized cannot be reliably interpreted.
  4. [§4, Monotonicity] The discussion says 'Monotonicity was consistently included as a key phenomenon through all diagnostics datasets for decades.' However, the specialized TE dataset, as described in Table 4 and Figure 4, does not list Monotonicity among its categories. Unless the authors can point to monotonicity examples within one of its macro-categories (e.g., Reasoning), this statement is an overgeneralization that undermines confidence in the comparative analysis. Please provide the evidence or revise the claim.
minor comments (4)
  1. [Throughout] The spelling 'FraCas' and 'FraCaS' is used inconsistently; please choose one form and use it consistently.
  2. [Figure 8] Figure 8 is described as showing the class distribution in the studied diagnostics datasets, but it displays only four of the five datasets and omits CLUE; either add CLUE or adjust the caption and the surrounding text.
  3. [References] References [4] and [9] duplicate the same RTE-5 entry, and references [7] and [13] duplicate the same Fourth PASCAL RTE Challenge entry; these should be merged or distinguished clearly.
  4. [§3 and §4] The paper uses 'SoTA' without expanding the abbreviation at first use; please define it in the introduction or in a list of abbreviations.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the survey contains no derivations or fitted predictions; its central claim is an inductive observation about five external datasets, and the sole self-citation (ArNLI) is a non-load-bearing list entry.

full rationale

This paper is a literature survey and comparison; it contains no equations, no fitted parameters, and no predictions, so there is no derivation chain that could reduce to its own inputs. The central claim — that 'there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered' — is presented as an observed gap derived from a qualitative comparison of five diagnostics datasets (FraCaS, TE specialized, GLUE, ALUE, CLUE) in Section 3 and Tables 3-4. That is an inductive empirical generalization, not a circular construction: it is not defined into existence, since 'diagnostics dataset' is defined independently as 'a specialized evaluation dataset that is used by humans to pinpoint specific areas where models struggle,' and the macro/micro category labels are taken from the external datasets themselves. The only self-citation with overlapping authorship is [35] (ArNLI), which appears in Section 2.1 in a plain list of available Arabic NLI datasets ('some available datasets are: ArbTEDS corpus [34], ArNLI dataset [35], ArEntil Dataset [36]'); it supports none of the paper's comparative claims, statistics, or conclusions, so it is not load-bearing. No uniqueness theorem, ansatz, or renamed result is imported from the authors' prior work. The paper even flags its own limitation ('we have just made initial statistics on macro/micro categories counts... but it still lacks important evaluation criteria'), showing the statistics are offered as descriptive, not as validated predictions. The skeptic's concern — that the negative universal claim rests on only five hand-selected datasets with no systematic search or inclusion criteria — is a coverage and correctness critique, not a circularity one, and per the review rules it cannot raise the circularity score without an exhibited reduction. Consequently the circularity burden is nil and the score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper makes no derivations and fits no parameters, so the free parameter and invented entity lists are empty. Its central claim relies on the representativeness of the selected benchmarks and on the accuracy of the original papers' category labels.

assumptions (2)
  • domain assumption The five studied diagnostics datasets are representative of all NLU diagnostics benchmarks in English, Arabic, and multilingual settings.
    The survey selects these datasets without a stated systematic methodology; the no-standard conclusion depends on this coverage assumption. Section 3 and Tables 3 and 4.
  • domain assumption The macro and micro category labels and their counts, as reported by the original benchmark papers, are accurate and can be compared across datasets.
    The statistics in Table 4 and Figures 7 and 8 are drawn from the source benchmarks without independent reannotation, so comparability relies on cross-dataset label consistency. Section 3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?." pith.science (2026). https://pith.science/paper/3R6CHEEB

@misc{pith2026250720419,
  author       = {Pith},
  title        = {Pith review of: Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3R6CHEEB}},
  note         = {Machine review of arXiv:2507.20419}
}
read the original abstract

Natural Language Understanding (NLU) is a basic task in Natural Language Processing (NLP). The evaluation of NLU capabilities has become a trending research topic that attracts researchers in the last few years, resulting in the development of numerous benchmarks. These benchmarks include various tasks and datasets in order to evaluate the results of pretrained models via public leaderboards. Notably, several benchmarks contain diagnostics datasets designed for investigation and fine-grained error analysis across a wide range of linguistic phenomena. This survey provides a comprehensive review of available English, Arabic, and Multilingual NLU benchmarks, with a particular emphasis on their diagnostics datasets and the linguistic phenomena they covered. We present a detailed comparison and analysis of these benchmarks, highlighting their strengths and limitations in evaluating NLU tasks and providing in-depth error analysis. When highlighting the gaps in the state-of-the-art, we noted that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered. Consequently, we formulated a research question regarding the evaluation metrics of the evaluation diagnostics benchmarks: "Why do not we have an evaluation standard for the NLU evaluation diagnostics benchmarks?" similar to ISO standard in industry. We conducted a deep analysis and comparisons of the covered linguistic phenomena in order to support experts in building a global hierarchy for linguistic phenomena in future. We think that having evaluation metrics for diagnostics evaluation could be valuable to gain more insights when comparing the results of the studied models on different diagnostics benchmarks.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

94 extracted references · 45 canonical work pages

  1. [33]

    SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,

    F. Iandola, A. Shaw, R. Krishna, and K. Keutzer, “SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,” in Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, N. S. Moosavi, A. Fan, V. Shwartz, G. Glavaš, S. Joty, A. Wang, and T. Wolf, Eds., Online: Association for Computational Linguistics,...

  2. [1]

    Principles of Evaluation in Natural Language Processing,

    P. Paroubek, S. Chaudiron, and L. Hirschman, “Principles of Evaluation in Natural Language Processing,” in Traitement Automatique des Langues, Volume 48, Numéro 1 : Principes de l’évaluation en Traitement Automatique des Langues [Principles of Evaluation in Natural Language Processing], P. Paroubek, S. Chaudiron, and L. Hirschman, Eds., France: ATALA (Ass...

  3. [2]

    Natural Language Inference ,

    Y. A. Wilks, “Natural Language Inference ,” Aug. 1973

  4. [3]

    PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,

    I. Dagan and O. Glickman, “PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,” 2004. [Online]. Available: https://api.semanticscholar.org/CorpusID:17200692

  5. [5]

    Probabilistic textual entailment: Generic applied modeling of language variability,

    O. Glickman and I. Dagan, “Probabilistic textual entailment: Generic applied modeling of language variability,” in Proceedings of the Workshop on Learning Methods for Text Understanding and Mining, 2004

  6. [6]

    The Seventh PASCAL Recognizing Textual Entailment Challenge,

    L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo, “The Seventh PASCAL Recognizing Textual Entailment Challenge,” Theory and Applications of Categories, 2011, [Online]. Available: https://api.semanticscholar.org/CorpusID:5791809

  7. [8]

    The Sixth PASCAL Recognizing Textual Entailment Challenge,

    L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo, “The Sixth PASCAL Recognizing Textual Entailment Challenge,” in Text Analysis Conference, 2009. [Online]. Available: https://api.semanticscholar.org/CorpusID:858065

  8. [9]

    The Fifth PASCAL Recognizing Textual Entailment Challenge,

    L. Bentivogli, B. Magnini, I. Dagan, H. T. Dang, and D. Giampiccolo, “The Fifth PASCAL Recognizing Textual Entailment Challenge,” in Proceedings of the Second Text Analysis Conference, TAC 2009, Gaithersburg, Maryland, USA, November 16-17, 2009, NIST, 2009. [Online]. Available: https://tac.nist.gov/publications/2009/additional.papers/RTE5_overview.proceedings.pdf

Show all 94 references
  1. [10]

    The Third PASCAL Recognizing Textual Entailment Challenge,

    D. Giampiccolo, B. Magnini, I. Dagan, and B. Dolan, “The Third PASCAL Recognizing Textual Entailment Challenge,” in Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, S. Sekine, K. Inui, I. Dagan, B. Dolan, D. Giampiccolo, and B. Magnini, Eds., Prag...

  2. [11]

    The Second PASCAL Recognising Textual Entailment Challenge,

    R. Bar-Haim et al., “The Second PASCAL Recognising Textual Entailment Challenge,” 2006. [Online]. Available: https://api.semanticscholar.org/CorpusID:13385138

  3. [12]

    The PASCAL Recognising Textual Entailment Challenge,

    O. and M. B. Dagan Ido and Glickman, “The PASCAL Recognising Textual Entailment Challenge,” in Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, I. and M. B. and d’Alché-B. F. Quiñonero-Candela Joa...

  4. [13]

    The Fourth PASCAL Recognizing Textual Entailment Challenge,

    D. Giampiccolo, H. T. Dang, B. Magnini, I. Dagan, E. Cabrio, and W. B. Dolan, “The Fourth PASCAL Recognizing Textual Entailment Challenge,” in Text Analysis Conference, 2008. [Online]. Available: https://api.semanticscholar.org/CorpusID:12381965

  5. [14]

    Recognizing Textual Entailment: Models and Applications,

    I. Dagan, D. Roth, M. Sammons, and F. M. Zanzotto, “Recognizing Textual Entailment: Models and Applications,” Synthesis Lectures on Human Language Technologies, vol. 6, pp. 1–220, Jul. 2013, doi: 10.2200/S00509ED1V01Y201305HLT023

  6. [15]

    The Winograd Schema Challenge,

    H. J. Levesque, E. Davis, and L. Morgenstern, “The Winograd Schema Challenge,” in AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, 2011. [Online]. Available: https://api.semanticscholar.org/CorpusID:15710851

  7. [16]

    A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,

    N. and B. S. Williams Adina and Nangia, “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1...

  8. [17]

    A large annotated corpus for learning natural language inference,

    S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning, “A large annotated corpus for learning natural language inference,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, L. Màrquez, C. Callison-Burch, and J. Su, Eds., Lisbon, Portugal...

  9. [18]

    GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,

    A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Brussels, Belgium...

  10. [19]

    PaLM: Scaling Language Modeling with Pathways,

    A. Chowdhery et al., “PaLM: Scaling Language Modeling with Pathways,” ArXiv, vol. abs/2204.02311, 2022, [Online]. Available: https://api.semanticscholar.org/CorpusID:247951931

  11. [20]

    Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE,

    Q. Zhong et al., “Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE,” ArXiv, vol. abs/2212.01853, 2022, [Online]. Available: https://api.semanticscholar.org/CorpusID:254246784

  12. [21]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach,

    Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” 2019

  13. [22]

    Semantics-aware BERT for Language Understanding,

    Z. Zhang et al., “Semantics-aware BERT for Language Understanding,” in AAAI Conference on Artificial Intelligence, 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:202539891

  14. [23]

    XLNet: Generalized Autoregressive Pretraining for Language Understanding,

    Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V Le, “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” in Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alché-Buc, E. Fox, and R....

  15. [24]

    AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model,

    S. Soltan et al., “AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model,” ArXiv, vol. abs/2208.01448, 2022, [Online]. Available: https://api.semanticscholar.org/CorpusID:251253416

  16. [25]

    BloombergGPT: A Large Language Model for Finance,

    S. Wu et al., “BloombergGPT: A Large Language Model for Finance,” ArXiv, vol. abs/2303.17564, 2023, [Online]. Available: https://api.semanticscholar.org/CorpusID:257833842

  17. [26]

    First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI,

    S. Banerjee, A. Mahajan, A. Agarwal, and E. Singh, “First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI,” CoRR, vol. abs/2412.09263, 2024, doi: 10.48550/ARXIV.2412.09263

  18. [27]

    ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,

    L. Xue et al., “ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,” Trans Assoc Comput Linguist, vol. 10, pp. 291–306, 2022, doi: 10.1162/tacl_a_00461

  19. [28]

    Rethinking embedding coupling in pre-trained language models,

    H. W. Chung, T. Févry, H. Tsai, M. Johnson, and S. Ruder, “Rethinking embedding coupling in pre-trained language models,” ArXiv, vol. abs/2010.12821, 2020, [Online]. Available: https://api.semanticscholar.org/CorpusID:225067567

  20. [29]

    mGPT: Few-Shot Learners Go Multilingual,

    O. Shliazhko, A. Fenogenova, M. Tikhonova, A. Kozlova, V. Mikhailov, and T. Shavrina, “mGPT: Few-Shot Learners Go Multilingual,” Trans Assoc Comput Linguist, vol. 12, pp. 58–79, 2024, doi: 10.1162/tacl_a_00633

  21. [30]

    DeBERTa: Decoding-enhanced BERT with Disentangled Attention,

    P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with Disentangled Attention,” 2021

  22. [31]

    Available: https://aclanthology.org/2007.tal-1.1

    [Online]. Available: https://aclanthology.org/2007.tal-1.1

  23. [32]

    SpanBERT: Improving Pre-training by Representing and Predicting Spans,

    M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “SpanBERT: Improving Pre-training by Representing and Predicting Spans,” Trans Assoc Comput Linguist, vol. 8, pp. 64–77, 2020, doi: 10.1162/tacl_a_00300

  24. [34]

    DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,

    V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” ArXiv, vol. abs/1910.01108, 2019, [Online]. Available: https://api.semanticscholar.org/CorpusID:203626972

  25. [35]

    A Dataset for Arabic Textual Entailment,

    M. Alabbas, “A Dataset for Arabic Textual Entailment,” in Proceedings of the Student Research Workshop associated with RANLP 2013, I. Temnikova, I. Nikolova, and N. Konstantinova, Eds., Hissar, Bulgaria: INCOMA Ltd. Shoumen, BULGARIA, Sep. 2013, pp. 7–13. [Online]. Available: ...

  26. [36]

    ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,

    K. Al Jallad and N. Ghneim, “ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,” Computer Science, vol. 24, no. 2, Mar. 2023, doi: 10.7494/csci.2023.24.2.4378

  27. [37]

    ArEntail: manually-curated Arabic natural language inference dataset from news headlines,

    R. Obeidat, Y. Al-Harahsheh, M. Al-Ayyoub, and M. Gharaibeh, “ArEntail: manually-curated Arabic natural language inference dataset from news headlines,” Lang Resour Eval, 2024, doi: 10.1007/s10579-024-09731-1

  28. [38]

    Baselines and Test Data for Cross-Lingual Inference,

    Ž. Agić and N. Schluter, “Baselines and Test Data for Cross-Lingual Inference,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), N. Calzolari, K. Choukri, C. Cieri, T. Declerck, S. Goggi, K. Hasida, H. Isahara, B. Maegaa...

  29. [39]

    XNLI: Evaluating Cross-lingual Sentence Representations,

    A. Conneau et al., “XNLI: Evaluating Cross-lingual Sentence Representations,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds., Brussels, Belgium: Association for Computational ...

  30. [40]

    Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,

    W. Yin, J. Hay, and D. Roth, “Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pro...

  31. [41]

    BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,

    M. Lewis et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault...

  32. [42]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach,

    Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” ArXiv, vol. abs/1907.11692, 2019, [Online]. Available: https://api.semanticscholar.org/CorpusID:198953378

  33. [43]

    MiniLM: Deep Self-Attention Distillation for Task- Agnostic Compression of Pre-Trained Transformers,

    W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “MiniLM: Deep Self-Attention Distillation for Task- Agnostic Compression of Pre-Trained Transformers,” ArXiv, vol. abs/2002.10957, 2020, [Online]. Available: https://api.semanticscholar.org/CorpusID:211296536

  34. [44]

    DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient- Disentangled Embedding Sharing,

    P. He, J. Gao, and W. Chen, “DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient- Disentangled Embedding Sharing,” ArXiv, vol. abs/2111.09543, 2021, [Online]. Available: https://api.semanticscholar.org/CorpusID:244346093

  35. [45]

    Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,

    M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers, “Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,” Political Analysis, vol. 32, no. 1, pp. 84–100, Jan. 2024, doi: 10.1017/pan.2023.20

  36. [46]

    GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,

    M. T. I. Khondaker, A. Waheed, E. M. B. Nagoudi, and M. Abdul-Mageed, “GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore...

  37. [47]

    CLSE: Corpus of Linguistically Significant Entities,

    A. Chuklin, J. Zhao, and M. Kale, “CLSE: Corpus of Linguistically Significant Entities,” in Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM), A. Bosselut, K. Chandu, K. Dhole, V. Gangal, S. Gehrmann, Y. Jernite, J. Novikova, and L. ...

  38. [48]

    The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,

    S. Gehrmann et al., “The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,” in Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), A. Bosselut, E. Durmus, V. P. Gangal, S. Gehrmann, Y. Jernite, L. Perez-Belt...

  39. [49]

    GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,

    S. Gehrmann et al., “GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,” in EMNLP 2022 - 2022 Conference on Empirical Methods in Natural Language Processing, W. Che and E. Shutova, Eds., Association for Computational Linguistics (ACL), Dec. 2022, pp. 266–281. doi: ...

  40. [50]

    IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,

    A. Kumar et al., “IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds., Abu Dhabi, United Arab Emirates: Asso...

  41. [51]

    MTG: A Benchmark Suite for Multilingual Text Generation,

    Y. Chen et al., “MTG: A Benchmark Suite for Multilingual Text Generation,” in Findings of the Association for Computational Linguistics: NAACL 2022, M. Carpuat, M.-C. de Marneffe, and I. V. Meza Ruiz, Eds., Seattle, United States: Association for Computational Linguistics, Jul...

  42. [52]

    IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,

    S. Cahyawijaya et al., “IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W. Yih, Eds., Online and Punta C...

  43. [53]

    Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,

    E. M. B. Nagoudi, A. Elmadany, A. El-Shangiti, and M. Abdul-Mageed, “Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,” in Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore: Association for Compu...

  44. [54]

    TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,

    E. M. B. Nagoudi, A. Elmadany, and M. Abdul-Mageed, “TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,” in Proceedinsg of the 5th Workshop on Open-Source Arabic Corpora and Processing Tools with Shared Tasks on Qur`an QA and Fine-Grained Hate Speech Detection, ...

  45. [55]

    AraBench: Benchmarking Dialectal Arabic-English Machine Translation,

    H. Sajjad, A. Abdelali, N. Durrani, and F. Dalvi, “AraBench: Benchmarking Dialectal Arabic-English Machine Translation,” in Proceedings of the 28th International Conference on Computational Linguistics, D. Scott, N. Bel, and C. Zong, Eds., Barcelona, Spain (Online): Internatio...

  46. [56]

    BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,

    A. Bhattacharjee, T. Hasan, W. U. Ahmad, and R. Shahriyar, “BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,” in Findings of the Association for Computational Linguistics: EACL 2023, A. Vlachos and I. Augenstei...

  47. [57]

    AraT5: Text-to-Text Transformers for Arabic Language Generation,

    E. M. B. Nagoudi, A. Elmadany, and M. Abdul-Mageed, “AraT5: Text-to-Text Transformers for Arabic Language Generation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio...

  48. [58]

    CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark,

    Y. Yao et al., “CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark,” 2022. [Online]. Available: https://arxiv.org/abs/2112.13610

  49. [59]

    PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,

    L. Doan, L. T. Nguyen, N. L. Tran, T. Hoang, and D. Q. Nguyen, “PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang...

  50. [60]

    Benchmarking Multidomain English-Indonesian Machine Translation,

    T. W. Guntara, A. F. Aji, and R. E. Prasojo, “Benchmarking Multidomain English-Indonesian Machine Translation,” in Proceedings of the 13th Workshop on Building and Using Comparable Corpora, R. Rapp, P. Zweigenbaum, and S. Sharoff, Eds., Marseille, France: European Language Res...

  51. [61]

    LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,

    J. Guan et al., “LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,” Trans Assoc Comput Linguist, vol. 10, pp. 434–451, 2022, doi: 10.1162/tacl_a_00469

  52. [62]

    XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,

    Y. Liang et al., “XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association f...

  53. [63]

    GLGE: A New General Language Generation Evaluation Benchmark,

    D. Liu et al., “GLGE: A New General Language Generation Evaluation Benchmark,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, C. Zong, F. Xia, W. Li, and R. Navigli, Eds., Online: Association for Computational Linguistics, Aug. 2021, pp. 408–420...

  54. [64]

    XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,

    S. Ruder et al., “XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W. Yih, Eds., Online and Punta Cana, Dominican Republi...

  55. [65]

    SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,

    A. Wang et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA: Curran Associates Inc., 2019

  56. [66]

    XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,

    J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, “XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,” 2020

  57. [67]

    ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,

    M. Abdul-Mageed, A. Elmadany, and E. M. B. Nagoudi, “ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Proces...

  58. [68]

    ORCA: A Challenging Benchmark for Arabic Language Understanding,

    A. Elmadany, E. M. B. Nagoudi, and M. Abdul-Mageed, “ORCA: A Challenging Benchmark for Arabic Language Understanding,” in Annual Meeting of the Association for Computational Linguistics, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:254926875

  59. [69]

    ALUE: Arabic Language Understanding Evaluation,

    H. Seelawi et al., “ALUE: Arabic Language Understanding Evaluation,” in Proceedings of the Sixth Arabic Natural Language Processing Workshop, N. Habash, H. Bouamor, H. Hajj, W. Magdy, W. Zaghouani, F. Bougares, N. Tomeh, I. Abu Farha, and S. Touileb, Eds., Kyiv, Ukraine (Virtu...

  60. [70]

    LAraBench: Benchmarking Arabic AI with Large Language Models,

    A. Abdelali et al., “LAraBench: Benchmarking Arabic AI with Large Language Models,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Y. Graham and M. Purver, Eds., St. Julian’s, Malta: Assoc...

  61. [71]

    Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks,

    J. Weston, A. Bordes, S. Chopra, and T. Mikolov, “Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks,” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2- 4, 2016, Conference Track Proceedings, Y. Bengio an...

  62. [72]

    ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,

    F. Koto et al., “ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Srikumar, Eds., Bangkok, Thailand: Association for Computational Linguistics, Aug. 2...

  63. [73]

    KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,

    J. Ham, Y. J. Choe, K. Park, I. Choi, and H. Soh, “KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association for Computat...

  64. [74]

    CLUE: A Chinese Language Understanding Evaluation Benchmark,

    L. Xu et al., “CLUE: A Chinese Language Understanding Evaluation Benchmark,” in Proceedings of the 28th International Conference on Computational Linguistics, D. Scott, N. Bel, and C. Zong, Eds., Barcelona, Spain (Online): International Committee on Computational Linguistics, ...

  65. [75]

    JGLUE: Japanese General Language Understanding Evaluation,

    K. Kurihara, D. Kawahara, and T. Shibata, “JGLUE: Japanese General Language Understanding Evaluation,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, Marseille, France: European Language Resources Association, Jun. 2022, pp. 2957–2966. [Online]....

  66. [76]

    FlauBERT: Unsupervised Language Model Pre-training for French,

    H. Le et al., “FlauBERT: Unsupervised Language Model Pre-training for French,” in Proceedings of the Twelfth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. M...

  67. [77]

    KLUE: Korean Language Understanding Evaluation,

    S. Park et al., “KLUE: Korean Language Understanding Evaluation,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. [Online]. Available: https://openreview.net/forum?id=q-8h8-LZiUm

  68. [78]

    UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,

    V. Basile, L. Bioglio, A. Bosca, C. Bosco, and V. Patti, “UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), D. Bollegala, R. Hu...

  69. [79]

    SuperGLEBer: German Language Understanding Evaluation Benchmark,

    J. Pfister and A. Hotho, “SuperGLEBer: German Language Understanding Evaluation Benchmark,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gom...

  70. [80]

    IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,

    B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natura...

  71. [81]

    IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,

    D. Kakwani et al., “IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu, Eds., Online: Association for...

  72. [82]

    Using the Framework,

    R. Cooper et al., “Using the Framework,” Mar. 1996. Accessed: Aug. 02, 2024. [Online]. Available: https://files.ifi.uzh.ch/cl/hess/classes/seminare/interface/framework.pdf

  73. [83]

    An extended model of natural logic

    C. D. Manning and B. MacCartney, “An extended model of natural logic”

  74. [84]

    VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,

    P. N.-T. Do, S. Q. Tran, P. G. Hoang, K. Van Nguyen, and N. L.-T. Nguyen, “VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,” in Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gome...

  75. [85]

    Analysis of identifying linguistic phenomena for recognizing inference in text,

    M.-Y. Day and Y.-J. Wang, “Analysis of identifying linguistic phenomena for recognizing inference in text,” Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration, IEEE IRI 2014, pp. 607–612, Mar. 2015, doi: 10.1109/IRI.2014.7051945

  76. [86]

    EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,

    A. Ravichander, A. Naik, C. Rose, and E. Hovy, “EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,” Mar. 2019, pp. 349–361. doi: 10.18653/v1/K19-1033

  77. [87]

    Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,

    S. Sugawara, Y. Kido, H. Yokono, and A. Aizawa, “Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,” Mar. 2017, pp. 806–817. doi: 10.18653/v1/P17-1075

  78. [88]

    NaturalLI: Natural Logic Inference for Common Sense Reasoning,

    G. Angeli and C. D. Manning, “NaturalLI: Natural Logic Inference for Common Sense Reasoning,” in Conference on Empirical Methods in Natural Language Processing, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:2854390

  79. [89]

    Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,

    L. Bentivogli, E. Cabrio, I. Dagan, D. Giampiccolo, M. Leggio, and B. Magnini, “Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,” Mar. 2010

  80. [90]

    Otmakhova, T

    Y. Otmakhova, T. Baldwin, T. Cohn, K. Verspoor, and J. Lau, Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation. 2022. doi: 10.48550/arXiv.2210.03256

  81. [91]

    Neural Networks and Textual Inference: How did we get here and where do we go now?,

    I. Cases and L. Karttunen, “Neural Networks and Textual Inference: How did we get here and where do we go now?,” Sep. 2017

  82. [94]

    On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,

    A. Poliak, Y. Belinkov, J. Glass, and B. Van Durme, “On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: H...

  83. [520]

    Available: https://aclanthology.org/2024.eacl-long.30/

    [Online]. Available: https://aclanthology.org/2024.eacl-long.30/

  84. [1422]

    doi: 10.18653/v1/2023.findings-emnlp.98

  85. [4961]

    doi: 10.18653/v1/2020.findings-emnlp.445

  86. [6018]

    doi: 10.18653/v1/2020.emnlp-main.484

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.