Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:01.264257Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.20419.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:01.264257Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
94 of 94 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 689c45d2-87cc-4baa-a256-e26be0f53ed6 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Principles of Evaluation in Natural Language Processing,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7680c4c5-ab05-40de-bfb9-5b3bdd50f0df · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Natural Language Inference ,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1259cf48-029a-46e1-a7d5-840d469b0b86 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927f39e9-f3cc-454e-bf7d-d4a33ded723b · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Probabilistic textual entailment: Generic applied modeling of language variability,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f96c17a0-920b-4d4d-8f46-579190916b18 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Seventh PASCAL Recognizing Textual Entailment Challenge,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55544bb1-66cf-4be4-b593-6d34d9c0c05d · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Sixth PASCAL Recognizing Textual Entailment Challenge,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe02df0-38d3-44dc-8c54-1e59a0b1d7b2 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Fifth PASCAL Recognizing Textual Entailment Challenge,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ece044d-c086-4df3-bf6b-0e44cb3cd691 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Third PASCAL Recognizing Textual Entailment Challenge,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ba844e-0949-47ea-addb-9ab860166dde · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Second PASCAL Recognising Textual Entailment Challenge,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b9991c9-4ae9-4ed4-88b2-123b54dcb7e0 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The PASCAL Recognising Textual Entailment Challenge,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae6e79b-fc41-44fd-9fbe-c77460cc6ebf · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Fourth PASCAL Recognizing Textual Entailment Challenge,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78a0d00-f5f3-4b3f-b020-b63fb381a23b · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Recognizing Textual Entailment: Models and Applications,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebe93d99-08dc-4bbf-b71a-f2a226fd7535 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Winograd Schema Challenge,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08857552-76f4-423c-8d58-10b41f0d7ba5 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66995a61-e64d-4e3c-9ff5-05e584fbf5a5 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A large annotated corpus for learning natural language inference,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3597f193-4c39-413e-ad2a-facdc16169b5 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d0f05a-eadd-4e8d-b8ea-4d6e4d976f47 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PaLM: Scaling Language Modeling with Pathways
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c1d34e-f5b0-49df-a804-4d570799d029 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3368c16d-2b2d-4928-be27-0ec4dcc1c3e8 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? RoBERTa: A Robustly Optimized BERT Pretraining Approach,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb834d31-58b1-488c-a015-781ebdc0d131 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Semantics-aware BERT for Language Understanding,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d1c5b9b-b0dd-4cf2-8f26-47ae3d0b1789 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XLNet: Generalized Autoregressive Pretraining for Language Understanding,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92347395-5567-476b-b319-f3dd66da57aa · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c865e660-91f3-44ca-8312-1150c6491cd4 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BloombergGPT: A Large Language Model for Finance
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff88a52-1719-405a-80ac-5075153207c8 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6ec2638-a7ae-49e6-968c-5deaf8a1b8fe · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa32c45-f2d4-4d46-9df6-b83f036f226a · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Rethinking embedding coupling in pre-trained language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3cc215-56f3-4f27-9ab0-4497e504c3bd · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? mGPT: Few-Shot Learners Go Multilingual,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c363be4e-caae-41c6-acfb-d19496db7f62 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DeBERTa: Decoding-enhanced BERT with Disentangled Attention,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9666e37-80af-4a35-9a5a-f5a4ba0d4a8e · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Available: https://aclanthology.org/2007.tal-1.1
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f610143-a1b9-4151-adf0-e3e4041827f1 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SpanBERT: Improving Pre-training by Representing and Predicting Spans,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 282848ce-90e3-4530-9594-25e0b0ed6caa · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a9672f9-0a89-4189-b7b4-577db920ad26 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6d3f64-5603-453d-9f90-1a7a75c00cf4 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A Dataset for Arabic Textual Entailment,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3278f0e6-242a-41c4-b4e6-b98a3facfe33 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd9dd812-c1de-45af-9eb6-6662e3eb03f9 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ArEntail: manually-curated Arabic natural language inference dataset from news headlines,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a673ff0-d2b4-49a6-98f6-a8169e2a117e · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Baselines and Test Data for Cross-Lingual Inference,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72509850-399e-4755-a1a4-301479c834bc · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XNLI: Evaluating Cross-lingual Sentence Representations,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eef9a04-89dc-4f4f-915a-623b072993da · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed28c6d-80b1-4c58-a42d-1a3e195df645 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 134ae3ce-000c-4867-b024-63f1140fcd0f · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4fdf83-36d2-47a1-9a3a-bf5c34d8e745 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8521c3-0535-431a-adb0-cf7b87e61d38 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c9bc1a-6cad-42d4-b394-b7fbd6ae360d · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667d3618-e2f9-4aad-aed6-1437fbae2362 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40349f99-cd76-426c-8308-65af190c6935 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CLSE: Corpus of Linguistically Significant Entities,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6780fe5d-533e-4a91-b726-fad1dc3b9773 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9bdb60a-d4bf-4124-a838-90cf873f7543 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d70fd07-bf64-4b58-9148-ee1014d64153 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cefb55ee-4493-4e6d-a886-cfd1123f63fa · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? MTG: A Benchmark Suite for Multilingual Text Generation,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9de5286-75d9-4675-8b07-727b8602c4a2 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bcbe44a-2669-4ef9-a89d-6eb813094b39 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8720e990-6952-456b-85f3-0f4c21cbe7cb · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddab1a97-4f81-4959-a22c-a35ea5dffcae · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AraBench: Benchmarking Dialectal Arabic-English Machine Translation,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71743e22-ff20-481c-9453-e97fd137ca06 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f11951-f937-49e1-9aee-a16f37f9f3d7 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AraT5: Text-to-Text Transformers for Arabic Language Generation,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfb6a474-923b-401a-b713-bd14cc433a19 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 697a2908-5568-4ec5-aa20-637db9fab701 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6026dce6-d303-4a06-95d5-120a540ddcfb · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Benchmarking Multidomain English-Indonesian Machine Translation,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 733fc5c9-80fc-4805-99d5-8ee6d3a4d02c · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec315073-34a0-4046-908d-7d5c555da901 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 834cf8cc-1dd7-40e3-a06a-18e24224514f · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GLGE: A New General Language Generation Evaluation Benchmark,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd68d2a5-e9de-4fd4-84b3-337170cf595c · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d447d3c-7d5e-4330-9a75-7b9517c6d21f · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b22bb799-41d6-4867-8b87-9c6c7755d77d · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91ba2886-2bae-4cb1-b74f-6be8ae3ec06f · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c24f0c1-3e5d-44a9-bc2a-ea910c5188b6 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ORCA: A Challenging Benchmark for Arabic Language Understanding,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b32b011a-7ce8-4541-b6bd-ae0eff8c7655 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ALUE: Arabic Language Understanding Evaluation,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa3195d2-a51d-4dc0-b2af-9a6328055803 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? LAraBench: Benchmarking Arabic AI with Large Language Models,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d53b659a-8b8c-4b69-b0d3-297b7cee87ca · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a748ce81-d50b-4170-9c70-a9b4a5157d34 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9877f715-1110-4b79-8c3f-227868f75d03 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a5d9f2e-eccc-4638-a16a-58b6e0d13866 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CLUE: A Chinese Language Understanding Evaluation Benchmark,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c0eb517-fdff-4864-8015-0a2649c2db02 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? JGLUE: Japanese General Language Understanding Evaluation,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a814e55-d8bb-433c-aca1-1e236472a633 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? FlauBERT: Unsupervised Language Model Pre-training for French,
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41de1071-0512-47c3-a8fe-401ca88f2706 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? KLUE: Korean Language Understanding Evaluation,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f389aa9-9204-45a9-b2f4-5567479ee156 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7a1f880-dcd6-4806-b4b3-97735c014b95 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SuperGLEBer: German Language Understanding Evaluation Benchmark,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d1de291-c078-4bfd-a723-c1c8c9e3acb6 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76900af7-102a-4306-8a0b-67ccdec9f9c5 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e878b29-8a63-4574-a305-5822fc162d67 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Using the Framework,
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba872268-3e9f-40d9-9560-5bbaf7689c15 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? An extended model of natural logic
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 200d0496-5d0a-4469-b32f-05bc9bee392d · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b17cb167-b484-401e-a616-35348e3c564f · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Analysis of identifying linguistic phenomena for recognizing inference in text,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2351178e-b1e2-4ac3-aae2-2b68cae1df7e · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40b6093e-fe1b-4e39-b590-8aa0b747a202 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acffebb4-f9ff-4e31-b663-cfa23187988a · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? NaturalLI: Natural Logic Inference for Common Sense Reasoning,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dec987b-ab00-4b50-ba09-6af7a885574b · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1106c391-1b24-4171-b2f6-dfb5fc5e9610 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4943d82-5393-44fc-b3cd-d9f8ea6645d3 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Neural Networks and Textual Inference: How did we get here and where do we go now?,
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fdf345c-55e3-41bd-8aec-bfd2f400b9e4 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40049495-c160-45a7-b31a-d5732925ded0 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Available: https://aclanthology.org/2024.eacl-long.30/
Reference 520
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 066998a0-309a-4784-bc5e-efbdddc4b793 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work
Reference 1422
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cab572-fdec-491a-aa9c-c7017d1b647a · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work
Reference 4961
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b815d477-6ea3-4d30-bf6b-0f182e0c0899 · outbound
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work
Reference 6018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.