Pith. sign in

Paper Citation Record · LEDGER

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?

As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.20419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20419 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:01.264257Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact20
  • verified fuzzy25
  • unresolved47
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 689c45d2-87cc-4baa-a256-e26be0f53ed6 · outbound

This paper cites Principles of Evaluation in Natural Language Processing,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Principles of Evaluation in Natural Language Processing,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:51.994781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:51.994781Z digest=sha256:5cc6030c97351b60a256221ddd000d6ecb79bbcf00f7bd22380db3fdf2207b7c

Observation 7680c4c5-ab05-40de-bfb9-5b3bdd50f0df · outbound

This paper cites Natural Language Inference ,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Natural Language Inference ,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.365652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.365652Z digest=sha256:499a6ccbe57484285a4bd00298748b6c8f74303ed2c99f94e77ae7b4cd05f820

Observation 1259cf48-029a-46e1-a7d5-840d469b0b86 · outbound

This paper cites PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PROBABILISTIC TEXTUAL ENTAILMENT: GENERIC APPLIED MODELING OF LANGUAGE VARIABILITY,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.487079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.487079Z digest=sha256:44681907f5702408ea2b031a6b2b4ded7b1cf8a4bdb778844d87639ca4d61814

Observation 927f39e9-f3cc-454e-bf7d-d4a33ded723b · outbound

This paper cites Probabilistic textual entailment: Generic applied modeling of language variability,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Probabilistic textual entailment: Generic applied modeling of language variability,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.725915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.725915Z digest=sha256:c94ac11d771388c6697718ccf90d5f3f9a09864e82d1e97c6d50b08070b0742a

Observation f96c17a0-920b-4d4d-8f46-579190916b18 · outbound

This paper cites The Seventh PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Seventh PASCAL Recognizing Textual Entailment Challenge,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.872893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.872893Z digest=sha256:7de8ce6bfab0bbfeafccd1d1f89df1ac2ac4fc1310e0da592d8e99af2f2f4522

Observation 55544bb1-66cf-4be4-b593-6d34d9c0c05d · outbound

This paper cites The Sixth PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Sixth PASCAL Recognizing Textual Entailment Challenge,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.153624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.153624Z digest=sha256:9dbdbd7d961a8387cf0883988900f80e9a1fbac7c21c15997fc10b99c2df4983

Observation 7fe02df0-38d3-44dc-8c54-1e59a0b1d7b2 · outbound

This paper cites The Fifth PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Fifth PASCAL Recognizing Textual Entailment Challenge,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.279800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.279800Z digest=sha256:3bcfe902594513eca329867ca3e200a2b1bd40e815b8e375457eecb3338311f2

Observation 9ece044d-c086-4df3-bf6b-0e44cb3cd691 · outbound

This paper cites The Third PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Third PASCAL Recognizing Textual Entailment Challenge,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.425491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.425491Z digest=sha256:d2a3c9f51f026cdd999f4e0d50841c669aa0519cf4e7a7b42d624ca67d2a2c0e

Observation 92ba844e-0949-47ea-addb-9ab860166dde · outbound

This paper cites The Second PASCAL Recognising Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Second PASCAL Recognising Textual Entailment Challenge,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.601192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.601192Z digest=sha256:27cbcd0a5fde3d8c9601f1ca37206036c79ecd5d821a1089994ba8b22759a2c6

Observation 1b9991c9-4ae9-4ed4-88b2-123b54dcb7e0 · outbound

This paper cites The PASCAL Recognising Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The PASCAL Recognising Textual Entailment Challenge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.769882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.769882Z digest=sha256:a63a22b6e1a9e3dbcdc43a7eb8ad50cfe8f348a3b680d3ddc31067bbfb86a1ee

Observation cae6e79b-fc41-44fd-9fbe-c77460cc6ebf · outbound

This paper cites The Fourth PASCAL Recognizing Textual Entailment Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Fourth PASCAL Recognizing Textual Entailment Challenge,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:53.910758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:53.910758Z digest=sha256:8345d3b0aae8831a0ddf206b7b269e33e32aaa8fe7e2b5090856ded2e8891abb

Observation f78a0d00-f5f3-4b3f-b020-b63fb381a23b · outbound

This paper cites Recognizing Textual Entailment: Models and Applications,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Recognizing Textual Entailment: Models and Applications,

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T13:38:04.882356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:54.002707Z digest=sha256:2ec9f411d753c845c9f1e9927eeccbcab3871382929722509144df13a0b4ff86

Observation ebe93d99-08dc-4bbf-b71a-f2a226fd7535 · outbound

This paper cites The Winograd Schema Challenge,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The Winograd Schema Challenge,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.120712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.120712Z digest=sha256:e605ad128786935c0c65534cbdc9739318d360cff7ef14aaf48d6fef535353d0

Observation 08857552-76f4-423c-8d58-10b41f0d7ba5 · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.187851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.187851Z digest=sha256:33f678f0c9d21465424a12b9c6ca23bf7685eec6f28c74a9ace55c633f739ee0

Observation 66995a61-e64d-4e3c-9ff5-05e584fbf5a5 · outbound

This paper cites A large annotated corpus for learning natural language inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A large annotated corpus for learning natural language inference,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.254334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.254334Z digest=sha256:1e11e6ba212b8f43e9cd0188ca1255a8b2e5923be9bc541479d8d880098259b1

Observation 3597f193-4c39-413e-ad2a-facdc16169b5 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.336819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.336819Z digest=sha256:ca364f9605195513df9e7987a92b874de7fbfbe8fba807103a89b6f0e9809616

Observation f1d0f05a-eadd-4e8d-b8ea-4d6e4d976f47 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PaLM: Scaling Language Modeling with Pathways

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.375277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.375277Z digest=sha256:bad67230f80ea7e10a4fe9e92d9d18959f0312aec74de95ae096644bd1a7f7a7

Observation 77c1d34e-f5b0-49df-a804-4d570799d029 · outbound

This paper cites Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:06.437083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:54.473408Z digest=sha256:702b595148f35fca82f8f179b9142d8dba924b8efdc3fc6ea1a7f0662c77e2b6

Observation 3368c16d-2b2d-4928-be27-0ec4dcc1c3e8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? RoBERTa: A Robustly Optimized BERT Pretraining Approach,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.816504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:54.552526Z digest=sha256:fddddeefa0541291737fa0f7bda9994f4cf8068b7c4b6c9a26cd5b03d543b197

Observation bb834d31-58b1-488c-a015-781ebdc0d131 · outbound

This paper cites Semantics-aware BERT for Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Semantics-aware BERT for Language Understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.803775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:54.618700Z digest=sha256:61e1ab8904388d4759970e845e6ca96f487b26de06186f61955230d3b1881c44

Observation 6d1c5b9b-b0dd-4cf2-8f26-47ae3d0b1789 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XLNet: Generalized Autoregressive Pretraining for Language Understanding,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.791260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:54.677424Z digest=sha256:1d137ac1b87cffa84aeff4b54b45b1f7bc1c05dc4f179e326d899c609b912ee7

Observation 92347395-5567-476b-b319-f3dd66da57aa · outbound

This paper cites AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.767217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.767217Z digest=sha256:8f7229fea030db1604a02487c52abe83beb5b57cac25d794a37a03d2a35fcdf4

Observation c865e660-91f3-44ca-8312-1150c6491cd4 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BloombergGPT: A Large Language Model for Finance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.862401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.862401Z digest=sha256:8279cff7651f291ea3a36b275ed2870dc8a1df54c5a33ccac2e2ffab3f8c6f02

Observation 3ff88a52-1719-405a-80ac-5075153207c8 · outbound

This paper cites First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:04.472952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:54.934359Z digest=sha256:4e3a9ca59c7a0e212cf41a09d39212cab025c30b4cb2d3b14c0da70bbde459b4

Observation a6ec2638-a7ae-49e6-968c-5deaf8a1b8fe · outbound

This paper cites ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:54.981376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:54.981376Z digest=sha256:44089c2534e93ad40d92a5c65338ab1698bd4232598ca946fa06ff84b2e88038

Observation 3fa32c45-f2d4-4d46-9df6-b83f036f226a · outbound

This paper cites Rethinking embedding coupling in pre-trained language models.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Rethinking embedding coupling in pre-trained language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.076184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.076184Z digest=sha256:7d01c0bb577b762c5068fab52a85e6c01aa74b2c547b29d7e32498497d335309

Observation 5b3cc215-56f3-4f27-9ab0-4497e504c3bd · outbound

This paper cites mGPT: Few-Shot Learners Go Multilingual,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? mGPT: Few-Shot Learners Go Multilingual,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.158177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.158177Z digest=sha256:20b34d0e2684a96608e22a9648161af0759317f31d19496ea695f65b44eb3a7c

Observation c363be4e-caae-41c6-acfb-d19496db7f62 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DeBERTa: Decoding-enhanced BERT with Disentangled Attention,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.778137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:55.202142Z digest=sha256:3f59fc503884adbb53baad61b54ec21136d2b0d65c2162222442d2e950965ccf

Observation c9666e37-80af-4a35-9a5a-f5a4ba0d4a8e · outbound

This paper cites Available: https://aclanthology.org/2007.tal-1.1.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Available: https://aclanthology.org/2007.tal-1.1

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:52.167222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:52.167222Z digest=sha256:2ca689831c8b36fb1bace28d844b19a360c56b065fd14d2e074743d3ae369bd2

Observation 5f610143-a1b9-4151-adf0-e3e4041827f1 · outbound

This paper cites SpanBERT: Improving Pre-training by Representing and Predicting Spans,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SpanBERT: Improving Pre-training by Representing and Predicting Spans,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.271734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.271734Z digest=sha256:03c4cab2828032af78ea83c72e6216d9875e83687ee7071dab389c326fa0b73d

Observation 282848ce-90e3-4530-9594-25e0b0ed6caa · outbound

This paper cites SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SqueezeBERT: What can computer vision teach NLP about efficient neural networks?,

Reference 33

Resolution
verified exact
doi, observed 2026-08-06T13:38:04.069737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:55.343223Z digest=sha256:3d89008ecf1a320b3d5e26b5eb49a32893e0b08bd356ac15798acc343a2f54b7

Observation 9a9672f9-0a89-4189-b7b4-577db920ad26 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.387682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.387682Z digest=sha256:57f1028aebd7b4fd40abc76e9221892e0cd71a4a77122ed8938c8407f9101acf

Observation ee6d3f64-5603-453d-9f90-1a7a75c00cf4 · outbound

This paper cites A Dataset for Arabic Textual Entailment,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? A Dataset for Arabic Textual Entailment,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.765397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:55.479023Z digest=sha256:e8561735acfd40b5ae74b9ee984c289fe7a6794e35500b290200451e4405d35e

Observation 3278f0e6-242a-41c4-b4e6-b98a3facfe33 · outbound

This paper cites ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ARNLI: ARABIC NATURAL LANGUAGE INFERENCE ENTAILMENT AND CONTRADICTION DETECTION,

Reference 36

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.913885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:55.613730Z digest=sha256:ebe707e2852bc9c2cb42b4beb132e819f894623770a1977347314561d26b758e

Observation bd9dd812-c1de-45af-9eb6-6662e3eb03f9 · outbound

This paper cites ArEntail: manually-curated Arabic natural language inference dataset from news headlines,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ArEntail: manually-curated Arabic natural language inference dataset from news headlines,

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.723052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:55.710306Z digest=sha256:1097e6daad2b753d76e5e59eb58c63357254b05130b8c42e20e727a4a062cb2c

Observation 1a673ff0-d2b4-49a6-98f6-a8169e2a117e · outbound

This paper cites Baselines and Test Data for Cross-Lingual Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Baselines and Test Data for Cross-Lingual Inference,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.752969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:55.827796Z digest=sha256:ee82b4c2ed3d8651ed3df6fb8d8e0f7ef040c6132cd967ae9b4bca0a87805a51

Observation 72509850-399e-4755-a1a4-301479c834bc · outbound

This paper cites XNLI: Evaluating Cross-lingual Sentence Representations,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XNLI: Evaluating Cross-lingual Sentence Representations,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.900030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.900030Z digest=sha256:e57c66121af234cc6e32877c0c31ceb98fb15b7d7e0602e69ffbe8d29bc76cc8

Observation 5eef9a04-89dc-4f4f-915a-623b072993da · outbound

This paper cites Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:55.976655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:55.976655Z digest=sha256:be4ec8a72f36e0ffefa03515d9593548b0273f7e670c4f548a6c8f0c6145014f

Observation fed28c6d-80b1-4c58-a42d-1a3e195df645 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.072296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.072296Z digest=sha256:d5ff409b3be2983bb70dd35a2c59441608e26363925ea2fcba48db5e0e3d8887

Observation 134ae3ce-000c-4867-b024-63f1140fcd0f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.192293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.192293Z digest=sha256:0bfb578b45a9b4b3da7d0ff3f3852ca5d7294ce5824e1f20c9a23dceff60f29f

Observation 4a4fdf83-36d2-47a1-9a3a-bf5c34d8e745 · outbound

This paper cites MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.242266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.242266Z digest=sha256:25d1e6cc2fe663319e910518bc33b4df24ec4eea81dea4e64c7d6c089408e419

Observation 6e8521c3-0535-431a-adb0-cf7b87e61d38 · outbound

This paper cites DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.318397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.318397Z digest=sha256:796205ec3e12c13ce2a3f9155c80d04caded55e54c8f648b77d0d66a5087f307

Observation 16c9bc1a-6cad-42d4-b394-b7fbd6ae360d · outbound

This paper cites Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.411095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.411095Z digest=sha256:b39722bca1184610bc061ea1105b51f60bf3741b43a0a763bf8ff5ce2cea8f19

Observation 667d3618-e2f9-4aad-aed6-1437fbae2362 · outbound

This paper cites GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.506650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.506650Z digest=sha256:e533ad0af71da51df850054ff72e7fa7e9652e2610a05742dd7063b2fc2f0d7a

Observation 40349f99-cd76-426c-8308-65af190c6935 · outbound

This paper cites CLSE: Corpus of Linguistically Significant Entities,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CLSE: Corpus of Linguistically Significant Entities,

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.524186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:56.576929Z digest=sha256:e19d559e3eafffc4ac37f715c6986769b580022e4fafedff2f2c12b19746cff3

Observation 6780fe5d-533e-4a91-b726-fad1dc3b9773 · outbound

This paper cites The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.721456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.721456Z digest=sha256:5b6f2e7a1cbd773040156b13f53b7b3ddc9b43dfdbeb59535a293c4240a6b32e

Observation d9bdb60a-d4bf-4124-a838-90cf873f7543 · outbound

This paper cites GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GEMv2: Multilingual NLG Benchmarking in a Single Line of Code,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:56.833242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:56.833242Z digest=sha256:a73ae39d2826ee5c930d709e13d15de5f9e10053c09d8ff2ae266cb5f9cb1a42

Observation 8d70fd07-bf64-4b58-9148-ee1014d64153 · outbound

This paper cites IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages,

Reference 50

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.343289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:56.920313Z digest=sha256:30b5d9cd46c0a6f7fba593647b93e1b67b3783751c2e6fe3e8639ef1785ab44c

Observation cefb55ee-4493-4e6d-a886-cfd1123f63fa · outbound

This paper cites MTG: A Benchmark Suite for Multilingual Text Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? MTG: A Benchmark Suite for Multilingual Text Generation,

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:37:57.001448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.001448Z digest=sha256:9866225be5111f88b33e2edecfbcfdeb990a8036ad12d949df087d70ccad548e

Observation d9de5286-75d9-4675-8b07-727b8602c4a2 · outbound

This paper cites IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.121190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.121190Z digest=sha256:41f80e8dcb503d7ecbad850b0b0265e89cecc1c198f190420b30b145ee3ee56c

Observation 5bcbe44a-2669-4ef9-a89d-6eb813094b39 · outbound

This paper cites Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Dolphin: A Challenging and Diverse Benchmark for Arabic NLG,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.740001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:57.234294Z digest=sha256:700024d4836de014b9c72886fea2ad77a3d43b8a96cad6b951fb28f807133891

Observation 8720e990-6952-456b-85f3-0f4c21cbe7cb · outbound

This paper cites TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.727123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:57.449397Z digest=sha256:16c9d5cf713f7aeb900bde06b09a7143f81506a65ac505c4c774901be38b599e

Observation ddab1a97-4f81-4959-a22c-a35ea5dffcae · outbound

This paper cites AraBench: Benchmarking Dialectal Arabic-English Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AraBench: Benchmarking Dialectal Arabic-English Machine Translation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.376864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.376864Z digest=sha256:0a463af7a2694b8a90d64c31c026bf9a124e619436d1c7cea3aa06798509b8af

Observation 71743e22-ff20-481c-9453-e97fd137ca06 · outbound

This paper cites BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.717646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.717646Z digest=sha256:ea3fbf8f7082826d26072af31b831045d57f82cfcf987f27b17a079cbd6053e6

Observation e9f11951-f937-49e1-9aee-a16f37f9f3d7 · outbound

This paper cites AraT5: Text-to-Text Transformers for Arabic Language Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? AraT5: Text-to-Text Transformers for Arabic Language Generation,

Reference 57

Resolution
verified exact
doi, observed 2026-08-06T13:38:03.156193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:57.586733Z digest=sha256:186f451b7cdbce7c011855aa1936e378833226850990ce97fb0a905fe9ead64b

Observation dfb6a474-923b-401a-b713-bd14cc433a19 · outbound

This paper cites CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:05.964795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:57.924811Z digest=sha256:f15a84f75b86a4d3202eefc83197308a931f82fac76202b75fb14b7a56199b30

Observation 697a2908-5568-4ec5-aa20-637db9fab701 · outbound

This paper cites PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation,

Reference 59

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.957624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:57.828392Z digest=sha256:7c31efaf2607a97ab006af75035f52da54a1d0c3cb297282656d9309b58737e7

Observation 6026dce6-d303-4a06-95d5-120a540ddcfb · outbound

This paper cites Benchmarking Multidomain English-Indonesian Machine Translation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Benchmarking Multidomain English-Indonesian Machine Translation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.714116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.142558Z digest=sha256:65d2263474e1c154c154ce55a4b9ef40fd15ca63f5fe5cc9ed4d3b9702a627c9

Observation 733fc5c9-80fc-4805-99d5-8ee6d3a4d02c · outbound

This paper cites LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.998337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.998337Z digest=sha256:ed126a3eb07c5d64810333cf8aa9d29bac2a13f35678094207aeb27ebe263e9a

Observation ec315073-34a0-4046-908d-7d5c555da901 · outbound

This paper cites XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.701733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.317146Z digest=sha256:1346fd0bfd968f82b4ef822b4bc3172b2f3a529ba522567bfe2cccb48db37e6c

Observation 834cf8cc-1dd7-40e3-a06a-18e24224514f · outbound

This paper cites GLGE: A New General Language Generation Evaluation Benchmark,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? GLGE: A New General Language Generation Evaluation Benchmark,

Reference 63

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.737908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.240465Z digest=sha256:be56a3fa7999481d8af4767e3c6d712d9f49077b08ddf5eec37dfed479e7194a

Observation cd68d2a5-e9de-4fd4-84b3-337170cf595c · outbound

This paper cites XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:58.638320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:58.638320Z digest=sha256:e662fabd1f6e1a3169b71f3edf82af1215486b026aaf9d63e3ba0d20d87e4230

Observation 5d447d3c-7d5e-4330-9a75-7b9517c6d21f · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.676377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.735056Z digest=sha256:2f134948b1f30d559c96afb8ea8856c3279418f1cde93d4aa4e2fea5424b4502

Observation b22bb799-41d6-4867-8b87-9c6c7755d77d · outbound

This paper cites XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.689029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.530760Z digest=sha256:271ff627fd5fd8025f9a9cc6c7513d2c9da6d65f605e59d3ee7507f459135efc

Observation 91ba2886-2bae-4cb1-b74f-6be8ae3ec06f · outbound

This paper cites ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:58.891569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:58.891569Z digest=sha256:bf5cc7a77536830932fa97a6217b0ad53e27c7317b3bc1ba6327934f802382b7

Observation 6c24f0c1-3e5d-44a9-bc2a-ea910c5188b6 · outbound

This paper cites ORCA: A Challenging Benchmark for Arabic Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ORCA: A Challenging Benchmark for Arabic Language Understanding,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.651628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.994414Z digest=sha256:4ef39e58e09ef4226df00c4c450f1dc161224a890a1e588653293fc52af41c4f

Observation b32b011a-7ce8-4541-b6bd-ae0eff8c7655 · outbound

This paper cites ALUE: Arabic Language Understanding Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ALUE: Arabic Language Understanding Evaluation,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.663782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:58.815279Z digest=sha256:1de448e60c83cdd9972ef7305d2cad0a913f6aaea86c55f32a850364a2d0f525

Observation aa3195d2-a51d-4dc0-b2af-9a6328055803 · outbound

This paper cites LAraBench: Benchmarking Arabic AI with Large Language Models,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? LAraBench: Benchmarking Arabic AI with Large Language Models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.638619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:59.258663Z digest=sha256:3765495123fec6e38802d7a2eb72a1c9d62e7215a75a2a6047ee675d8d825b66

Observation d53b659a-8b8c-4b69-b0d3-297b7cee87ca · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:59.448235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.448235Z digest=sha256:18822729e8fbfa5fb10b6ca1e26f0446ebb48dd5facfb1ffc060b708faf4b22f

Observation a748ce81-d50b-4170-9c70-a9b4a5157d34 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic,

Reference 72

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:37:59.160290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.160290Z digest=sha256:afd217814d32d45cafed67bc6c710aab3bb12e0477a928d3fd5b777fc3df7f08

Observation 9877f715-1110-4b79-8c3f-227868f75d03 · outbound

This paper cites KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding,

Reference 73

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.509856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:59.675564Z digest=sha256:44452064a920874cff07c9f762d3d12961617f91d0de142b1d9270fe76b9d0bb

Observation 9a5d9f2e-eccc-4638-a16a-58b6e0d13866 · outbound

This paper cites CLUE: A Chinese Language Understanding Evaluation Benchmark,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? CLUE: A Chinese Language Understanding Evaluation Benchmark,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:59.782498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.782498Z digest=sha256:770a1bdedaf9cf9887fbadfa0732a5ac9ab0284c2fcd16b91a19ebdd94b0e0e2

Observation 3c0eb517-fdff-4864-8015-0a2649c2db02 · outbound

This paper cites JGLUE: Japanese General Language Understanding Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? JGLUE: Japanese General Language Understanding Evaluation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.597294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:59.854991Z digest=sha256:3c915be7b37167f16fd0242bf66e1dee951daafd88618c4a27a0e84ff3adc6dc

Observation 1a814e55-d8bb-433c-aca1-1e236472a633 · outbound

This paper cites FlauBERT: Unsupervised Language Model Pre-training for French,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? FlauBERT: Unsupervised Language Model Pre-training for French,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.611359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:59.569273Z digest=sha256:23df0c7cf11b7152dc83736e27f99cab8d3d6a5978fab4ebc08c4516e1cc764d

Observation 41de1071-0512-47c3-a8fe-401ca88f2706 · outbound

This paper cites KLUE: Korean Language Understanding Evaluation,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? KLUE: Korean Language Understanding Evaluation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.583677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.038452Z digest=sha256:534e210464c2026664c0cf6468eff23cdf499631524d64cf34ef46a8aaf192e6

Observation 1f389aa9-9204-45a9-b2f4-5567479ee156 · outbound

This paper cites UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? UINAUIL: A Unified Benchmark for Italian Natural Language Understanding,

Reference 78

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.321727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.166860Z digest=sha256:c8735e5402a7b1aa3de6c3e95549dcb386d4b6a3f646c09956d7d94443f21a8e

Observation f7a1f880-dcd6-4806-b4b3-97735c014b95 · outbound

This paper cites SuperGLEBer: German Language Understanding Evaluation Benchmark,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? SuperGLEBer: German Language Understanding Evaluation Benchmark,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:00.244011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:00.244011Z digest=sha256:73d950b338fd3db04b710622c0939f6df8083509b952bb2df2d599d91d5f5d24

Observation 8d1de291-c078-4bfd-a723-c1c8c9e3acb6 · outbound

This paper cites IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:59.916339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:59.916339Z digest=sha256:f72abe209b69ea36d2a1cf5c472493f7e0ef0f8fe08524d809acf8ebd637b9bd

Observation 76900af7-102a-4306-8a0b-67ccdec9f9c5 · outbound

This paper cites IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.569610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.370586Z digest=sha256:7412d423643ba582a9e49b022e76e0d14210bf2e927b87e1fbdab4c182e4e35a

Observation 1e878b29-8a63-4574-a305-5822fc162d67 · outbound

This paper cites Using the Framework,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Using the Framework,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.555603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.526646Z digest=sha256:c8736715c3b0fb4979a032a6693fd691cd4afea2a41c7437706497f62b85e737

Observation ba872268-3e9f-40d9-9560-5bbaf7689c15 · outbound

This paper cites An extended model of natural logic.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? An extended model of natural logic

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.542099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.643216Z digest=sha256:677021dd54a4af6c8d9956c14d10d8cf0b06d60340c56323b4c7e09190e7d3af

Observation 200d0496-5d0a-4469-b32f-05bc9bee392d · outbound

This paper cites VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding,

Reference 84

Resolution
verified exact
doi, observed 2026-08-06T13:38:02.179538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.314390Z digest=sha256:258d1f1d5f371470af802b756b1896657caa7a083884e317ccb1476b54e4c502

Observation b17cb167-b484-401e-a616-35348e3c564f · outbound

This paper cites Analysis of identifying linguistic phenomena for recognizing inference in text,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Analysis of identifying linguistic phenomena for recognizing inference in text,

Reference 85

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-06T13:38:05.500337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.819109Z digest=sha256:bacf1624fb11a16ee7be36061d54bf61b2034e843dcac385ebc2a8a65b157fff

Observation 2351178e-b1e2-4ac3-aae2-2b68cae1df7e · outbound

This paper cites EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference,

Reference 86

Resolution
verified exact
doi, observed 2026-08-06T13:38:01.970328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.875892Z digest=sha256:00e0da0fb8f2f20883cb236fba1a710b795e4ce11207cb6a4c2959909de7374e

Observation 40b6093e-fe1b-4e39-b590-8aa0b747a202 · outbound

This paper cites Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and Readability,

Reference 87

Resolution
verified exact
doi, observed 2026-08-06T13:38:01.777765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.932270Z digest=sha256:38213f40dd073ff5b53d962f0065c92d5816a125a65b30bfdd06c832ea5ed128

Observation acffebb4-f9ff-4e31-b663-cfa23187988a · outbound

This paper cites NaturalLI: Natural Logic Inference for Common Sense Reasoning,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? NaturalLI: Natural Logic Inference for Common Sense Reasoning,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.125392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.991401Z digest=sha256:49c9fe42601cf20d68afaaa005a0ee977187d3a72a6f15eb200a95cb1c3981dd

Observation 5dec987b-ab00-4b50-ba09-6af7a885574b · outbound

This paper cites Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference.,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.450314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:00.733334Z digest=sha256:872a862a23e670889aef820145af48ff46c00c82c21a7d1567f36065f28a5084

Observation 1106c391-1b24-4171-b2f6-dfb5fc5e9610 · outbound

This paper cites Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:01.441104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:01.153777Z digest=sha256:573de48f6514ae7b9fa97f0c0c6fcc53ffec14caedd6e143de68955eacdff317

Observation c4943d82-5393-44fc-b3cd-d9f8ea6645d3 · outbound

This paper cites Neural Networks and Textual Inference: How did we get here and where do we go now?,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Neural Networks and Textual Inference: How did we get here and where do we go now?,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:06.760253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:01.264257Z digest=sha256:2f62ea6813c00b9bb8c87c1c25ea76b77580e92e04e3df79f9b41b1955601676

Observation 8fdf345c-55e3-41bd-8aec-bfd2f400b9e4 · outbound

This paper cites On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference,

Reference 94

Resolution
verified exact
doi, observed 2026-08-06T13:38:01.605861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:01.049741Z digest=sha256:429fef4a0f62d5d8e3e0bff261a9c8d2e371c04daf17227f03a6deeb17e92272

Observation 40049495-c160-45a7-b31a-d5732925ded0 · outbound

This paper cites Available: https://aclanthology.org/2024.eacl-long.30/.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Available: https://aclanthology.org/2024.eacl-long.30/

Reference 520

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:07.625044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:37:59.366997Z digest=sha256:4f8f69bf87a0a24d2a69ba781313a67d64388b35b3ffbcdff26c7de03cfffe46

Observation 066998a0-309a-4784-bc5e-efbdddc4b793 · outbound

This paper cites an unresolved cited work.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work

Reference 1422

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:57.304442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:57.304442Z digest=sha256:e589ea1512cfda887b5b6ecce762c98b547cb9f683f66e2518835f8451067b48

Observation f7cab572-fdec-491a-aa9c-c7017d1b647a · outbound

This paper cites an unresolved cited work.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work

Reference 4961

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:00.432949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:00.432949Z digest=sha256:9acaf4dd556533f83d675bd7687b4790d6ee54a05f08e955dc14085c4f5af2c8

Observation b815d477-6ea3-4d30-bf6b-0f182e0c0899 · outbound

This paper cites an unresolved cited work.

Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? Unresolved cited work

Reference 6018

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:58.401646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:58.401646Z digest=sha256:956943411da630f655df11ff5038231e7602422f7e4ae95adf113ad60da31bab

Pith citing papers

No inbound Pith citation observations are available.