Pith. sign in

Paper Citation Record · LEDGER

Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2111.02840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.02840 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:03.461424Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.523597Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 299597fc-8011-4ba3-900f-bd62627e7964 · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.495253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:866c0d883fe3d94d2a989ae0b94d6e966b6c280bd714c469d8984249bc8a445c

Observation 5eda0ea9-d41e-4ad7-83a9-452544fbe937 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 267

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.774235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:f71100e0b65ba88ee657ed9cbec69d9fa41cd46a6b0bcd59d98a62f162b6e926

Observation e0c194f5-dff8-434a-98bb-03e0a6e15350 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.098865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:34b56403d6e14847d41211e9730bd243a6555fa312a8305a96e412cbd0325c3e

Observation 6f6c6654-9fed-443f-8b78-24f80ffffc80 · inbound

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models cites this paper.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.461424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.461424Z digest=sha256:55e39e6e159fb9053fd25ca059eae64cffdbfb2a6ca5926dc3652523c86eefb3

Observation fc9f0b3b-194d-4e02-932d-47bbb20ff156 · inbound

Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations cites this paper.

Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:15.210109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:15.210109Z digest=sha256:d1473b49233b0550bfe0068e819407ad691f5f363db9caa93b6dd127bfc877a6

Observation 0e1e92a1-6ee0-4d9a-bb6d-3d55a5a0d33b · inbound

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI cites this paper.

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:31.072680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T12:17:59.458633Z digest=sha256:000af59abcf17348816f7367f8c5a6e3036e65be752724c63d3f6b5e9e7eccb8

Observation 1581c38e-115f-4014-b61b-b00e5933eef6 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.496124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.496124Z digest=sha256:2af6a19afbdffe46b6e8c3bd5fc982c92af9123f7f01c9b60c2fb86e449dd716

Observation 770cff7c-c045-40a4-a587-8bee484c7ae1 · inbound

SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds cites this paper.

SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:34.612856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:34.612856Z digest=sha256:b287a1c3fa2b76f41bf30b64ccc790f6c232109b5fcf618df06cfb0c4ef7a85c

Observation 906feaee-1255-47a8-909b-097d40f3347d · inbound

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation cites this paper.

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:07:44.414501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:07:44.414501Z digest=sha256:db4ab43cc315da4acc27d0959b000b518246d6544d3034d150beedd8060b643f

Observation 057b9cff-7b10-40c0-bb0b-d2e1fde96b51 · inbound

Understanding the Prompt Sensitivity cites this paper.

Understanding the Prompt Sensitivity Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:05:23.299039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:41:03.416429Z digest=sha256:50eba077503acdc94b1ab26d98da88e53a404e1ae50231b40ab7e0ef800d4d4f

Observation 12759f72-b577-4243-8e03-70be1f0978ec · inbound

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades cites this paper.

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.367930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T02:31:18.183715Z digest=sha256:bd0342b87c12e49f05e683ee4e01fe2e6abd8f0cb38b013a637f38b8fcba050f

Observation a0a3fae3-2bc8-407c-be7c-f78a3993fab6 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.929065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:3c8d35af1441b16fb2e5163d61d454440f0a4a79ae6efeb8f8485d95f65f3229

Observation d6314ba4-b8cf-4cc9-b9b0-f08fb46c7267 · inbound

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks cites this paper.

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.135367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:27:30.923556Z digest=sha256:ea0e2dd6234aa8213a1d92da176dab209a273623de2fa43016e774a241c12517

Observation 4fbfb9ed-1b90-4717-92e7-8a60808cc76d · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.524920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:c29a9a32bde69641df05186626bdd69fdf400b89858045f9f9b39ce0d8591d88

Observation ab15359e-50ef-4264-af87-7108b2756630 · inbound

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption cites this paper.

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:47:27.205626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T23:42:23.695930Z digest=sha256:097a4eb03308cafbf1cc99e6a4b1ac9e27591ed7e27e04e117285add9511f507