Pith. sign in

Paper Citation Record · LEDGER

Resurrecting saturated LLM benchmarks with adversarial encoding

As of 22 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 0 inbound Pith citation observations for arXiv:2502.06738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06738 v1

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:33:17.337479Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

7 of 7 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 501ad800-ebb2-4b3b-9227-a0ed7aec9f4e · outbound

This paper cites Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation.

Resurrecting saturated LLM benchmarks with adversarial encoding Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:33:17.200609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.200609Z digest=sha256:7ea30b506e9818ca8a921da800731dccefe3fe9a56bd73ea17611b5626f82777

Observation bda0e274-b3b8-4f1b-8a82-d6efe9ec049d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Resurrecting saturated LLM benchmarks with adversarial encoding Measuring Massive Multitask Language Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:33:17.232253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.232253Z digest=sha256:ac440b081b14f31d041cfd086803bb2a22b8d0c2673e7182b77a9a132af5cf58

Observation e618dcf9-7b2c-46e6-ba19-ceb623dd2756 · outbound

This paper cites an unresolved cited work.

Resurrecting saturated LLM benchmarks with adversarial encoding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:33:17.482433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:33:17.263517Z digest=sha256:1efff4693e6a98a00e9e0c2f53637e1d29b82d3ca471a34612bd553c0127a32f

Observation b637f7dc-40ba-476a-aae5-ac67b0dab463 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Resurrecting saturated LLM benchmarks with adversarial encoding GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:33:17.281392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.281392Z digest=sha256:18fc53f2f6211539489e5bdd8bb7955f4c7162891f2b31e75a4c32df0b387624

Observation c5045eaf-401a-42b3-9f11-3c5aa05e49cd · outbound

This paper cites M., Weber, L., Choshen, L., Sun, Y., Xu, G., & Yurochkin, M.

Resurrecting saturated LLM benchmarks with adversarial encoding M., Weber, L., Choshen, L., Sun, Y., Xu, G., & Yurochkin, M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:33:17.468072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:33:17.303773Z digest=sha256:9d9a8e8dcba62587e4d523e1b16534eca5c2c1721d7776cbc118f96e87a463b6

Observation 833f2b1b-0d99-4dff-b062-285a0902cf16 · outbound

This paper cites Evaluating LLMs with Multiple Problems at once.

Resurrecting saturated LLM benchmarks with adversarial encoding Evaluating LLMs with Multiple Problems at once

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:33:17.424768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:33:17.319689Z digest=sha256:2a0077c7cfc5f06ae2a7db24db62579451be2718fe1f8fa8aa3b4875aa2c1b51

Observation a1caf82e-42e9-4443-a364-aec25bb9945c · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

Resurrecting saturated LLM benchmarks with adversarial encoding Large Language Models Are Not Robust Multiple Choice Selectors

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-08T14:33:17.337479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.337479Z digest=sha256:8790f544f02a32de7e1e09e5aa895c334ef1428cd40ed742c21bd0cd5d39a053

Pith citing papers

No inbound Pith citation observations are available.