Pith. sign in

Paper Citation Record · LEDGER

Resurrecting saturated LLM benchmarks with adversarial encoding

As of 10 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 0 inbound Pith citation observations for arXiv:2502.06738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06738 v1

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:33:17.337479Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

7 of 7 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 501ad800-ebb2-4b3b-9227-a0ed7aec9f4e · outbound

This paper cites Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation.

Resurrecting saturated LLM benchmarks with adversarial encoding Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:33:17.200609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.200609Z digest=sha256:c3a67f26dee17d35a89a7d3872d5ba0d49ff43e5b6645e6fc65c4bbef49c42c5

Observation bda0e274-b3b8-4f1b-8a82-d6efe9ec049d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Resurrecting saturated LLM benchmarks with adversarial encoding Measuring Massive Multitask Language Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:33:17.232253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.232253Z digest=sha256:607c2ac9475ace40ec1164f2ca37e9907a6338c4ea70d064ff6d146aa2fe8012

Observation e618dcf9-7b2c-46e6-ba19-ceb623dd2756 · outbound

This paper cites an unresolved cited work.

Resurrecting saturated LLM benchmarks with adversarial encoding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:33:17.482433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:33:17.263517Z digest=sha256:28e34a5b69eb952444af3f9553ef7ebf52c358632eac627553f1984b086466ae

Observation b637f7dc-40ba-476a-aae5-ac67b0dab463 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Resurrecting saturated LLM benchmarks with adversarial encoding GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:33:17.281392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.281392Z digest=sha256:cad8f14211ddfde12fa666739eeb242b57097b6545082be788cbdb6d35376bc8

Observation c5045eaf-401a-42b3-9f11-3c5aa05e49cd · outbound

This paper cites M., Weber, L., Choshen, L., Sun, Y., Xu, G., & Yurochkin, M.

Resurrecting saturated LLM benchmarks with adversarial encoding M., Weber, L., Choshen, L., Sun, Y., Xu, G., & Yurochkin, M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:33:17.468072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:33:17.303773Z digest=sha256:fc66ffed85f0add1f97ded634a61ad02b708f3ddad756bb905485b5e48f27292

Observation 833f2b1b-0d99-4dff-b062-285a0902cf16 · outbound

This paper cites Evaluating LLMs with Multiple Problems at once.

Resurrecting saturated LLM benchmarks with adversarial encoding Evaluating LLMs with Multiple Problems at once

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:33:17.424768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:33:17.319689Z digest=sha256:3f21a7cbd9329db6b86309c6e5312cd31024dec1b1e305fd01896b820174cf3b

Observation a1caf82e-42e9-4443-a364-aec25bb9945c · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

Resurrecting saturated LLM benchmarks with adversarial encoding Large Language Models Are Not Robust Multiple Choice Selectors

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-08T14:33:17.337479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:33:17.337479Z digest=sha256:d1fe1adb9135ca59f4709130358e5372f7801283f0caf82dded85767909392f2

Pith citing papers

No inbound Pith citation observations are available.