Pith. sign in

Paper Citation Record · LEDGER

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2404.18923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.18923 v5

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T02:06:53.585629Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact7
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 32ebd788-32dd-46d3-aed3-b31474d611b4 · outbound

This paper cites Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.107794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:4ff73dc59699278058a7ff4c90a2392d81220b221957d619b37790896c0ab17b

Observation 0d95b0e6-e4c0-4889-87bf-176f003f54ee · outbound

This paper cites Lossless and Near-Lossless Compression for Foundation Models.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Lossless and Near-Lossless Compression for Foundation Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.114043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:00bd24dc1cab3f3a5a6f556d249bc48e324aa834a19b1ec2190ef18ffd919544

Observation ea843537-bd90-4f2c-afb7-5a1fd2146f67 · outbound

This paper cites In Proceedings of the 2021 Confer- ence of the North American Chapter of the As- sociation for Computational Linguistics: Hu- man Language Technologies, pages 3849–3864, Online.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models In Proceedings of the 2021 Confer- ence of the North American Chapter of the As- sociation for Computational Linguistics: Hu- man Language Technologies, pages 3849–3864, Online

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:08:46.294415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:f6750ba08f1e061426a57c6a4909809837cce5ff5b8466898b529d3703b15067

Observation d52b7aab-3f4a-4ef9-b7ea-75b40fb9e81c · outbound

This paper cites In Proceed- ings of the 58th Annual Meeting of the As- sociation for Computational Linguistics , pages 7871–7880, Online.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models In Proceed- ings of the 58th Annual Meeting of the As- sociation for Computational Linguistics , pages 7871–7880, Online

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:08:46.728908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:e8f001d8c5cce66b453fd8d1e4234529a1a247e8c533535703b12b73190eb052

Observation 586211ca-043b-41c1-b3d3-7be86b9c6c8e · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T02:08:45.148408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:62755b02b6390d8de53d2ad955cc080f671ca69fc6db9b5e649bfdcd15ec86e1

Observation 38d6e5f6-cda5-4569-90f7-5b4ef3f5d6d4 · outbound

This paper cites Are Emergent Abilities in Large Language Models just In-Context Learning?.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Are Emergent Abilities in Large Language Models just In-Context Learning?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.153483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:2ec88161080b1507a88ec896986d8d7768b68957621facc74dfe639e07ca3055

Observation e7fe15f2-ee82-42fd-9a94-9449eb690554 · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.143108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:cb9752b3baf71d34f8cd48c6715ebfcbf17c221958578cb5f329e604477af694

Observation aec9314c-16e6-4227-94d2-049a26902099 · outbound

This paper cites Efficient Benchmarking of Language Models.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Efficient Benchmarking of Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.131464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:4ece3d308e1b955bbb38b0c810de09b622401781bd0c06e7a8619a5ecd5c25ea

Observation 5ceec790-30b7-42db-8968-3a32bb47a0a5 · outbound

This paper cites OpenAI blog, 1(8):9.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models OpenAI blog, 1(8):9

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T02:08:46.298378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:98b19f4428cf9ef39a34b67c7699276ac936296bef6aacb87735e69eb9df1ae7

Observation 1a557266-7cd3-4818-929d-f9a28db54faa · outbound

This paper cites The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T02:08:45.125550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:8f17eb05534a8273a35feb2fcffc058ba4b885d51ec9aa98f76895634f4249a4

Observation c9cf5e73-4b8d-4422-b065-beeb432dbc00 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T02:08:45.119104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:2752b4f6cb210aaa89f4f21c7937c5ccea8bd8b4f9886f3b5c55121cb46f5fa1

Observation ed773011-57f7-4197-bad1-87acc9795a4d · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:08:45.137093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:b71d029d1bc9bf968feb70c997dbad7dcae32b1c7639147824d9dd2e103cce80

Observation b3d357df-2617-4f84-9b46-f20eaa6c3ec9 · outbound

This paper cites PRobELM: Plausibility Ranking Evaluation for Language Models.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models PRobELM: Plausibility Ranking Evaluation for Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.158812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:784a6ba2375544076b9b5ba2a9b90e1eacf3db142de677a1b3879e03aab1b929

Pith citing papers

No inbound Pith citation observations are available.