Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T02:06:53.585629Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2404.18923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T02:06:53.585629Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 32ebd788-32dd-46d3-aed3-b31474d611b4 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d95b0e6-e4c0-4889-87bf-176f003f54ee · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Lossless and Near-Lossless Compression for Foundation Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea843537-bd90-4f2c-afb7-5a1fd2146f67 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models In Proceedings of the 2021 Confer- ence of the North American Chapter of the As- sociation for Computational Linguistics: Hu- man Language Technologies, pages 3849–3864, Online
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d52b7aab-3f4a-4ef9-b7ea-75b40fb9e81c · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models In Proceed- ings of the 58th Annual Meeting of the As- sociation for Computational Linguistics , pages 7871–7880, Online
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 586211ca-043b-41c1-b3d3-7be86b9c6c8e · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 38d6e5f6-cda5-4569-90f7-5b4ef3f5d6d4 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Are Emergent Abilities in Large Language Models just In-Context Learning?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7fe15f2-ee82-42fd-9a94-9449eb690554 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aec9314c-16e6-4227-94d2-049a26902099 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Efficient Benchmarking of Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ceec790-30b7-42db-8968-3a32bb47a0a5 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models OpenAI blog, 1(8):9
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a557266-7cd3-4818-929d-f9a28db54faa · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9cf5e73-4b8d-4422-b065-beeb432dbc00 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed773011-57f7-4197-bad1-87acc9795a4d · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3d357df-2617-4f84-9b46-f20eaa6c3ec9 · outbound
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models PRobELM: Plausibility Ranking Evaluation for Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.