Pith. sign in

Paper Citation Record · LEDGER

OLMES: A Standard for Language Model Evaluations

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.08446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08446 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:21:06.683770Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:48:39.417321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6bcbf22-ed82-44d4-a177-9ab5992b46aa · inbound

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models cites this paper.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models OLMES: A Standard for Language Model Evaluations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.222181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.222181Z digest=sha256:5c2426648f19162c2e88979fe6744f66cece7e87f02b6c2d2b088e46c0f04523

Observation 91353968-e45e-429d-91ec-25358e755dfb · inbound

BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism cites this paper.

BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism OLMES: A Standard for Language Model Evaluations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:11:53.008559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:11:53.008559Z digest=sha256:bc92e7f838a08f30afc54102e6012c83cabe288b03312d853aad3d5cd021e4cb

Observation ce66fda4-371b-4587-868c-ba7650f917ec · inbound

Metadata Conditioning Accelerates Language Model Pre-training cites this paper.

Metadata Conditioning Accelerates Language Model Pre-training OLMES: A Standard for Language Model Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:21:29.542029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:21:29.542029Z digest=sha256:ab3e9fab2e2d9f4d5e2db9c01dd6f8f497f2353ca7655889b24a08d7160fac3a

Observation 9345277d-b21b-4c45-884f-28de0bdde6b4 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model OLMES: A Standard for Language Model Evaluations

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:30:02.953489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:4511ee4b7206bbbe87e9f8e5269d7259d03d0b1895cde67866ee4b4f89cc8e1d

Observation ec5e14aa-2d17-4068-8c94-99f3c1b3ecea · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.219629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.219629Z digest=sha256:47a5fa84e5c325d29ba02a0e87679e813e6ee2844965864da3c03f24c7e842e8

Observation f8ebd05b-1524-4a6b-bb3e-532e0e596892 · inbound

Typhoon T1: An Open Thai Reasoning Model cites this paper.

Typhoon T1: An Open Thai Reasoning Model OLMES: A Standard for Language Model Evaluations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T22:52:54.314631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:52:54.314631Z digest=sha256:45ff8cc0810e51b4c2b6794ec4a34092aab9f72de9c51eb6af6f05e44a9c8bac

Observation 5bf3399b-c513-4526-9484-535745367f70 · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text OLMES: A Standard for Language Model Evaluations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.082470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:44.082470Z digest=sha256:af2910bdd0579049ee56e4aa4d1215fa102e1d7d31a65d2c63fc0a5ae9389295

Observation b5b39e20-9c79-4cd0-a001-af7ac9eb3b26 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models OLMES: A Standard for Language Model Evaluations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:57.931539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:57.931539Z digest=sha256:53849eb3382bb42523a6eedda161008fd137b12a1307708fac5d59e9017bc4ff

Observation 7072c36e-25f9-4ad5-9295-6e4aa128b307 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language OLMES: A Standard for Language Model Evaluations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.365791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.365791Z digest=sha256:925854188cd6ed7656c3b21e632d64588d6c70a6e9e02962bbdc29b870caed5b

Observation df52f2b2-e137-46b5-945d-8776a15126ea · inbound

On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator? cites this paper.

On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator? OLMES: A Standard for Language Model Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:55.651822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:55.651822Z digest=sha256:e60b89aae237f9b692ecca53c9b57c16df5318959ea0617721c042f77f4160f0

Observation d2334a67-9cef-41e9-bad2-19cb0492beaa · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training OLMES: A Standard for Language Model Evaluations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.130029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.130029Z digest=sha256:f23947e9e99c16d4a2b816b5cb740e4a156215a658f4c834f4c11baee2ca59d7

Observation 2f589883-c4bb-4b03-a752-766fe9ab3b97 · inbound

Pre-Training LLMs on a budget: A comparison of three optimizers cites this paper.

Pre-Training LLMs on a budget: A comparison of three optimizers OLMES: A Standard for Language Model Evaluations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.555619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.555619Z digest=sha256:0cbfbb7e9ccd6297313b6de65cdc0d0b9eee98b1a11c22fbeae88b8c3116397e

Observation 552e56c2-64e9-408b-b04b-301bcfcf2e40 · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report OLMES: A Standard for Language Model Evaluations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.236081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.236081Z digest=sha256:5b0246d8e5ef5ba79fa8b3459e03a3a5e426177ad2ecfd163bee243b2f466b03

Observation eacdda79-bdab-4511-b107-0f9eb807efe3 · inbound

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation cites this paper.

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation OLMES: A Standard for Language Model Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:06.683770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:06.683770Z digest=sha256:cded2bcccf103a211d9eb32a4f051c7db4aead42b4dcb8fa3a362cb4b9090bbe

Observation ce3cce6a-28b0-4179-8219-13eb2352f3c5 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking OLMES: A Standard for Language Model Evaluations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.414644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.414644Z digest=sha256:2196aebd1b933bbdbbb042b05e1fe8c622d12fd7878c278fd0bcf835bd2cec8b

Observation b5c2de02-3dc2-43de-a798-ef1e165f5fbc · inbound

PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning cites this paper.

PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning OLMES: A Standard for Language Model Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:54:45.332313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:54:45.332313Z digest=sha256:07d7ebe4d07132f913d99988642fc40c2a2828ab740b3d8488a6a0cfb47202bd

Observation 73aa1acf-9932-48fc-a83b-11a2c10866ef · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety OLMES: A Standard for Language Model Evaluations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:27533d6c56210bcb6d4e9d5dd469332a5d8d6542ccb8bba5dad4834a426ef021

Observation fc53d5b3-acd6-471f-82db-093ba7b84250 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:28.316987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:19:12.696712Z digest=sha256:deae03bd7f2a9403314c2d1c303d4e7debf24b8235f59cb44eeb56051459380d

Observation 19f4615f-ead4-43ca-8814-52ef30da456b · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.159520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:48c6c168304ed5da8039181d70402b063897066486cba6dcd40d52b709538b46

Observation e179736f-a2fd-408d-afde-1ea02f97e1af · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models OLMES: A Standard for Language Model Evaluations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.120033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:e03a50e5f55a116288522be224afaf3285858628ec8c9f8cb505c8faa933f5d2

Observation 26808366-0af1-4f91-908e-62e1eafe78eb · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures OLMES: A Standard for Language Model Evaluations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.418701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:4a6f74407a54c4ba8fcc342a3281d3ecf89a4eb5b3f15f54665ecc3f86431845