Pith. sign in

Paper Citation Record · LEDGER

OLMES: A Standard for Language Model Evaluations

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.08446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08446 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:20:38.219629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:48:39.417321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9345277d-b21b-4c45-884f-28de0bdde6b4 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model OLMES: A Standard for Language Model Evaluations

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:30:02.953489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:94b892b253a9d7102a5b6a2722d8678216c0e3108af04ff0c59dfdb1d99efa1c

Observation ec5e14aa-2d17-4068-8c94-99f3c1b3ecea · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.219629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.219629Z digest=sha256:8eb5c1dc0cd9a1205e7f85d3bef03bce44b1a4ff6e8d8b583a7af605abae4830

Observation f8ebd05b-1524-4a6b-bb3e-532e0e596892 · inbound

Typhoon T1: An Open Thai Reasoning Model cites this paper.

Typhoon T1: An Open Thai Reasoning Model OLMES: A Standard for Language Model Evaluations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T22:52:54.314631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:52:54.314631Z digest=sha256:7e21b715c2e5c3d4e9d2a3b6ab38dd57ee28126d00726356b579b94177a0bc8a

Observation 5bf3399b-c513-4526-9484-535745367f70 · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text OLMES: A Standard for Language Model Evaluations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.082470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:44.082470Z digest=sha256:208cbdacba0103c46e5989ca89ec8d5fd548e51281063cf38086ed256535cafd

Observation b5b39e20-9c79-4cd0-a001-af7ac9eb3b26 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models OLMES: A Standard for Language Model Evaluations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:57.931539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:57.931539Z digest=sha256:bfb2925ee256b999171b92a0b868298344b45e0db4de4b573832b7b9c51dc0a4

Observation 7072c36e-25f9-4ad5-9295-6e4aa128b307 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language OLMES: A Standard for Language Model Evaluations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.365791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.365791Z digest=sha256:92efc5bbb9e4809c86c228292b63e2e0800da01addf41f88a57705eca4cf2941

Observation df52f2b2-e137-46b5-945d-8776a15126ea · inbound

On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator? cites this paper.

On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator? OLMES: A Standard for Language Model Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:55.651822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:55.651822Z digest=sha256:f34b2adb755ffaea783646f910203dcf1f0beed0f9ab46f3047f5d8ef936fef9

Observation d2334a67-9cef-41e9-bad2-19cb0492beaa · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training OLMES: A Standard for Language Model Evaluations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.130029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.130029Z digest=sha256:f85ae5248e338b3fff41e0cb55108bb3fb3a3ddd612051bef018647a4b9d6d10

Observation 2f589883-c4bb-4b03-a752-766fe9ab3b97 · inbound

Pre-Training LLMs on a budget: A comparison of three optimizers cites this paper.

Pre-Training LLMs on a budget: A comparison of three optimizers OLMES: A Standard for Language Model Evaluations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.555619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.555619Z digest=sha256:e880149d1a1484197e05e046c4b757e40634b206481d25d0a2c7ca0c9e614a83

Observation 552e56c2-64e9-408b-b04b-301bcfcf2e40 · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report OLMES: A Standard for Language Model Evaluations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.236081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.236081Z digest=sha256:3a1374c0941a9b3536a299d224f1e8b05c6848327b2449c75dc3fdc4f0fa1b30

Observation ce3cce6a-28b0-4179-8219-13eb2352f3c5 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking OLMES: A Standard for Language Model Evaluations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.414644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.414644Z digest=sha256:262733e5c0a61058613baf175e436b4836885f3df7635ef7b8fb4a3ad62faa01

Observation b5c2de02-3dc2-43de-a798-ef1e165f5fbc · inbound

PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning cites this paper.

PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning OLMES: A Standard for Language Model Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:54:45.332313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:54:45.332313Z digest=sha256:c347abb36bb6d5268f4873c5a525e99cc666ca35a08ada822a9f9fffdfe7ee59

Observation 73aa1acf-9932-48fc-a83b-11a2c10866ef · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety OLMES: A Standard for Language Model Evaluations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:eaea910c646ef5a44ff0437546a1d97a65b744ca2783cbe5daf129d13c857e5d

Observation fc53d5b3-acd6-471f-82db-093ba7b84250 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:28.316987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:19:12.696712Z digest=sha256:7f1e1ced338e6703476d14d905494cee1cf9a55f7b2e775277dc7d70f5a06bae

Observation 19f4615f-ead4-43ca-8814-52ef30da456b · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.159520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:a53db13c84c59b43242f9be07c6462c16466049317bf77f605fd7fc8c84b8eae

Observation e179736f-a2fd-408d-afde-1ea02f97e1af · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models OLMES: A Standard for Language Model Evaluations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.120033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:80e2569bb12d1e79c0adf44b345b1a79bbeb51cbe128c8c1467f7f0b6edecc66

Observation 26808366-0af1-4f91-908e-62e1eafe78eb · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures OLMES: A Standard for Language Model Evaluations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.418701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:0e57f9e3ab39af5b0316784fa0d37ccc2e88b8b0950eee1bb066df34d2d8836d