Pith. sign in

Paper Citation Record · LEDGER

The MiniPile Challenge for Data-Efficient Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2304.08442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.08442 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:52:07.511745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T14:03:20.499861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ed6fd1fb-cf23-465d-a8f6-d8ba531c2c82 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models The MiniPile Challenge for Data-Efficient Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:17.035055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:8e5ab5c4ed5c6b5a5d69e314e85c71d5351564c5759b8520efacd89b8e9f8f13

Observation fb9a4ed6-9ec1-47c4-a061-9f419192f01c · inbound

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing cites this paper.

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing The MiniPile Challenge for Data-Efficient Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:52:07.511745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:52:07.511745Z digest=sha256:7afffd02eccbc96c919ef42355e3ccb9e00ecac7d297814c24e5bbb13417f7bc

Observation 7ee45e5a-4790-48c3-9e16-75edeca7d294 · inbound

Sparsified State-Space Models are Efficient Highway Networks cites this paper.

Sparsified State-Space Models are Efficient Highway Networks The MiniPile Challenge for Data-Efficient Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:45.649233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:45.649233Z digest=sha256:095e2ce3ae901ea5ce162807c53e23fd30c04ce44658a8361f0fdb65e4e5f8ce

Observation 05fcb154-0386-485f-ba5a-f8a97e565261 · inbound

Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation cites this paper.

Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation The MiniPile Challenge for Data-Efficient Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:59.498544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:59.498544Z digest=sha256:d7607f9bfd3ce47728f70677907b0e940a1522a4a7dfbb9e736d89d1dedc831e

Observation dc16a7db-3828-4ed7-92ae-c9be9a81ac67 · inbound

Causal Estimation of Tokenisation Bias cites this paper.

Causal Estimation of Tokenisation Bias The MiniPile Challenge for Data-Efficient Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:42.992139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:15:42.992139Z digest=sha256:bd1b6805c3b43fc907abb80c0dfe04f05557378c319f4b9ca77546d2845d3e3a

Observation a69bde22-87cd-4b7b-9563-486cfafb349e · inbound

ByteSpan: Information-Driven Subword Tokenisation cites this paper.

ByteSpan: Information-Driven Subword Tokenisation The MiniPile Challenge for Data-Efficient Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.816970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.816970Z digest=sha256:c85b2db24537d3532a5339b6b3c43cd0b3df856db11eb0c8ffaa40f00a787b56

Observation f96c99fe-5eda-4be2-9f39-2bc3de2d429d · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding The MiniPile Challenge for Data-Efficient Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:26.663156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:26.663156Z digest=sha256:46290bc31f9c9c5bd6c06f7407495863aa38357cf019f22e2981923dec1be157

Observation e32e8cf1-5355-4bbf-8076-3d0e8174fe62 · inbound

Dataset Ownership Verification for Pre-trained Masked Models cites this paper.

Dataset Ownership Verification for Pre-trained Masked Models The MiniPile Challenge for Data-Efficient Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:04:35.614370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:04:35.614370Z digest=sha256:0da058938b0fab5a42dc714482e84e85edc1913cf36493630d69b2d0694312aa

Observation ff07b9ad-9655-48f1-862f-f70633ea58ac · inbound

Faster Superword Tokenization cites this paper.

Faster Superword Tokenization The MiniPile Challenge for Data-Efficient Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:52.364277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:51:20.043728Z digest=sha256:f470493a122e3b7b640dff46497680375651c14c3008aa9811c2bb3a83534a51

Observation 12e9c6f7-2127-4056-bb17-b3408965877c · inbound

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer cites this paper.

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer The MiniPile Challenge for Data-Efficient Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:16.118183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:31:33.619232Z digest=sha256:c458a48ae9e59dd8716af70192c4f2f48486050c105d86abb1d0fbab79e8d333

Observation 98073453-e083-40c0-afba-e9e15e0897d5 · inbound

LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models cites this paper.

LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models The MiniPile Challenge for Data-Efficient Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:03:20.502894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:58:55.899958Z digest=sha256:945c32f24d6239bb63930158d3817e7c99486d28bcb96469df684ed818c67080

Observation 5414b1d7-748e-44f5-81c1-a6a9601f70c5 · inbound

Joint Optimization for Greedy Longest-match Tokenization cites this paper.

Joint Optimization for Greedy Longest-match Tokenization The MiniPile Challenge for Data-Efficient Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T23:47:05.969487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:47:05.969487Z digest=sha256:477cd6cc87fb7db5596663f5bdeb8e0f464135900f28dc294bf77532650107b5