Pith. sign in

Paper Citation Record · LEDGER

Training Compute-Optimal Large Language Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2203.15556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.15556 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 606 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 64d1c609-479c-44ad-82b8-69e05d983e04 · inbound

PaLM: Scaling Language Modeling with Pathways cites this paper.

PaLM: Scaling Language Modeling with Pathways Training Compute-Optimal Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:07.228435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T23:45:06.755839Z digest=sha256:49e3dce67c33def1cfe71da39128fe70f693a5e8fd460e3b61cb414fed49f785

Observation aa725d81-d872-4190-8027-e3d12dab574a · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Training Compute-Optimal Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:34:28.405246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:2f04d636e08b51be3c7fa2546848e4ab698cd0e36cccf139cf365748abeeed44

Observation dd258193-3a6c-41fc-8511-69f5ebe27dcd · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Training Compute-Optimal Large Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:22:30.370781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:7b662291939c2063a4600dcba97e781a2bc6a4012b7b46c3dbef9f11b495848a

Observation e1cf7fc7-07f7-4438-8da1-e2a00847bef0 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Training Compute-Optimal Large Language Models

Reference 259

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:53:17.086220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:439559b6f20266ae6b93879b0c743df2267f9b93c16575ee87160dcbb6416858

Observation 8d8ebddc-7e88-4459-b6d1-194b8b3fbebd · inbound

A Generalist Agent cites this paper.

A Generalist Agent Training Compute-Optimal Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:24:50.014544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:24:49.833638Z digest=sha256:adffb4760edfb186c71b642483083e55f9d2ba7950b1162846d08e5c223b82a6

Observation 9db5af05-c262-46bb-8acc-55079721d816 · inbound

Scaling Laws and Interpretability of Learning from Repeated Data cites this paper.

Scaling Laws and Interpretability of Learning from Repeated Data Training Compute-Optimal Large Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T15:52:40.515825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T15:52:40.335080Z digest=sha256:b0deb3bcf3ac80283861cfd0739bed4c9e69af157379be3abdbe8e1ad5956fe3

Observation 5d746373-c98f-49ae-9ccd-bde0f9bb1437 · inbound

Teaching Models to Express Their Uncertainty in Words cites this paper.

Teaching Models to Express Their Uncertainty in Words Training Compute-Optimal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:36:08.608477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T17:36:08.566658Z digest=sha256:8d023706f6ebe4b00ac8c94868c8fd532cb144334967fb330f49c04a8286e3d3

Observation 97d62675-bd67-4532-b373-062df090a1e8 · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Training Compute-Optimal Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:38:38.342977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:118de90f2fc9ab2a5b8adacddda26221fa0f63dc0c9af821d6bb213187597c86

Observation 7de39fc6-bd99-4eab-bd79-7b2d76bb6058 · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Training Compute-Optimal Large Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.020705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:70ad6afc9ab63c7c22b5635190233dcfb352289f4de668da00f9513063a2046b

Observation 46655cd7-ecda-42f1-ba82-91df5b74823e · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Training Compute-Optimal Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:1fe20e76520d503bbf383415f478c1c634cb0cb2676fd9763b4f4fc445e1e922

Observation 9e36862b-b4b7-47e3-9794-7709d116ea52 · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle Training Compute-Optimal Large Language Models

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:40:41.917817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:5f193518c96c982728d0ea6795e9878c0cd1e2bef9cd1a53b89ae5f0a9341241

Observation f691636c-c0f5-4cb4-aaee-70218c1e6cda · inbound

Atlas: Few-shot Learning with Retrieval Augmented Language Models cites this paper.

Atlas: Few-shot Learning with Retrieval Augmented Language Models Training Compute-Optimal Large Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:48:43.193411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T13:48:43.024120Z digest=sha256:28c237abd4d66889a7531edc44719ec527aae74855721ca23108fa379163aaaf

Observation 3fa06da8-67e6-462e-ab03-eec923b1017e · inbound

Atlas: Few-shot Learning with Retrieval Augmented Language Models cites this paper.

Atlas: Few-shot Learning with Retrieval Augmented Language Models Training Compute-Optimal Large Language Models

Reference 198

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:48:43.399942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T13:48:43.024120Z digest=sha256:6c219b2e42c48f6d480f365d7f95b468b966bd8b915529221b9b9ea955324fa9

Observation bd46aea6-f635-4995-b44e-618234880b36 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Training Compute-Optimal Large Language Models

Reference 139

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:35:36.095713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:0132c9050cc4d28ac69e707d358bd26593a61d18f9c88800485e29895a75e3f0

Observation facb64b0-41e1-457c-98ec-e9595fbc23a8 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Training Compute-Optimal Large Language Models

Reference 167

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:29:06.203954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:b54e3afd5649a07b777a741a6bbcaebecfdb88bea99da8a0f273db208c019328

Observation db24b1b4-0fcd-41f4-a370-2b0ba4560c7e · inbound

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them cites this paper.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.829606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:b5b715636421c4457460f9ffa94fcb97523412d5c286f740090ff0d4814cd789

Observation 7ecdc3f2-8d1f-4838-aa08-560d3eb87ad7 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Training Compute-Optimal Large Language Models

Reference 248

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.593102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:4a531494e0af5cfef165b53d4b144b028c26d90c0cd85a5f278466b385580392

Observation f89287c7-824d-40a7-b916-58d09275a5f2 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Training Compute-Optimal Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:53:21.934761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:71c38cde524aed946bd52ec73b1732fca7570d3639ffbaf7da7be6197c4ef8d6

Observation b057a6b7-f00c-4452-b867-ba59bfccf88a · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Training Compute-Optimal Large Language Models

Reference 176

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:53:22.164130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:d90b2fd907c2285b1f4fa3f43426dfd0fc0b7098c72337ac8ce0365dab012726

Observation e7a157a1-fe20-4df7-8171-411e64e2c0b9 · inbound

Solving math word problems with process- and outcome-based feedback cites this paper.

Solving math word problems with process- and outcome-based feedback Training Compute-Optimal Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-24T11:14:23.102902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T11:10:40.864420Z digest=sha256:6e09604478597be357b6b1c692ef2d719c81abb2404bfee78cfa7b2219048c1e

Observation a9fbe08a-bec6-4fde-9d37-ff1daf047068 · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Training Compute-Optimal Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:09:12.992496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:1e447fedaf1fe402b152a3f85394ec3dbfe799f1265bda71b7030b3d27b3bfc0

Observation b907b3e8-0575-46df-807d-07c014390d37 · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Training Compute-Optimal Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:10:49.405653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:467507b79365c95b8f8420eacb50e90f6fecdca7bb80d7392df9fa2088de9092

Observation d6ca6895-0a4a-4767-bcc2-c6898a3ef7f7 · inbound

REPLUG: Retrieval-Augmented Black-Box Language Models cites this paper.

REPLUG: Retrieval-Augmented Black-Box Language Models Training Compute-Optimal Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:41:54.119942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T12:41:53.833754Z digest=sha256:54f4a85f0a9d7205647fe1f51c7462227754cd8eed6a399f3f1a179154fe84f3

Observation a2645ea3-9151-4172-b565-948c48458d92 · inbound

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning cites this paper.

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning Training Compute-Optimal Large Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:14:16.417815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T09:13:30.054153Z digest=sha256:7c85934186c33f0247e1a61061cde70ee8bfb911c7c1b42b7fae97bc532c8b37

Observation b27ea298-d4ef-492a-a789-7905ad802049 · inbound

Accelerating Large Language Model Decoding with Speculative Sampling cites this paper.

Accelerating Large Language Model Decoding with Speculative Sampling Training Compute-Optimal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:29:36.298332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T07:29:36.205026Z digest=sha256:857a2c6a725a6e74a6e456124be4eefa6170a436aa93dac8409a600e7ce42806

Observation c7f612a5-867d-43ee-8088-9cfae7f0d2fd · inbound

Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation cites this paper.

Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation Training Compute-Optimal Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:00:05.374302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T18:00:05.294144Z digest=sha256:12d083ab8888d20405fe4e485a63cdcb42ca298341cb4a1e96e7f4bd278e9a5c

Observation b10c6990-fff9-46a3-a839-0c07a42045ec · inbound

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models cites this paper.

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models Training Compute-Optimal Large Language Models

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:11:21.831689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T15:11:21.589770Z digest=sha256:07e6b63fcdd7ced8cb3d8c0b9ee2a6e1f0f7610d92710f10278eae36eb7516c9

Observation b9fcbaf5-5361-4f22-b31a-8a9fb04123ae · inbound

SemDeDup: Data-efficient learning at web-scale through semantic deduplication cites this paper.

SemDeDup: Data-efficient learning at web-scale through semantic deduplication Training Compute-Optimal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:43:30.936629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T02:43:30.851915Z digest=sha256:d3704d3dcd9615ce342da944afde4dda04582769bce3d9a7af9d68d2970f6867

Observation e0607f81-b746-4939-8f3e-86e75aac89e0 · inbound

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society cites this paper.

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society Training Compute-Optimal Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:40:53.778903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T01:40:53.351795Z digest=sha256:41447e3da17142c992499aa6492842f372cacd380a38abfaafd7d77cfb00fa22

Observation fa94eff9-bf58-4286-bfd0-e7c204c116c8 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Training Compute-Optimal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:46:39.705991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:7011ee9202cbb1df8d2956fd01f92adc3d8683bc5902d8b3e709ac37072f660e

Observation c3a28a78-2b98-4d1c-beb6-d178e62256ad · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Training Compute-Optimal Large Language Models

Reference 125

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:45:17.981519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:d33204f81555abf605b7578456cb95022fb8d0cf1cd21ffe7522630515a8e9c1

Observation b37f237e-fe82-4f10-ac72-a9d7e079c571 · inbound

Segment Anything cites this paper.

Segment Anything Training Compute-Optimal Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:14:20.214356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T06:14:19.815357Z digest=sha256:0f57a0e718ef2255936bdefc0b6d93d2a0f54c26adb133ebabd5fc5077110b9c

Observation 4af3b7ae-2251-40ed-b1ce-d47a43a565c1 · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Training Compute-Optimal Large Language Models

Reference 113

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:46:56.773494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:10b927e5f11165139a5676e81991f925710bb8479bc4a47cd7505a754302d6cc

Observation 6da0582a-384f-406c-b65b-f9bed07efe79 · inbound

DINOv2: Learning Robust Visual Features without Supervision cites this paper.

DINOv2: Learning Robust Visual Features without Supervision Training Compute-Optimal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:a554a16e1a729fffead8254600b8ea1fc50a45dddce9698941434d1ca66d39ee

Observation 2e535302-3e64-4efe-95c1-86138b82378a · inbound

How Secure is Code Generated by ChatGPT? cites this paper.

How Secure is Code Generated by ChatGPT? Training Compute-Optimal Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T10:04:18.968874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T10:02:22.456932Z digest=sha256:76acab98bdc9b25c0aa8ff76206b6173783e3900a3475a247c0689cb7e944e80

Observation 3fad26bb-b1e8-48c9-86a5-a9c9c38bdf0e · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models Training Compute-Optimal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-10T20:37:01.852070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:06b4c70766bb0bf7bfe4406548a6bb7da4caa9260aa02b929d44b2a1ecf9010d

Observation a1e2c194-0da6-42c0-85cb-3561303ece12 · inbound

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes cites this paper.

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes Training Compute-Optimal Large Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:50:09.426631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T20:50:09.265838Z digest=sha256:613cd2ed5b1eb7c38ce212190ec38b553df0ec8b2b8cffccb5afc7bcde3f8066

Observation dba7ddbd-db91-49af-a6ad-9d90a26c2571 · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Training Compute-Optimal Large Language Models

Reference 199

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:33:00.893432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:d65266d2d97df809c2166becc4706a5bb5f09cdf9402f934dbfc644d7df8cc49

Observation 857a1af4-2bb3-44f1-b81e-2944113ff07d · inbound

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? cites this paper.

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:36:55.142743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T07:36:55.087443Z digest=sha256:0dc0281b63b6f671ace04abb04a978d153d4a857bb317a35ee0256d66e2aff8f

Observation c1a50e12-f52b-4b81-83fd-812b4a4c07ff · inbound

Towards Expert-Level Medical Question Answering with Large Language Models cites this paper.

Towards Expert-Level Medical Question Answering with Large Language Models Training Compute-Optimal Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T04:32:33.397004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T04:32:33.271634Z digest=sha256:3cc776015e492c755b4c254f23f2d73a955eb42a396fdb959096045a6a00cc2b

Observation 85bc7124-dea5-42e9-b05e-e4d7e9bafd55 · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Training Compute-Optimal Large Language Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:59:27.164344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:106a62776ef446bc66453c2fde4958861247f8365aaef5e03266a6ded5ab5a4c

Observation 3a58ab4c-492a-43da-92c6-7bf8d5a8c918 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Training Compute-Optimal Large Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:35:21.644705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:f486fd1df4580ef25ed769f226549bd5c1295723ed2799fe118053415dc4c4c5

Observation 1d1c0c9d-8ad7-4973-9d17-b7251f69fe4c · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Training Compute-Optimal Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:43:45.848361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:ea797b86ef64b8a7c07fda8f53e7411a8087891e916b58e5500479d46e13e5f2

Observation fca11763-1ee7-481e-8a5d-0a044294e676 · inbound

Objaverse-XL: A Universe of 10M+ 3D Objects cites this paper.

Objaverse-XL: A Universe of 10M+ 3D Objects Training Compute-Optimal Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:02:11.610496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T13:02:11.512409Z digest=sha256:fd6b440e1713e667faf2c62afe04c498084b37480c2dbff38072b3894102f2f0

Observation 49b3f0d3-f2cf-46ae-9c6b-d6f1a001fc31 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Training Compute-Optimal Large Language Models

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.339978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:bc33c955587d97024c4c710d52293ac3934fd2465890aad6d94a9a4639018285

Observation a7e7866a-25a7-452f-8f09-7270dc0f3be9 · inbound

C-Pack: Packed Resources For General Chinese Embeddings cites this paper.

C-Pack: Packed Resources For General Chinese Embeddings Training Compute-Optimal Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:24:32.209644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T13:24:32.084878Z digest=sha256:79897dfb68ae5ce9b910237d59740e4f9fbe7828a2d837cd6950d3708fc73d1f

Observation 47e0fbcd-31fb-4c15-a0d2-d5d89badc8c7 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models Training Compute-Optimal Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.567927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:4a003a8a87b419e59be10ff078f90f99cc847d0e626a82f5c5fa6ef344f218fd

Observation d4fa4b8b-6765-4c85-b423-171e5b585cbb · inbound

Language Modeling Is Compression cites this paper.

Language Modeling Is Compression Training Compute-Optimal Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:36:13.455055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T22:36:13.392250Z digest=sha256:870d3122f56fa054f9cdb2dae568b9578d6e0ba28b75c7d3ccfe74edcbc9d6ce

Observation 3b16d98c-d430-4f1b-b779-5c1a9f45a305 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Training Compute-Optimal Large Language Models

Reference 150

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:12:31.599118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:356fe5f4d7f8327d9f16cfea9fd3b85f021fdd69c30841fc5b80ad68b8a418ee

Observation df0020b1-7652-4c94-ad9d-d41d6bc9403d · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Training Compute-Optimal Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.316183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:cfed8d8dbb6be5bcd7ce53f1085057d1cfee305476e14dac59d48febcf60d68c

Observation 7b858d8b-11f9-4169-9382-e54d61608af6 · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Training Compute-Optimal Large Language Models

Reference 229

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T02:15:18.479900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:65cbe4f43a69cfb9d9e80f8d338611fb071cd219fa5bed26b0ce5881e2378a33

Observation 49d8ca2a-ebf9-405c-b1e6-7a8641dc9994 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:13:09.005969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:e862e2213598cba6ab4510463cd384888244d31106f437d98bda6936bec80b2e

Observation f6b58561-b258-4259-b2ed-5e681fa55d3d · inbound

Llemma: An Open Language Model For Mathematics cites this paper.

Llemma: An Open Language Model For Mathematics Training Compute-Optimal Large Language Models

Reference 149

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:17:46.477044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T08:17:46.055279Z digest=sha256:b1fe68da41c604da4f331501f084d7b7badc3ffae5f9f722a8258927555dba4e

Observation 1c6e46a8-f7b0-4b3d-9a34-47f45169629e · inbound

A decoder-only foundation model for time-series forecasting cites this paper.

A decoder-only foundation model for time-series forecasting Training Compute-Optimal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:07:21.307923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T18:07:21.246053Z digest=sha256:cfcb72547693eda86b29374984d201b2ceb531233a8c3533ebd331e4d7d0fc86

Observation b76660d8-9d32-4a5a-8395-61918991be23 · inbound

BitNet: Scaling 1-bit Transformers for Large Language Models cites this paper.

BitNet: Scaling 1-bit Transformers for Large Language Models Training Compute-Optimal Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T05:43:56.559339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T05:41:15.544164Z digest=sha256:709913dcdb91f32d153ca1fb760bd1eaf6da8057125a4f8e4fbbd79eb979675b

Observation 3d964446-55f3-435c-9e3d-a8675e6846e2 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Training Compute-Optimal Large Language Models

Reference 292

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:46:10.131134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:4a38203d3b4544b55657e9dec235a6efc147055ed402360805a2819373f0eef7

Observation 1642cbfc-a9b9-41e9-b731-157add39abe5 · inbound

A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA cites this paper.

A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA Training Compute-Optimal Large Language Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:05:50.603625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T22:05:50.453130Z digest=sha256:78ad03e7c51e92e7da1d06740b3d9c56cde71ac17cb1f678f1e2a9ccefae4744

Observation 84fc2006-19e4-4456-8f1d-a88b48ac71d3 · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Training Compute-Optimal Large Language Models

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T05:03:55.325041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:18aa3cade742656c3fa48458ad17ceae0fe7b4832c72cec2ad735ff164ba8bdd

Observation b4df69e2-8710-478e-8ae2-a72a8df3a3e3 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Training Compute-Optimal Large Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.028539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:6d85c2de7944a7ced16f58bdf9385630f60d37ee3130e5ba36d3219bad8cd3c0

Observation 31e4c1fd-1109-4958-8f00-b131e8a797fe · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Training Compute-Optimal Large Language Models

Reference 144

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:08:05.926341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:eef8c31d8671e4e48668ef8e8590db3153d8404d8be0013a1ea8ec9156f7944e

Observation 576152b0-6f81-4d4b-97a0-89d797b42952 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Training Compute-Optimal Large Language Models

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:17:08.467578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:7f0ebaf09f370884a9a7487cabd51876a42e6b6614e8121f962f900c19c21916

Observation 417b9257-a802-4c92-9b22-7dd90ed1baf7 · inbound

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models cites this paper.

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Training Compute-Optimal Large Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:50:09.661617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T22:50:06.399707Z digest=sha256:c3244e5938ce37503ae23c60ff545f6a1154f3c65459ea9e668b647c14dd2b5f

Observation d895236b-47cc-41f3-a560-a8ee720904e2 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Training Compute-Optimal Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:36:18.077733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:9ab975e9d37f694c02401a20cd5273b36718a76edeeb4401814dba3d8f7e264e

Observation 37d157d4-34fb-4407-a4ed-aa71288e74a6 · inbound

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval cites this paper.

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval Training Compute-Optimal Large Language Models

Reference 138

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T13:07:16.448046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T13:07:16.151160Z digest=sha256:984d808cf99a73a5ba76dc65ceb80d16bb8502d8bdbc22adb613b9ff31ceb564

Observation 7ae8a20e-a170-4570-9d03-5a72b6a02dfd · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey Training Compute-Optimal Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:22:54.480517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:e087cd79282a87984b3c5ccabfd5fe386befc51066ee8a566252870fc812a049

Observation 88e56c38-da17-4c57-9681-d19afd23902e · inbound

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models cites this paper.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Training Compute-Optimal Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.513427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:b1ca778d97b818919082666b09a6ee06891a8661a9f67ad2140019f71afe07cd

Observation 5f7487fd-3218-4f37-86bc-05556fdba3e9 · inbound

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts cites this paper.

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts Training Compute-Optimal Large Language Models

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T03:08:48.180582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T03:07:32.103513Z digest=sha256:58cc14cb033c5e19800847517671eab4a6fcc79a2858d7d562e2680820536669

Observation 31181b02-1b02-4ddc-a044-6156c5fb73e5 · inbound

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders cites this paper.

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:05:56.983966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T03:03:51.053556Z digest=sha256:5b1bf764ddf03b59549154d86b9e855f6e8362ba212e48f7291c2acd1ff93be0

Observation 22adb4d1-4cce-4831-83bc-a2569fe6c6c4 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Training Compute-Optimal Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:47:27.856225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:b66b1b4a9ce5000afaa2a9f8436afb3fe1adc377b665e621aaa76cee1f5321a5

Observation 4c611a33-95e1-4524-befe-dfc59eac2804 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Training Compute-Optimal Large Language Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T18:00:53.423895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:0dec0796a1baf1c98c0cde8bb77c767aae2dc5b8944f6b709ede46f90b78ffb9

Observation 8650ba1a-f5b5-4608-968d-a33c0593582c · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone Training Compute-Optimal Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:19:27.344354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:6d63fcb9cfaa50c6804e5569debf4d567a5a4043cfae2a443b404c55d7557305

Observation b902fd47-4ea1-4d87-a98f-204138b27bcc · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Training Compute-Optimal Large Language Models

Reference 139

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.635339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:a5fa3aca4a8c44bac6efc58fdc2b0ff785e388b59a2f524a610d84dcf1f9d979

Observation d370648a-a0b7-4758-842f-b5984f2273ea · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis Training Compute-Optimal Large Language Models

Reference 147

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:03:56.499180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:f57fc48b80837b5836f75ca780d195dfd862d0a2af12abc1e6fd6f50335bef0b

Observation fc1ecde2-0d2d-4a9b-b434-0781c3ccb0e7 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation Training Compute-Optimal Large Language Models

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:06.484041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:1cd954d5b83d0091667d74190d42d38434fa88b1ae32eec2cab57a917fd8752a

Observation 74cac238-ba98-4e31-bebb-2bed1a78721e · inbound

Scaling and evaluating sparse autoencoders cites this paper.

Scaling and evaluating sparse autoencoders Training Compute-Optimal Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:47:23.225720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T17:47:23.089288Z digest=sha256:b88b487a85a46aec87ce3da91fc5cda437ac48ef48daebec55bc2f6e1023cbe8

Observation cca4f1e8-62db-459b-8c09-95105b3d1d91 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Training Compute-Optimal Large Language Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T23:10:40.969748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:569398f463708b2245285cfebcbfbd8905bfda8962511fe5156155d68538b5b9

Observation 706efc08-1160-48bc-a18f-47494a64b170 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Training Compute-Optimal Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:09:16.900368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:4b1e7069d734c6d4a57c29677313ed4ec57915f57deaa23b0ecef62637415bc5

Observation 4bc80752-f624-4bd0-9b7b-7b9bd7e6b4c3 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Training Compute-Optimal Large Language Models

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:58:17.005898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:d82c2893d90964d4f970d200e47fcbd92906f34617fa57394da67aab214bb1fc

Observation f7503109-87b3-4183-af31-c2bbbfb18611 · inbound

Retrieval-Augmented Generation for Natural Language Processing: A Survey cites this paper.

Retrieval-Augmented Generation for Natural Language Processing: A Survey Training Compute-Optimal Large Language Models

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T23:08:35.719443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T23:06:41.081461Z digest=sha256:38d980368e7936da6e5c1e8db755100d31b4d68cb9b15c48d0752403e9cfedeb

Observation 0c9edd13-1fe8-477a-ba4c-7433b7a29ae9 · inbound

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling cites this paper.

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling Training Compute-Optimal Large Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:42:23.571168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:42:23.297389Z digest=sha256:ec1a57101fa5bad9f19320a15efced3f25b9064164a1a8db3b905188eff49754

Observation b4d8f6dd-a4ad-4c80-b38a-a7c82448d183 · inbound

Gemma 2: Improving Open Language Models at a Practical Size cites this paper.

Gemma 2: Improving Open Language Models at a Practical Size Training Compute-Optimal Large Language Models

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T12:11:16.326752Z digest=sha256:a0b86a82e62df9cc208fced7b8aa3ea1b73f86b0251ee62c069a5a202a99f91e

Observation 36a23d8a-381f-43fa-b873-1f7e453df19a · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Training Compute-Optimal Large Language Models

Reference 146

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:36.878637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:b1e5f37c19c8779247b8a81f4778b3073aae9c90f62f9f991aa1de962ba1b1a1

Observation 6a91c791-6dcb-4108-9cd7-332b2482e400 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Training Compute-Optimal Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.561356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:591f91cf9583cc8ae3eb6dac7ceeeeb84be9f4d9bd27a814a4ba9ac955453fa4

Observation 6672f655-71d9-4a21-9daa-bfff810d6d1f · inbound

Optimization Hyper-parameter Laws for Large Language Models cites this paper.

Optimization Hyper-parameter Laws for Large Language Models Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.808999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:5eced31528d94b279a67810fb47a134f61f64de351fa274225de8aa0f2e4e340

Observation 8b1f3197-1f63-4abf-9e52-873a5ab2f18f · inbound

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio cites this paper.

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:43:25.300044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:42:38.782232Z digest=sha256:9a56da37251cd13727aff714c909a59b3f8102cf35dde7add61ce75cb659f77c

Observation d048c94f-d0c4-4984-abc9-4f37295afb07 · inbound

Moshi: a speech-text foundation model for real-time dialogue cites this paper.

Moshi: a speech-text foundation model for real-time dialogue Training Compute-Optimal Large Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:13:22.130283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T08:13:21.962488Z digest=sha256:a8a01cbcba9d42d0d4bcfd64a82766617f1ee2de7effbc5b74f5b38dd5b6e3c3

Observation d920897a-5533-4b8d-bb5a-9aea289eefde · inbound

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs cites this paper.

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs Training Compute-Optimal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:35:47.087940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:35:29.917096Z digest=sha256:8330536281b883557c92b00f780f460f198403c2a7276e03d91d0fcdaff72e44

Observation 3748e25a-ea43-4cb6-b32a-6aaf71edd70f · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Training Compute-Optimal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.385871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:a72287922fd4401c34e3e7407dac622aa9147f47a607c33165ec79851d4e3292

Observation d1719ae1-6120-4019-b2f1-392737d9254a · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Training Compute-Optimal Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:25:27.889490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:602f7341ec40b7b859895d8828668ffcb152e5f7fcfa649954e0827cc7b182f8

Observation 053fea96-b57e-44d2-9e8f-a157ac0ff533 · inbound

SEDD: Scalable and Efficient Dataset Deduplication with GPUs cites this paper.

SEDD: Scalable and Efficient Dataset Deduplication with GPUs Training Compute-Optimal Large Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T06:47:39.612475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T06:45:53.685914Z digest=sha256:7968330b8fba59c62fda89eaf17f87d371be37c0bbebca709134b205eaac965b

Observation e4356763-83f6-4bd1-aac7-3d20e124f756 · inbound

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps cites this paper.

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Training Compute-Optimal Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:45:17.645576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:45:17.473970Z digest=sha256:0d4a204672ed9770e4f9f1f7723a44e8b4312b0ab56df56235436af67387ad8f

Observation e61f9261-cd1e-4ecb-8fd6-3803923e571f · inbound

Exact Sequence Interpolation with Transformers cites this paper.

Exact Sequence Interpolation with Transformers Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:45:21.762074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T03:43:01.071221Z digest=sha256:8e3a7b0f7ed1c21a71efc3f425708d65bbf66727c017f99e94a4d5499d22e229

Observation dac5b05d-fb2a-4628-aaa8-5f840127c8fe · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Training Compute-Optimal Large Language Models

Reference 180

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:30:02.964296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:ac41635d78e4cf819b95a934cb634b8dcde0e17afe55290c8d9b794d45b9f301

Observation b5d53e79-90eb-449c-bcee-cf838155a9d5 · inbound

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models cites this paper.

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models Training Compute-Optimal Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:31.021149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T04:16:04.110552Z digest=sha256:9d616437d3a8b87970b65aff68816ee481fe9e26bc8907819ca79433ded6564b

Observation 8414832e-bc04-4502-b6b5-da98f555a9df · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models Training Compute-Optimal Large Language Models

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:42:54.834679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:5bcf3bfed68706b1b436b43927495e7f154867e5f762b63c8592450f72e1d336

Observation 826614d5-10cd-4aef-a68f-0fd3d0483e1f · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Training Compute-Optimal Large Language Models

Reference 198

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.345446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6ff5f75674d7e637f21a753ad86665a73ee0986a4f4679880ff8d6c3c714b0cb

Observation 6b70340b-707e-4ee4-955a-5fa53d4ac107 · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Training Compute-Optimal Large Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.077172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:b07e753853a9b485b5abc963d147c3ea3d8df4cb79e129bf8a309170d264ebe1

Observation 4c8bd24e-30ee-4361-9f20-a39bb6978cf8 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Training Compute-Optimal Large Language Models

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.127680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:c93028e01bada405768d72a8463324e8670da28c8ba75ff6beb10e397adb68e8

Observation d5d4ec90-366a-4a8d-8a5b-13a577a8c122 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Training Compute-Optimal Large Language Models

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:44.656600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:251a25039f1fda3c2a20fc5ebfec5d8a34aeac314e431d50077d388da0620c43

Observation 94999a6e-dad0-46d3-8871-95e63684b8b0 · inbound

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices cites this paper.

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices Training Compute-Optimal Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:05:16.456726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:03:26.037233Z digest=sha256:a9ee27a23f23051eff607c1049747916f549c5b256228415a486f21a033e7740