Pith. sign in

Paper Citation Record · LEDGER

Training Compute-Optimal Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2203.15556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.15556 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 100 of 932 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T22:02:41.636083Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 64d1c609-479c-44ad-82b8-69e05d983e04 · inbound

PaLM: Scaling Language Modeling with Pathways cites this paper.

PaLM: Scaling Language Modeling with Pathways Training Compute-Optimal Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:07.228435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T23:45:06.755839Z digest=sha256:b78a1fc08d6d47956d1f8ef6fbc873c01836536c1759256bf436c37f63e91b2a

Observation aa725d81-d872-4190-8027-e3d12dab574a · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Training Compute-Optimal Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:34:28.405246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:40c0bee98d6f9994ea730020d6cd3641472708910989208ede07fb42ea456802

Observation dd258193-3a6c-41fc-8511-69f5ebe27dcd · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Training Compute-Optimal Large Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:22:30.370781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:2011ce46235401dbda66babbf8fb0934448d37d4943163c1a0d6ec1d29e0548d

Observation e1cf7fc7-07f7-4438-8da1-e2a00847bef0 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Training Compute-Optimal Large Language Models

Reference 259

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:53:17.086220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:5bc277ae52c1bd2ecc421b0c7a2c5093db40eba008e072553b9d5f6b3e89445c

Observation 8d8ebddc-7e88-4459-b6d1-194b8b3fbebd · inbound

A Generalist Agent cites this paper.

A Generalist Agent Training Compute-Optimal Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:24:50.014544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T06:24:49.833638Z digest=sha256:ecb066e85a13f9f4c18ab5a5e756386871a091b82ec9ee9614c4d19b8aa17418

Observation 9db5af05-c262-46bb-8acc-55079721d816 · inbound

Scaling Laws and Interpretability of Learning from Repeated Data cites this paper.

Scaling Laws and Interpretability of Learning from Repeated Data Training Compute-Optimal Large Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T15:52:40.515825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T15:52:40.335080Z digest=sha256:87decb6e78410f3cf26d56bb3f53a6636f9313ebcefbe335c0d8745a3df3cdce

Observation 5d746373-c98f-49ae-9ccd-bde0f9bb1437 · inbound

Teaching Models to Express Their Uncertainty in Words cites this paper.

Teaching Models to Express Their Uncertainty in Words Training Compute-Optimal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:36:08.608477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T17:36:08.566658Z digest=sha256:9a5c8f21b871f024ee075972f7c6309fb224a6a995ff4eaacaef851cd9299f94

Observation 97d62675-bd67-4532-b373-062df090a1e8 · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Training Compute-Optimal Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:38:38.342977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:45f138446f1bbce9bddadeabd26246207fbea493dea7ce0f139571bc606051d1

Observation 7de39fc6-bd99-4eab-bd79-7b2d76bb6058 · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Training Compute-Optimal Large Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.020705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:c6fe44c71b49cebf7c0c766256d8f428ddcc216138398360c4a901d4de1662b5

Observation 46655cd7-ecda-42f1-ba82-91df5b74823e · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Training Compute-Optimal Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:a07dac36162449c08f96cd891c6808bbab696565a1d95c658eae6529e7abc782

Observation 9e36862b-b4b7-47e3-9794-7709d116ea52 · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle Training Compute-Optimal Large Language Models

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:40:41.917817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:7e47cb39e1ef197d40c4a0fabcc452f7479c3d35ce01974c8afbf74de8aa3b9b

Observation f691636c-c0f5-4cb4-aaee-70218c1e6cda · inbound

Atlas: Few-shot Learning with Retrieval Augmented Language Models cites this paper.

Atlas: Few-shot Learning with Retrieval Augmented Language Models Training Compute-Optimal Large Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:48:43.193411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T13:48:43.024120Z digest=sha256:8174deaf93ac8f694c08198bc54423f32df09bfaf311910f436aacfb93c59585

Observation 3fa06da8-67e6-462e-ab03-eec923b1017e · inbound

Atlas: Few-shot Learning with Retrieval Augmented Language Models cites this paper.

Atlas: Few-shot Learning with Retrieval Augmented Language Models Training Compute-Optimal Large Language Models

Reference 198

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:48:43.399942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T13:48:43.024120Z digest=sha256:31c06518ea910ee081328f47fd13b3e2572257787911d67b7f4f26f528c31c6f

Observation bd46aea6-f635-4995-b44e-618234880b36 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Training Compute-Optimal Large Language Models

Reference 139

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:35:36.095713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:5ac04d04cc9f059aabbcba0a13d750436946474dde3b991fc3ce2a5a93acc5f7

Observation facb64b0-41e1-457c-98ec-e9595fbc23a8 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Training Compute-Optimal Large Language Models

Reference 167

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:29:06.203954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:74e8e75e39ef421b7a9896e10f98f37db4c4b3c756eb86f95c9853a2d07ebd4b

Observation db24b1b4-0fcd-41f4-a370-2b0ba4560c7e · inbound

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them cites this paper.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.829606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:f6b705499b0445a1881ffa773a42a16f78c043d80e9ee591c6a91794d33a8a4b

Observation 7ecdc3f2-8d1f-4838-aa08-560d3eb87ad7 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Training Compute-Optimal Large Language Models

Reference 248

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.593102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:d450d7ae9d2fdbd91badd1d3eabfbd6bf0df5621da6c8a939646d6415e9aee8c

Observation f89287c7-824d-40a7-b916-58d09275a5f2 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Training Compute-Optimal Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:53:21.934761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:f3c4de7190d9b7ec53a764d942f9919ceb6816cea8dbf423899ea1bbc5927941

Observation b057a6b7-f00c-4452-b867-ba59bfccf88a · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Training Compute-Optimal Large Language Models

Reference 176

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:53:22.164130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:84422d94a82d60c58235ed05f95608dd9e021259158bd161a195066b58b9f636

Observation e7a157a1-fe20-4df7-8171-411e64e2c0b9 · inbound

Solving math word problems with process- and outcome-based feedback cites this paper.

Solving math word problems with process- and outcome-based feedback Training Compute-Optimal Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-24T11:14:23.102902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T11:10:40.864420Z digest=sha256:19389742dc1a23bb142305df4d9a48b3ebdda42a18e6bbe5cc82f5d06f4008df

Observation a9fbe08a-bec6-4fde-9d37-ff1daf047068 · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Training Compute-Optimal Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:09:12.992496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:e444af888e5af24fd2a8205e209d689ab1b5236cb32d87567010e8a434c3d530

Observation b907b3e8-0575-46df-807d-07c014390d37 · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Training Compute-Optimal Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:10:49.405653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:0404af5ea6e7aecf10482ec3b7e0c3bc0d5e4c71c9cb711cbf39a77526645566

Observation d6ca6895-0a4a-4767-bcc2-c6898a3ef7f7 · inbound

REPLUG: Retrieval-Augmented Black-Box Language Models cites this paper.

REPLUG: Retrieval-Augmented Black-Box Language Models Training Compute-Optimal Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:41:54.119942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T12:41:53.833754Z digest=sha256:ee7e32a8e74e2831dcbb3036f8b352bb5033f269f2493a792d9ad71b0f8ca511

Observation a2645ea3-9151-4172-b565-948c48458d92 · inbound

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning cites this paper.

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning Training Compute-Optimal Large Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:14:16.417815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T09:13:30.054153Z digest=sha256:53d62635d92af6dca3d95ffce27987deba5ced321f3fc9163a4fb3e602ed5fbe

Observation b27ea298-d4ef-492a-a789-7905ad802049 · inbound

Accelerating Large Language Model Decoding with Speculative Sampling cites this paper.

Accelerating Large Language Model Decoding with Speculative Sampling Training Compute-Optimal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:29:36.298332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T07:29:36.205026Z digest=sha256:00866ff4749209de1516951d4a5b414aabbda691b4ea2838b25fb24c1a044aa9

Observation c7f612a5-867d-43ee-8088-9cfae7f0d2fd · inbound

Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation cites this paper.

Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation Training Compute-Optimal Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:00:05.374302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T18:00:05.294144Z digest=sha256:58f2b15116a80d2c45818641edce11c9b5f9a90394da7d0b76f73e9f1af287f4

Observation b10c6990-fff9-46a3-a839-0c07a42045ec · inbound

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models cites this paper.

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models Training Compute-Optimal Large Language Models

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:11:21.831689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T15:11:21.589770Z digest=sha256:bbda7708caad74b2b75262286e0e27f34777e0648c866525a6499adfc51d69a4

Observation b9fcbaf5-5361-4f22-b31a-8a9fb04123ae · inbound

SemDeDup: Data-efficient learning at web-scale through semantic deduplication cites this paper.

SemDeDup: Data-efficient learning at web-scale through semantic deduplication Training Compute-Optimal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:43:30.936629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T02:43:30.851915Z digest=sha256:27d64ff224965814bcda2f16f1ebd7039214756936595b7552c19407d62ab6a7

Observation e0607f81-b746-4939-8f3e-86e75aac89e0 · inbound

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society cites this paper.

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society Training Compute-Optimal Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:40:53.778903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T01:40:53.351795Z digest=sha256:97249e86e8ef95d3c48e7739f93a7dadbc9ecf497f26cd05680116f04dbdc37c

Observation fa94eff9-bf58-4286-bfd0-e7c204c116c8 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Training Compute-Optimal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:46:39.705991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:0278578ff1e7db7a14496c9b309017bc64a048ab3c014e4fc5299275f1abed97

Observation c3a28a78-2b98-4d1c-beb6-d178e62256ad · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Training Compute-Optimal Large Language Models

Reference 125

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:45:17.981519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:d666b54b0dcb3ae02c18d91ef6630304dca5b11f55d0e0905e47570b220b5a88

Observation b37f237e-fe82-4f10-ac72-a9d7e079c571 · inbound

Segment Anything cites this paper.

Segment Anything Training Compute-Optimal Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:14:20.214356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T06:14:19.815357Z digest=sha256:0cf60db12445217b58096c9169af02037e363ff36b0b22f9ddd624bd24457681

Observation 4af3b7ae-2251-40ed-b1ce-d47a43a565c1 · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Training Compute-Optimal Large Language Models

Reference 113

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:46:56.773494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:e98db33c8964b01f89b286824b1c79973e0de25abb5ad12b7e2d52643f4013ab

Observation 6da0582a-384f-406c-b65b-f9bed07efe79 · inbound

DINOv2: Learning Robust Visual Features without Supervision cites this paper.

DINOv2: Learning Robust Visual Features without Supervision Training Compute-Optimal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:dfdf784cc9b6c286ce8ee0f4ffd97ba0bb641857ba30f9119113116751ede138

Observation 2e535302-3e64-4efe-95c1-86138b82378a · inbound

How Secure is Code Generated by ChatGPT? cites this paper.

How Secure is Code Generated by ChatGPT? Training Compute-Optimal Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T10:04:18.968874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T10:02:22.456932Z digest=sha256:888ee490c19687d532a3c4358fa16b826c9ded8a70654c41146db146afdfaa46

Observation 3fad26bb-b1e8-48c9-86a5-a9c9c38bdf0e · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models Training Compute-Optimal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-10T20:37:01.852070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:0a3e31329689dd0dc97b7a120926341521ff0b7520babb36658e2ab597d09fae

Observation a1e2c194-0da6-42c0-85cb-3561303ece12 · inbound

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes cites this paper.

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes Training Compute-Optimal Large Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:50:09.426631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T20:50:09.265838Z digest=sha256:567ac4bbe02ea68fe126e5c069100b0c6b75fb339ebee87ba2c0ad4a44684a9b

Observation dba7ddbd-db91-49af-a6ad-9d90a26c2571 · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Training Compute-Optimal Large Language Models

Reference 199

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:33:00.893432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:a02c5a5a9b8234c29ed5f78394e85b0061b69c22ee56a8c1ca2c6b55435c9d49

Observation 857a1af4-2bb3-44f1-b81e-2944113ff07d · inbound

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? cites this paper.

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:36:55.142743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T07:36:55.087443Z digest=sha256:ea9de790a63f077578882a364e673d1a1590c39de15fc9ec6971b8816072bb76

Observation c1a50e12-f52b-4b81-83fd-812b4a4c07ff · inbound

Towards Expert-Level Medical Question Answering with Large Language Models cites this paper.

Towards Expert-Level Medical Question Answering with Large Language Models Training Compute-Optimal Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T04:32:33.397004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T04:32:33.271634Z digest=sha256:97062f5ef97f85fc95dd90dcd660cb8291051a5c0f5eb40f4ecc9dd6a1177342

Observation 85bc7124-dea5-42e9-b05e-e4d7e9bafd55 · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Training Compute-Optimal Large Language Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:59:27.164344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:903e0a46a82ed6cb1154b0f799ce9825cc4867715077818c5dfb43898f9700d6

Observation 3a58ab4c-492a-43da-92c6-7bf8d5a8c918 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Training Compute-Optimal Large Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:35:21.644705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:1f97455d2f85b377921e880bf801b240a1c3efbaf1493643e555d50eb0c11662

Observation 1d1c0c9d-8ad7-4973-9d17-b7251f69fe4c · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Training Compute-Optimal Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:43:45.848361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:1081ee7542b13d161aac9227cf719cb43cd04d48d2bfe56d3c7b11b3be1e5c85

Observation fca11763-1ee7-481e-8a5d-0a044294e676 · inbound

Objaverse-XL: A Universe of 10M+ 3D Objects cites this paper.

Objaverse-XL: A Universe of 10M+ 3D Objects Training Compute-Optimal Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:02:11.610496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T13:02:11.512409Z digest=sha256:eaef1c62ef319326a25a6c03b0846db7735aff59f7cb883e049b08d920193774

Observation 49b3f0d3-f2cf-46ae-9c6b-d6f1a001fc31 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Training Compute-Optimal Large Language Models

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.339978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:fb5f95dbeddfe55086016db81faa388c69e080e1bea9f40d9120d08f22f69709

Observation a7e7866a-25a7-452f-8f09-7270dc0f3be9 · inbound

C-Pack: Packed Resources For General Chinese Embeddings cites this paper.

C-Pack: Packed Resources For General Chinese Embeddings Training Compute-Optimal Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:24:32.209644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T13:24:32.084878Z digest=sha256:f45c92d3e5a9078bbbc9ea6bced73653e9bdc811d7c4fa5cbba9e1262dab9c4e

Observation 47e0fbcd-31fb-4c15-a0d2-d5d89badc8c7 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models Training Compute-Optimal Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.567927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:8cd4a1f08eab46d04bdfec355bdb33dcc8edb37ec03a655be61324f69d28b06e

Observation d4fa4b8b-6765-4c85-b423-171e5b585cbb · inbound

Language Modeling Is Compression cites this paper.

Language Modeling Is Compression Training Compute-Optimal Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:36:13.455055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T22:36:13.392250Z digest=sha256:c61f6b66fcd7887ec040ad130f889b4b5024594d30c320a5085486bf1e403b4a

Observation 3b16d98c-d430-4f1b-b779-5c1a9f45a305 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Training Compute-Optimal Large Language Models

Reference 150

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:12:31.599118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:8a51ff3314b76d062bf723b7ddc1652999a31e5889d76bd8033c504219092200

Observation df0020b1-7652-4c94-ad9d-d41d6bc9403d · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Training Compute-Optimal Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.316183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:bb0a7b0e0e683c63f14331343aa826e2cc50c9e8197c74d194cd2cb3b9fc0322

Observation 7b858d8b-11f9-4169-9382-e54d61608af6 · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Training Compute-Optimal Large Language Models

Reference 229

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T02:15:18.479900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:75ce136159568cb036f44e7df82b5a0303eff0baeba6bb8c238f2bfdae05175b

Observation 49d8ca2a-ebf9-405c-b1e6-7a8641dc9994 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:13:09.005969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:f59beac744a872432b5c30e06f3905fd683bd4485888de573c71165ba41f3d06

Observation f6b58561-b258-4259-b2ed-5e681fa55d3d · inbound

Llemma: An Open Language Model For Mathematics cites this paper.

Llemma: An Open Language Model For Mathematics Training Compute-Optimal Large Language Models

Reference 149

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:17:46.477044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T08:17:46.055279Z digest=sha256:80afe954ac10a13a556445f08f18fb9bde5f64cd05bfae889cfc13c1fc7d3f0f

Observation 1c6e46a8-f7b0-4b3d-9a34-47f45169629e · inbound

A decoder-only foundation model for time-series forecasting cites this paper.

A decoder-only foundation model for time-series forecasting Training Compute-Optimal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:07:21.307923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T18:07:21.246053Z digest=sha256:b8221a0acde121c182989e618d6e485c042ce7cdcbcdb580fc16ba9d0fab7573

Observation b76660d8-9d32-4a5a-8395-61918991be23 · inbound

BitNet: Scaling 1-bit Transformers for Large Language Models cites this paper.

BitNet: Scaling 1-bit Transformers for Large Language Models Training Compute-Optimal Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T05:43:56.559339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T05:41:15.544164Z digest=sha256:60d508bb7bd9592d6b5b3e09a6101124c3e38050639481be968437e01ef4bb02

Observation 3d964446-55f3-435c-9e3d-a8675e6846e2 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Training Compute-Optimal Large Language Models

Reference 292

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:46:10.131134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:c4017dc79cc01346e0ed8a63950127833fcae64f0192e705a6591831aa47a731

Observation 1642cbfc-a9b9-41e9-b731-157add39abe5 · inbound

A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA cites this paper.

A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA Training Compute-Optimal Large Language Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:05:50.603625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T22:05:50.453130Z digest=sha256:7e3ea92efa60f9e1cb2f064482109d3c4b0ca393252e99e90d4f013f8e3a96a6

Observation 84fc2006-19e4-4456-8f1d-a88b48ac71d3 · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Training Compute-Optimal Large Language Models

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T05:03:55.325041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:59b073f8e923d3103974b62febadf0780e0b2fc7953622fd3bfd16d3d2453f45

Observation b4df69e2-8710-478e-8ae2-a72a8df3a3e3 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Training Compute-Optimal Large Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.028539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:4c0f7743e4c7ab7652da4dd14e7857c211b047752191d0b25a0993228e854245

Observation 31e4c1fd-1109-4958-8f00-b131e8a797fe · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Training Compute-Optimal Large Language Models

Reference 144

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:08:05.926341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:69fac4c5c15b66189f8f701db98cf581f64c404aace3e238887b69352b93cbeb

Observation 576152b0-6f81-4d4b-97a0-89d797b42952 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Training Compute-Optimal Large Language Models

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:17:08.467578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:158b20e090ec17379c0bc6c8afbb902010639859c45a9893bbda82c232584a44

Observation 417b9257-a802-4c92-9b22-7dd90ed1baf7 · inbound

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models cites this paper.

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Training Compute-Optimal Large Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:50:09.661617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T22:50:06.399707Z digest=sha256:079279a50c7fc734f9aee33eb2ace5529427b7e896aa62819060d716fd60ef49

Observation d895236b-47cc-41f3-a560-a8ee720904e2 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Training Compute-Optimal Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:36:18.077733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:8d89b259a8e891b6638a44618528fd0783dec13976bec3968a691e282df7d9b9

Observation 37d157d4-34fb-4407-a4ed-aa71288e74a6 · inbound

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval cites this paper.

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval Training Compute-Optimal Large Language Models

Reference 138

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T13:07:16.448046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T13:07:16.151160Z digest=sha256:6d0fae73f74b0aa5d995d574121b5d3bd1c62a27f3c14ae1c9bd5f537e1c2def

Observation 7ae8a20e-a170-4570-9d03-5a72b6a02dfd · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey Training Compute-Optimal Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:22:54.480517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:3ebd9c8ffe514e6c343704f184ad7ecee21ceb49f05f5733fcf126ef7f00c98c

Observation 88e56c38-da17-4c57-9681-d19afd23902e · inbound

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models cites this paper.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Training Compute-Optimal Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.513427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:fd5878b6b93c3821469a4f9f47280db0a21bd21f16611c88f68ec5d76fa131d5

Observation 5f7487fd-3218-4f37-86bc-05556fdba3e9 · inbound

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts cites this paper.

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts Training Compute-Optimal Large Language Models

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T03:08:48.180582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T03:07:32.103513Z digest=sha256:4fb9725be867430807fae11cc96a6db70cf6984fe6825c0aa3ee1af908d0a933

Observation 31181b02-1b02-4ddc-a044-6156c5fb73e5 · inbound

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders cites this paper.

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:05:56.983966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T03:03:51.053556Z digest=sha256:d3345ade7188e2851e22b5a598bbacb3f9a08a8c4a2e3afd44b7ed0bb03ccd81

Observation 22adb4d1-4cce-4831-83bc-a2569fe6c6c4 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Training Compute-Optimal Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:47:27.856225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:79dedb3a18b42d58d75cff36818d60f1637e5b32ff307354f2c266f67082e43b

Observation 4c611a33-95e1-4524-befe-dfc59eac2804 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Training Compute-Optimal Large Language Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T18:00:53.423895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:de4dc87764f65d0c10b90814f2dd377d10b5f422cd159a201e3ed17b77579119

Observation 8650ba1a-f5b5-4608-968d-a33c0593582c · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone Training Compute-Optimal Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:19:27.344354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:d5d051e725208c83315b641ed7b534f4f0626c909e171ac024696692f6b9dcb7

Observation b902fd47-4ea1-4d87-a98f-204138b27bcc · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Training Compute-Optimal Large Language Models

Reference 139

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.635339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:1bbc69171fa11405c8025f3714e2e84ebae6b0d1516e92e91f65afdba9d0cc04

Observation d370648a-a0b7-4758-842f-b5984f2273ea · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis Training Compute-Optimal Large Language Models

Reference 147

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:03:56.499180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:b6e4599cfc0508e07917a5306def5517b516fb46e0d006d12221691c95f1e115

Observation fc1ecde2-0d2d-4a9b-b434-0781c3ccb0e7 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation Training Compute-Optimal Large Language Models

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:06.484041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:78a8ae074fd109c94f86d174bde5de7842b038a4d80b9b54244522bdd0fb4bd0

Observation 74cac238-ba98-4e31-bebb-2bed1a78721e · inbound

Scaling and evaluating sparse autoencoders cites this paper.

Scaling and evaluating sparse autoencoders Training Compute-Optimal Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:47:23.225720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T17:47:23.089288Z digest=sha256:d933abdc42ccba413512e4837869bd54832c24c4d501b24a977d4408e1a687e5

Observation cca4f1e8-62db-459b-8c09-95105b3d1d91 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Training Compute-Optimal Large Language Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T23:10:40.969748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:9952a8d7a1c08b711180830f97bd9c27a2444c65224d7173bf8a2418cf5bd23a

Observation 706efc08-1160-48bc-a18f-47494a64b170 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Training Compute-Optimal Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:09:16.900368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:4d43f48ed3d6836490b8b1cb7e887b369d06ec1336147c10b250d5d56b9d1335

Observation 4bc80752-f624-4bd0-9b7b-7b9bd7e6b4c3 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Training Compute-Optimal Large Language Models

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:58:17.005898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:6bb5a1001b3c96f818028f074425a8416dd813d714bf84d5aa435b9a545a30fc

Observation f7503109-87b3-4183-af31-c2bbbfb18611 · inbound

Retrieval-Augmented Generation for Natural Language Processing: A Survey cites this paper.

Retrieval-Augmented Generation for Natural Language Processing: A Survey Training Compute-Optimal Large Language Models

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T23:08:35.719443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T23:06:41.081461Z digest=sha256:d4264079684ae6a7160d3de69bd8e43dbf05ab94322d2c590a1900f33cd0032b

Observation 0c9edd13-1fe8-477a-ba4c-7433b7a29ae9 · inbound

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling cites this paper.

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling Training Compute-Optimal Large Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:42:23.571168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:42:23.297389Z digest=sha256:a8d3b4d022b31730d84d2693755cbaf4e2260488159cf5370b3e372e8d39c755

Observation b4d8f6dd-a4ad-4c80-b38a-a7c82448d183 · inbound

Gemma 2: Improving Open Language Models at a Practical Size cites this paper.

Gemma 2: Improving Open Language Models at a Practical Size Training Compute-Optimal Large Language Models

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T12:11:16.326752Z digest=sha256:0f0bf589b940d4c3d2d8a89893ac701634dac5ca1eef4c09ac66bf3dd9a0cee0

Observation 36a23d8a-381f-43fa-b873-1f7e453df19a · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Training Compute-Optimal Large Language Models

Reference 146

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:36.878637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:595eab09c5edaf7c5b71aa269494e05e41ac6047ea1f22b8acb6e09a1f18a897

Observation 6a91c791-6dcb-4108-9cd7-332b2482e400 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Training Compute-Optimal Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:07:31.561356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:2d84fe78a2bd65f0af8254cfe8c6540dd6a908430533c9e14728ace236470222

Observation 6672f655-71d9-4a21-9daa-bfff810d6d1f · inbound

Optimization Hyper-parameter Laws for Large Language Models cites this paper.

Optimization Hyper-parameter Laws for Large Language Models Training Compute-Optimal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.808999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:59506c00cba7b757e4745388788585510b820a534172bf9d82d630f31e5ebb83

Observation 8b1f3197-1f63-4abf-9e52-873a5ab2f18f · inbound

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio cites this paper.

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:43:25.300044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T20:42:38.782232Z digest=sha256:d117a0f0b3bb225da9d84c8f39ca84ffe5e0139e1febdc2cb910af57e97ee72d

Observation d048c94f-d0c4-4984-abc9-4f37295afb07 · inbound

Moshi: a speech-text foundation model for real-time dialogue cites this paper.

Moshi: a speech-text foundation model for real-time dialogue Training Compute-Optimal Large Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:13:22.130283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T08:13:21.962488Z digest=sha256:a787668972847fb9cd1bb1a2a9fdb5626e68f8bd028d4e6608187be5c87e289b

Observation d920897a-5533-4b8d-bb5a-9aea289eefde · inbound

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs cites this paper.

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs Training Compute-Optimal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:35:47.087940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T19:35:29.917096Z digest=sha256:b46fa8b0f142956790e7d844f103637d7305d31bd5aa40af8c5ea6dca4c1ed1d

Observation d1e52c40-edae-4719-8635-5ef299a8bc18 · inbound

Derivational Morphology Reveals Analogical Generalization in Large Language Models cites this paper.

Derivational Morphology Reveals Analogical Generalization in Large Language Models Training Compute-Optimal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T22:02:41.636083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T22:02:41.636083Z digest=sha256:ccb6ceac13bb80dbafed44ee133dad2b0ab07db6e126a79fb279c063ccd83123

Observation 1d27156f-b5e5-4e4e-b32e-9d2e1b96af54 · inbound

Safety case template for frontier AI: A cyber inability argument cites this paper.

Safety case template for frontier AI: A cyber inability argument Training Compute-Optimal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T22:01:11.365002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T22:01:11.365002Z digest=sha256:f7327b4d62c9d35719b9da53cc0a5eff85280e9da0999be58ba763517d3c1a7f

Observation 6e33ad3e-1564-41eb-b6c7-e96853034d81 · inbound

Sparse Upcycling: Inference Inefficient Finetuning cites this paper.

Sparse Upcycling: Inference Inefficient Finetuning Training Compute-Optimal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:22:22.010883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:22:22.010883Z digest=sha256:2de7a223000161363ed578b5827fadda7ace27c8bf10adc4c60ce3cdfbdff9a6

Observation eca3ac43-c4f2-47b4-8187-d818120db3e4 · inbound

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model cites this paper.

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model Training Compute-Optimal Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:14:04.895992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:14:04.895992Z digest=sha256:78c65cdd7002d87e0eeebb442db0ff074ed7477f9c77b0892920251fe135b655

Observation 9c46955e-b925-4e65-a53f-c7ee25ccb9ca · inbound

Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition cites this paper.

Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition Training Compute-Optimal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T20:49:48.165983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:49:48.165983Z digest=sha256:b5aca87d1ddda449776c82b3a86099610401b375db03e157fe535bdeefaa8c98

Observation 669e7edb-2385-4e0a-9f2a-360009bc108d · inbound

P$^2$ Law: Scaling Law for Post-Training After Model Pruning cites this paper.

P$^2$ Law: Scaling Law for Post-Training After Model Pruning Training Compute-Optimal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:55:21.732381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:55:21.732381Z digest=sha256:79477b2648e867e65e6c76e53a909732b6f35eade8c05a28b475604c0ac558a1

Observation 02ea0b3f-9bfe-434c-9ee3-cc6e9c44db9a · inbound

The Spatial Complexity of Optical Computing and How to Reduce It cites this paper.

The Spatial Complexity of Optical Computing and How to Reduce It Training Compute-Optimal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:44:49.064669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:44:49.064669Z digest=sha256:043516b0f1e18399eddf25c62e4e28ef4b8415738f08d0f421a5d3663ec4fabf

Observation 202efaaf-0d75-42a4-947f-20c2c779813d · inbound

How quantum computing can enhance biomarker discovery cites this paper.

How quantum computing can enhance biomarker discovery Training Compute-Optimal Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:45:07.141867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:45:07.141867Z digest=sha256:e38184b53808820ad24b82cbf5977e84a50e0b5939e2c2703269603f7a8524d5

Observation 005b9aa4-b0ca-4b00-a7ab-613b8dd1b82f · inbound

Scalable Autoregressive Monocular Depth Estimation cites this paper.

Scalable Autoregressive Monocular Depth Estimation Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:44:09.484370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:44:09.484370Z digest=sha256:9b658b1bb0e3097f1a89f991792fb02abb06806c8023112f785e2f8bf3849ac3

Observation 1620967c-09da-41a0-b344-83c135b367a6 · inbound

Data Pruning in Generative Diffusion Models cites this paper.

Data Pruning in Generative Diffusion Models Training Compute-Optimal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:30:09.045128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:30:09.045128Z digest=sha256:5fd332f4b7a1b155d6c7ca248bbe449f81ccb09cf133aa05ea5cfc83f70ab6cd

Observation bf45e740-35a5-4541-933d-85dff6fd3006 · inbound

Provable unlearning in topic modeling and downstream tasks cites this paper.

Provable unlearning in topic modeling and downstream tasks Training Compute-Optimal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:31:06.713376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:31:06.713376Z digest=sha256:31ee3dc70239a6aa8724413d6ed9e452e1679c9c9196cbbcf78140579debdc74

Observation 496f6329-48e3-42c8-8c26-ec0508bd2782 · inbound

Loss-to-Loss Prediction: Scaling Laws for All Datasets cites this paper.

Loss-to-Loss Prediction: Scaling Laws for All Datasets Training Compute-Optimal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:23.213770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:07:23.213770Z digest=sha256:0111616d151626fdc0fe1a1ca7d504fa7342a85664a16f18905da76210eef591

Observation 4f6d86ba-958d-4921-8dcf-c75dd51f8736 · inbound

Accelerated zero-order SGD under high-order smoothness and overparameterized regime cites this paper.

Accelerated zero-order SGD under high-order smoothness and overparameterized regime Training Compute-Optimal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:50:22.752797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:50:22.752797Z digest=sha256:be75f8b44ff1fc83bfbc7c8468673a1c8a81ad310f9062b09dc4a25cdade37c3