Pith. sign in

Paper Citation Record · LEDGER

Will we run out of data? Limits of LLM scaling based on human-generated data

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 77 inbound Pith citation observations for arXiv:2211.04325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.04325 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 77 of 77 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:53.570304Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

80
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5a6f788a-a0a6-4c36-8a87-1ebe2a3ecebe · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.424024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:170ab14dfc3bc5b9569350e340905765d45718359a17a1ac052a1c47615bd2cd

Observation 831c9e31-12e6-4466-a507-8f2a0986a6dd · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:45.887160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:9f07493b7688acfaa2803041a7ae2e9c78889872a2d1385d3cbd074ddf56352d

Observation f2b47ae6-3139-42ea-93f1-0e2283217e37 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 149

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:10.063514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:04f5da5ab13befa29409ea033a94bad0d09438bb77236a397d277f57bffdf1e4

Observation a2753f32-9be5-4eec-aef0-0b697da001d6 · inbound

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts cites this paper.

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:08:48.125057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-24T03:07:32.103513Z digest=sha256:5416050f4cb9bebaf7543e2f86efcacd10e84e73ad29a005f604842f243dc03d

Observation d6f10483-5db1-49f1-9f04-dc38322d68b6 · inbound

Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees cites this paper.

Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:32:25.049137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:32:25.049137Z digest=sha256:b3e9606cc0365785bcd95b82d33bc50a4f7c7774288e660c18c0fb356ca991fc

Observation 3f8817db-60fb-4152-8a6c-ff7825881fc9 · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:22:44.236980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:b24eaff8da439dba806c61587c7fd904615957979a0f54c01bf20f886ad47ad7

Observation 848b8176-fe63-4f73-a976-b0bed82d3e17 · inbound

Towards Data Governance of Frontier AI Models cites this paper.

Towards Data Governance of Frontier AI Models Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:54.256471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:54.256471Z digest=sha256:3caf9fd1cf56c66a4210539e3b6a601f4fe84b58122e3dd8804de527bc54be79

Observation 262a8a64-92e3-4621-8870-a4f8f5e65940 · inbound

A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities cites this paper.

A Federated Approach to Few-Shot Hate Speech Detection for Marginalized Communities Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:13:00.810802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:13:00.810802Z digest=sha256:e16dec7a880c0f7860e27fa8f7874802b4571e0b94cc7fe7e0a827449a173309

Observation 9e514f58-4e1f-46cd-a5d9-b580e990ceaa · inbound

Interactions Between Artificial Intelligence and Digital Public Infrastructure: Concepts, Benefits, and Challenges cites this paper.

Interactions Between Artificial Intelligence and Digital Public Infrastructure: Concepts, Benefits, and Challenges Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:15.876889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:26:15.876889Z digest=sha256:552271e7d234f660643c1fa36a4436092b77d5093630c6190e6b046199a3baf2

Observation 5903fae7-375d-45c2-93a3-53bde9a10cc1 · inbound

Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali cites this paper.

Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:10.127865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:46:10.127865Z digest=sha256:2582c4d8e9ea2661c42bdc10126f6ca7a6ff6921ccc536407d4c35fa9e5ba20b

Observation 96b4d63a-fe9f-419c-b244-d6e4497f3e7a · inbound

Language Models as Continuous Self-Evolving Data Engineers cites this paper.

Language Models as Continuous Self-Evolving Data Engineers Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:52.570456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:52.570456Z digest=sha256:f4b0ed6f3ed1e2c8f8d4b40094e095e4486581dd86e905c667aa06efab3221cf

Observation 1d68e051-9436-4202-af41-b71856f84b7b · inbound

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling cites this paper.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.027916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.027916Z digest=sha256:d786c8a879e4ea16bed1da64c958d871661a682943420979f900210ab60475a1

Observation 56457b2f-84b0-431a-9696-9d0fdfc5f58b · inbound

Exploring the sustainable scaling of AI dilemma: A projective study of corporations' AI environmental impacts cites this paper.

Exploring the sustainable scaling of AI dilemma: A projective study of corporations' AI environmental impacts Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:20:32.602995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:20:32.602995Z digest=sha256:db4feb0335069de2b9594a2d65399577667d02544b749a8a200cbb2ea586a1c4

Observation 842e2fa5-f0fd-4b7f-8c2c-bfcb8d5a9be8 · inbound

State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence cites this paper.

State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T23:52:06.037131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:52:06.037131Z digest=sha256:0ab3536cea23ed41923773a0f9de28d2c38ffc90ad65ac8a25905ee08988e2e1

Observation 4731627e-7fa4-4e11-a8e2-dc5e07e6b0df · inbound

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning cites this paper.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.337684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.337684Z digest=sha256:08fd707a6a99be2db7d3655d6981a3ac21a3e10c1ef9b7fa3f0b363fd7bea737

Observation 99900fa9-792c-4c23-ad14-0f4202cec919 · inbound

BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation cites this paper.

BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:09:15.405937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:09:15.405937Z digest=sha256:4acb412b3a06d2ba8e8ed9d643d19cfa5185353ecdc98710e608d52a034ad145

Observation 73549ae8-b156-4366-b838-528b5ee46227 · inbound

Toward Neurosymbolic Program Comprehension cites this paper.

Toward Neurosymbolic Program Comprehension Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:26:32.074862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:26:32.074862Z digest=sha256:590136baed5eda783a8ed74187b3a510ba95d0cd1f09997d1c09e83cd62ba05a

Observation 68c1292b-d7b5-4ff2-ae81-2af432f37b84 · inbound

Reformulation for Pretraining Data Augmentation cites this paper.

Reformulation for Pretraining Data Augmentation Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T23:08:35.521790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:08:35.521790Z digest=sha256:e1088e9ca93e7c00c13d4cef473df5d833647cbf5e2aa131c67f2b8d916ad00b

Observation c7ca237a-e0d2-4117-a607-fc3f3d0c8afa · inbound

Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks cites this paper.

Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T03:55:21.898749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T03:53:00.509414Z digest=sha256:8298d2b9be7cfc28f6be77f40f0aa4e2417b0f95bc9b817d432cb45a3272b709

Observation e5315d0c-680f-47d4-9503-49ccc0ac289d · inbound

Carbon- and Precedence-Aware Scheduling for Data Processing Clusters cites this paper.

Carbon- and Precedence-Aware Scheduling for Data Processing Clusters Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T20:51:22.558001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:51:22.558001Z digest=sha256:a56cd5ea44903c0e2c0022527a104e0a59a2917e634f85a27cbd5f1d1fd56f97

Observation 01efc45d-5d01-406b-948e-7755469047e7 · inbound

Governing AI Beyond the Pretraining Frontier cites this paper.

Governing AI Beyond the Pretraining Frontier Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:46.888906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:46.888906Z digest=sha256:6033006981a2e971d2407670d1cbaa18ac331ab85abe2cb3bde4b1be04cdff18

Observation 1ec157f3-0afd-4d78-a41f-0421b661ea15 · inbound

Will the Technological Singularity Come Soon? Modeling the Dynamics of Artificial Intelligence Development via Multi-Logistic Growth Process cites this paper.

Will the Technological Singularity Come Soon? Modeling the Dynamics of Artificial Intelligence Development via Multi-Logistic Growth Process Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:19.296369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:19.296369Z digest=sha256:b21e86e71eac966ed45c35fb625c067eb207de49132202bafe803ca0ae42de63

Observation a01342c8-52f5-49ce-9c6e-f7f9a1e75a02 · inbound

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices cites this paper.

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:05:16.560310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T01:03:26.037233Z digest=sha256:d202b8b5696b09694bb73de7b417696ac6179cd605b8249d4095ae836d8200dd

Observation cbf0bcf5-7040-43f4-ae36-73e83a968d80 · inbound

Position: The Most Expensive Part of an LLM should be its Training Data cites this paper.

Position: The Most Expensive Part of an LLM should be its Training Data Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.570304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.570304Z digest=sha256:6e526291f1a63434f7d716f770766991f411d9c92d3138e4403e00e6f59fa379

Observation 3080e3a3-abdd-42a7-bdea-6fe2c804baa1 · inbound

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training cites this paper.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.138230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.138230Z digest=sha256:a7887e7d977cb70454640be277463a6dd65b842011b02edb0cb2f606e6e6619f

Observation 4d011f77-b867-47aa-9bc5-21ab0a5ae992 · inbound

Dargana: fine-tuning EarthPT for dynamic tree canopy mapping from space cites this paper.

Dargana: fine-tuning EarthPT for dynamic tree canopy mapping from space Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:20.289891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:20.289891Z digest=sha256:904b5f73f081e1e2cb0d8997e357cc0d358175a73b577083d70eec2c6dde2e9e

Observation 14a84937-3fdd-4c8f-9cc0-ece3082d4d28 · inbound

Position: Enough of Scaling LLMs! Lets Focus on Downscaling cites this paper.

Position: Enough of Scaling LLMs! Lets Focus on Downscaling Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:14.117855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:14.117855Z digest=sha256:c62d2a3f968713432ca1d9476cbbdcb5ac17b5b2b5e2858ac92d7788c4e3dd36

Observation e6800c57-0937-404e-a74a-5918aac18046 · inbound

A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm cites this paper.

A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:55.898770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:59:55.898770Z digest=sha256:e408f441ccde4072110a3859262db7df4062d09bd73caf394a57c5d36a775663

Observation 8388d6d6-7ee7-410a-8e49-6e64b2408816 · inbound

Incentivizing Inclusive Contributions in Model Sharing Markets cites this paper.

Incentivizing Inclusive Contributions in Model Sharing Markets Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:58.571255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:58.571255Z digest=sha256:bc1ba2ffd463f221155b5b5c2ed4d3fd4f08a66f2f67f170b5982953b3f338ce

Observation 82466082-51e3-412c-8899-ae44c2b3692c · inbound

Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research cites this paper.

Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T22:31:28.861620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:31:28.861620Z digest=sha256:1a3634c0e8c83d16d61e64b111338f8e227f595b0d2aa65e34bec08ff6e4a678

Observation f065f91d-2f0f-48d1-840d-ce4b007bf996 · inbound

Parallel Scaling Law for Language Models cites this paper.

Parallel Scaling Law for Language Models Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:45.851502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:14:45.851502Z digest=sha256:f34db8144323deeff17f1cd6bb05754a1b438fddcc9756e262f80714e9ec9f04

Observation 53ce9448-2e7b-4bb0-a398-68f2c5612986 · inbound

Addressing memory bandwidth scalability in vector processors for streaming applications cites this paper.

Addressing memory bandwidth scalability in vector processors for streaming applications Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:39.292182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:39.292182Z digest=sha256:4d10275d5d9a742cdd4182edb291889cbfd07b9bea2a7a1da4f23337fbbd94d3

Observation bf8115c0-2833-498c-a89a-5b6722fdf237 · inbound

Granary: Speech Recognition and Translation Dataset in 25 European Languages cites this paper.

Granary: Speech Recognition and Translation Dataset in 25 European Languages Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:19:15.038184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:19:15.038184Z digest=sha256:0171aeb56aa6b3eadc2a5c56288450360840b8a2dfac26591b1e3e6c683508f9

Observation e03034ff-c716-4d16-b365-62a4cea802f2 · inbound

Multilingual Test-Time Scaling via Initial Thought Transfer cites this paper.

Multilingual Test-Time Scaling via Initial Thought Transfer Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:23.150570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:23.150570Z digest=sha256:089f45f6c86eaaf1acdc0af66003ae6e04371935334d5191f2cfbe4398b115e6

Observation ad7abbee-0a8e-40b3-9e02-0f15c998f13e · inbound

Transfer of Structural Knowledge from Synthetic Languages cites this paper.

Transfer of Structural Knowledge from Synthetic Languages Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.537389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:15:46.537389Z digest=sha256:0ff0309b6cc137e69afdf3b8f63900bd303609b3073ddba2a8ed620b974d4273

Observation 4612b370-df80-40a7-9d47-d1e2a4e3af3a · inbound

Dimension-adapted Momentum Outscales SGD cites this paper.

Dimension-adapted Momentum Outscales SGD Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:14.992638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:14.992638Z digest=sha256:df57140725ee61a57c5f8241b505400003ee8d4627bd6e606a5f112df3c4e8bd

Observation 14bb25b9-417e-4c09-80dd-5ba719829399 · inbound

Towards Anonymous Neural Network Inference cites this paper.

Towards Anonymous Neural Network Inference Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:48.589015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:35:48.589015Z digest=sha256:1355c9fd8aca30332a173ce281987c3746a10ec10f49f9b0d3d56df796b37637

Observation fbcb6a76-e838-4dc0-ad3a-c6b7899f51fc · inbound

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence cites this paper.

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:08.046459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:08.046459Z digest=sha256:0d8b3635d6d7891ec3a10e63a22461215417384f3f6d3902eab9a541c9042591

Observation 689319c3-678d-4da4-a70a-b982d24d5ee7 · inbound

Bayesian Inverse Physics for Neuro-Symbolic Robot Learning cites this paper.

Bayesian Inverse Physics for Neuro-Symbolic Robot Learning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:53.374794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:53.374794Z digest=sha256:08497ef151aac1fd04e1fb70578ddde32e4ff47d76b7f6c8a87a2c81e58a9cce

Observation 509c0d6b-3565-443d-ae07-842750376b53 · inbound

The Synthetic Mirror -- Synthetic Data at the Age of Agentic AI cites this paper.

The Synthetic Mirror -- Synthetic Data at the Age of Agentic AI Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:47:36.431460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:47:36.431460Z digest=sha256:6ed8e0ef5bb2b7f8c45d1db2df1077c3a5e6a3f2875e9f9fdbe869d0beb877a8

Observation 78185f72-cefe-4955-b62b-d7dafd7db555 · inbound

LLM Web Dynamics: Tracing Model Collapse in a Network of LLMs cites this paper.

LLM Web Dynamics: Tracing Model Collapse in a Network of LLMs Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:49.328946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:59:49.328946Z digest=sha256:907a7cd7c4369739b092ce3bc6c307b5356a91692146443470bdca8cfbaf2e30

Observation 565350d4-04dd-4b8b-a364-fc0f17141ba1 · inbound

A Comparative Study of Open-Source Libraries for Synthetic Tabular Data Generation: SDV vs. SynthCity cites this paper.

A Comparative Study of Open-Source Libraries for Synthetic Tabular Data Generation: SDV vs. SynthCity Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:03:27.375820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:03:27.375820Z digest=sha256:eb33fb0f62ada9eaa2fbece3e197df8492211c73d4d4746fa7215969feeeada3

Observation 8d551cc3-3d97-40c4-9a23-aa876ffa6622 · inbound

Lost in Retraining: Roaming the Parameter Space of Exponential Families Under Closed-Loop Learning cites this paper.

Lost in Retraining: Roaming the Parameter Space of Exponential Families Under Closed-Loop Learning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:57:07.185572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:57:07.185572Z digest=sha256:8a278c0520f2f03a03859dc086be3ae2a1b29dcc6a33558ab6cd6db6e1da62e5

Observation edcf4a8a-367e-4fc7-87cf-cc2df63f800d · inbound

Hierarchical Reasoning Model cites this paper.

Hierarchical Reasoning Model Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:58:03.457295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T04:58:03.283911Z digest=sha256:c9b5030054b58dddc6efcd4fe3f53e41aeaddad8d7ec3bf3450b4ea3e441c631

Observation ea500b1a-5515-4f4a-935e-322268ff2939 · inbound

Residual Matrix Transformers: Scaling the Size of the Residual Stream cites this paper.

Residual Matrix Transformers: Scaling the Size of the Residual Stream Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:42.018922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:10:42.018922Z digest=sha256:d8dbafad699d8c83a4147021d7be909e40cfedff12a2301c0719d42c9870df04

Observation 9abf2387-9409-4fbc-90b8-c5f0f70d8c5f · inbound

Energy-Based Transformers are Scalable Learners and Thinkers cites this paper.

Energy-Based Transformers are Scalable Learners and Thinkers Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:37.584623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:37.584623Z digest=sha256:ae4a7e4707a56aa4c16069987d0c3ce3968010c5ff9eeb13dee5d0fb9ca52401

Observation d31d45bc-27ea-4193-bbe7-042c3d85530c · inbound

Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models cites this paper.

Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:53:35.442569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:53:35.442569Z digest=sha256:87511804d1a08be3ef72d31c31067cc43e6f847c96b2955385038b45a3a039f4

Observation d0faab4c-2961-44a0-84f7-399bba3419db · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:40.536802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:40.536802Z digest=sha256:170ebbc5b05b379f4e7660f75cef17c4e7d1230fb6de3d9c5b783561c9ee2abb

Observation 5468e308-bcc6-4515-8258-05aa1c0a1120 · inbound

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need cites this paper.

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:43.060987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:43.060987Z digest=sha256:e04f949e6f1e63b72e7d0a7d6cf44a6f6adab01ae9c0d764d65202bf22f5f978

Observation 295e86fa-ded3-4644-bd65-a085b6d81857 · inbound

HuggingGraph: Understanding the Supply Chain of LLM Ecosystem cites this paper.

HuggingGraph: Understanding the Supply Chain of LLM Ecosystem Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:29:11.630290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:29:11.630290Z digest=sha256:73b4e7064778484ecb96085eeb4528992fc6560ba7617e1f296bdb3ff61756ec

Observation f0a212a6-4dde-4bc2-b561-d028cfc6cc4e · inbound

Meta CLIP 2: A Worldwide Scaling Recipe cites this paper.

Meta CLIP 2: A Worldwide Scaling Recipe Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.444125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.444125Z digest=sha256:9ad70ef57e793c70b794647ff97405cedc529f072f1aaf3884bb9d2498dd659b

Observation 5415c29a-ec84-4c8b-a45f-c5dfb6d20f83 · inbound

LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points cites this paper.

LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:48:28.385076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:48:28.385076Z digest=sha256:96224b77c04caa1d3ffc9a0228bf17f3289d413e2997fed72bf3528aba8ee9ac

Observation 7fbcc309-8d5e-43df-b604-0813db2ba0a9 · inbound

Learning Facts at Scale with Active Reading cites this paper.

Learning Facts at Scale with Active Reading Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:05:45.242496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:05:45.242496Z digest=sha256:fd5a54295cd98fc5041242ffdd675c4441602b5dc9dc9687110092e1a4b6f741

Observation 997fcac4-34ca-4ce7-bbc4-77fa54a4521f · inbound

HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents cites this paper.

HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T18:24:29.213666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:24:29.213666Z digest=sha256:88447a0abd2ef3ef0c4e7220daf4321ac2ebcd52c8040fd63105f8e8437e3fca

Observation 128ab968-9b7b-4d12-af2c-da3479e348b1 · inbound

User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies cites this paper.

User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:20.668234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T06:00:20.668234Z digest=sha256:ddd4fdc50984685d73b52624c740c104f14167d4f2aacb0aeea51ce411d0c610

Observation 5cae4786-ccbd-4b9e-869e-a917d914b6ae · inbound

Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution cites this paper.

Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:42:45.708509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:42:45.708509Z digest=sha256:4aa21a4f0a8c045b74ed6f06c7635464a87de08058122b4bd3f50e3b01b6805d

Observation dd5843fe-ecff-4e71-8c42-ee35b2df103a · inbound

Generative Data Refinement: Just Ask for Better Data cites this paper.

Generative Data Refinement: Just Ask for Better Data Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T20:24:40.771539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:24:40.771539Z digest=sha256:b5c4b3655b11a689133dcd0e19d9cc2e92033c017dd1d28d1e238b5dfd01bb69

Observation c4052d0f-c2d8-4173-bdaa-5c533ccb9e66 · inbound

Differentially-private text generation degrades output language quality cites this paper.

Differentially-private text generation degrades output language quality Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T17:03:13.329279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:03:13.329279Z digest=sha256:d125d81b44dc2469a0e87fd9161a62649f4dc1bf4dfb0a9b0627089a15a6779c

Observation 97c729b3-34fd-4d99-9081-b058935585ac · inbound

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency cites this paper.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.571487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.571487Z digest=sha256:2222950183e1fc2b629084ed71cfb68817da1b2244d91364ce2b7f4303404d73

Observation ffbb22eb-bf1b-46a4-9813-0f61801d8f95 · inbound

SenseAI: A Human-in-the-Loop Dataset for RLHF-Aligned Financial Sentiment Reasoning cites this paper.

SenseAI: A Human-in-the-Loop Dataset for RLHF-Aligned Financial Sentiment Reasoning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:50.338012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:42:18.737105Z digest=sha256:a41d8cef745b99db3bf6de25040d3225e25825a0bbe38e9e9c6122c73ccba717

Observation 1c452130-863e-4fb8-a6f4-a9b79ccea0bd · inbound

FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion cites this paper.

FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:06.138293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T02:35:40.593397Z digest=sha256:edbab1dd480feeaa553f1eb6a64ad2c1b60d3b8772efdfb99f509e2ac8afc440

Observation 9d682923-555e-4c88-88bb-c7f7f720d810 · inbound

AI Hastens Limits to Exponential Growth cites this paper.

AI Hastens Limits to Exponential Growth Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:13.249821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T09:16:54.056531Z digest=sha256:043915da6b49a855c042a7c41ccb4ac32eb755e4b8f32b3937780d55a35e681d

Observation cc489784-f5ab-4d48-9384-289a812752d2 · inbound

LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning cites this paper.

LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:22.448434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T18:44:05.775332Z digest=sha256:abd1525f238e02b115ff10f0b9cbab1e3193240db5816cdde9f009f71c00c854

Observation 513ba88d-9229-4d55-afb3-7dd8ed8c1bce · inbound

Submodular Ground-Set Pruning: Monotone Tightness and a Non-Monotone Separation cites this paper.

Submodular Ground-Set Pruning: Monotone Tightness and a Non-Monotone Separation Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T17:36:04.791771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:26:30.268643Z digest=sha256:06da49d16c9a65a0094f1f34742454ad678d5029808fdb2298dff2a37af91597

Observation 128389b6-8509-4ce8-a49d-dd509c92dca8 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:06.631428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T15:35:08.202464Z digest=sha256:edc6b0361bc2217f31bc506a13e1c01217fcb66c957e86a0a527c03c1b5ac9e1

Observation e83ca845-5e91-4fa9-9270-2aebd77450fd · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.984390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T23:25:55.375541Z digest=sha256:931c45c93b49753df86a373d2deb86d0af8ab2c6232254b592bca3498fbf2524

Observation 062cdf99-ddad-49ad-ba21-711e408a37c0 · inbound

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations cites this paper.

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.477138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:51:28.068087Z digest=sha256:c0dae873f031b99d647e29ba5a03ad4672407c419d212c4d9d8bf49a54d6f7f9

Observation 9c139432-a16c-4192-a058-7135c6ff2a80 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:54.418300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T02:30:14.693348Z digest=sha256:f20b10091f94a1aadd161e1c95f692dde777466a9a246cc768752633fbd199a6

Observation d814c45a-47fa-4772-96d6-058b6c4d1100 · inbound

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production cites this paper.

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:19.456297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:17:24.147248Z digest=sha256:a506a59315958c71d6cf8fa38f3df9b8959c6cf4d99c7fdaa9f4d79b5b19f83d

Observation e5f6a51d-979a-4857-af91-545854798ccf · inbound

Beyond Scaling: Agents Are Heading to the Edge cites this paper.

Beyond Scaling: Agents Are Heading to the Edge Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:53:14.918737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:51:14.999341Z digest=sha256:64fb00e32da07310d3b630a11da6e946f83a6a29323376c3f6154bbdd1319092

Observation cf2ae034-3a0b-4681-90bc-b36f63c2fda9 · inbound

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training cites this paper.

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:25.188002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:16:48.536850Z digest=sha256:2992711ab1c2df634dabd1b81aa532182eef2716d703669711b81ea45aa6634b

Observation c19a4932-413e-4e86-8f2e-1119549dad61 · inbound

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws cites this paper.

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:09.820598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T22:24:01.067388Z digest=sha256:c4dfc1b8b01699b596fbec1f8f1ddbaa1a7a60c9abb13c2bfcad7518fe2428ac

Observation 30d223e5-6b82-48a5-9b0d-44959168c99b · inbound

Data-Driven Automation cites this paper.

Data-Driven Automation Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.200386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:58:40.370152Z digest=sha256:3c7e5afcccb360260709d554863612ed68f38d01907d4b91320aa6bfc1e0d33d

Observation 9c6b9f74-2cf2-475d-a001-39fbe79f505c · inbound

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering cites this paper.

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:55:06.335237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T22:54:03.054871Z digest=sha256:961c30f40713ed2cf4b74dd1da9dfcb84123c0b5e944cd433aa01c801025ebb8

Observation 217dc43e-c9da-44b1-92a5-a52fff2b9e56 · inbound

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining cites this paper.

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:44.254033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T03:55:42.948380Z digest=sha256:35ea662cfd1d63c733779280a61936df0f0389af2c8d10fa6506f9482a3661c5

Observation 68a1ac45-c9db-4e92-bc33-7527ec0af216 · inbound

What Drives Interactive Improvement from Feedback? cites this paper.

What Drives Interactive Improvement from Feedback? Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:05:43.468362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T02:36:20.649687Z digest=sha256:09a5362737d23a0b2fbc347e5131c3342e1b8d9556e009753db18489054ccab9

Observation 8d06a739-dfc6-4a3a-8f19-c728b4707174 · inbound

LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning cites this paper.

LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T09:01:49.796306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:01:49.796306Z digest=sha256:febf96e00c6cce5f3154923bfac8c7832b75c96ebca82a901ca84160c5f1981c