Pith. sign in

Paper Citation Record · LEDGER

Sparse Layers are Critical to Scaling Looped Language Models

As of 20 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 3 inbound Pith citation observations for arXiv:2605.09165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09165 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:26:37.559459Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:26:53.829565Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T17:08:44.006954Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact19
  • verified fuzzy4
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af9ae7f5-b4c6-43fd-9394-60f4ef6c3fd7 · outbound

This paper cites Universal Transformers.

Sparse Layers are Critical to Scaling Looped Language Models Universal Transformers

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.141390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:71a237009cd07ec0c316b83546f47394733780fd5938c9e0029f91771d6b7f91

Observation 08f630f4-af7e-4624-b4fa-8085d07a2458 · outbound

This paper cites Scaling Latent Reasoning via Looped Language Models.

Sparse Layers are Critical to Scaling Looped Language Models Scaling Latent Reasoning via Looped Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.143836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:4ea88f1bce51a2da1544884e307a1b237726d4fef930548fdd162f3972101942

Observation 6cddd360-fa1d-4001-8120-9b26c997f321 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Sparse Layers are Critical to Scaling Looped Language Models Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:35:29.150966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:62361697bd2bc7939c9c0b08faade74e6e4c2fe2bc3b31840301ffdd305d4aeb

Observation ebb975bf-74c0-48fa-b4e2-f7c2f1b49285 · outbound

This paper cites Scaling Laws for Neural Language Models.

Sparse Layers are Critical to Scaling Looped Language Models Scaling Laws for Neural Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:35:29.146101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:58506ca5e768f94f6953f9617fb9c9a080d55d6f7c8947469e0b35576e2c8cf7

Observation 96cdd294-7655-4f0a-a445-07e850fc2128 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Sparse Layers are Critical to Scaling Looped Language Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.148683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:f94af8647edac16ee63650c8d5d0e3a68671b47ab5e125dab4ed406dd38009ca

Observation fc17ab31-d334-406e-a644-6bbe10371ba7 · outbound

This paper cites MoEUT: Mixture-of-Experts Universal Transformers.

Sparse Layers are Critical to Scaling Looped Language Models MoEUT: Mixture-of-Experts Universal Transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T15:32:41.166003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:25871844702700de09e75a68a1b351fb9e3fdf28ad0a6aeb3688889e226976b3

Observation 30f10f13-7b88-4333-b5e3-bfc05eb0cfd9 · outbound

This paper cites an unresolved cited work.

Sparse Layers are Critical to Scaling Looped Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-07-06T15:32:41.175489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:618718ba40781ae897172693892e53cb27a6432c9bc7dbee367408f4661616f6

Observation 8f2f846e-f696-481b-9a95-f94bf96ddf75 · outbound

This paper cites URL http://ieeexplore.ieee.org/document/ 7900006/.

Sparse Layers are Critical to Scaling Looped Language Models URL http://ieeexplore.ieee.org/document/ 7900006/

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.527666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:8fb3556e5d0b8c8a8b9a4b4db7d07e0bdb2bcf170ddb006c5ec4b01ed2d07d4b

Observation 12411a4e-616a-47d4-8169-a10b05fa6253 · outbound

This paper cites DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference.

Sparse Layers are Critical to Scaling Looped Language Models DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.162165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:47b047096cfa7028082783c1a6f0c528985b9ba26a64141c1539a75eda391ee1

Observation 7b3c31ce-afb1-40e2-bb03-9e3114b67417 · outbound

This paper cites Confident Adaptive Language Modeling.

Sparse Layers are Critical to Scaling Looped Language Models Confident Adaptive Language Modeling

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:35:29.157001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:52d70444d5625bfaa7114c64f2e272a023f7a257cabc9bd64143850fa4d57f2b

Observation 340b20d4-83df-4fb3-a089-43013b248b20 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Sparse Layers are Critical to Scaling Looped Language Models RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.144506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:3c7f140ef7aa07fa2a1ce5b1b9f74712261cf7d42a558c1c10605864e32f043e

Observation 4f4bb4fd-9db1-453f-85dd-52893aceb77f · outbound

This paper cites GLU Variants Improve Transformer.

Sparse Layers are Critical to Scaling Looped Language Models GLU Variants Improve Transformer

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.149543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:97859677c9435c2d69be7dadea8f797040bdfdcc5a1de82483e4b496ca9778bc

Observation 7e856c73-559c-4b1d-bff5-27469d4ea621 · outbound

This paper cites Root Mean Square Layer Normalization.

Sparse Layers are Critical to Scaling Looped Language Models Root Mean Square Layer Normalization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T15:32:41.171792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:4d81d8e65194d3e14bbc7a934b518a2341306bb0f131d409861df12b9d443de7

Observation 252f9fdf-be00-47a8-a6d7-a4206bc4ecd0 · outbound

This paper cites Switch transformers: scaling to trillion parameter models with simple and efficient sparsity.J.

Sparse Layers are Critical to Scaling Looped Language Models Switch transformers: scaling to trillion parameter models with simple and efficient sparsity.J

Reference 14

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T07:35:29.131604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:64d6c1160dc2e85a0b6880010b0616d265ea6e3260225117a6f053917291928b

Observation 6ae657e2-4a9d-45a9-b8fa-3a8e84edd563 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models, April.

Sparse Layers are Critical to Scaling Looped Language Models ST-MoE: Designing Stable and Transferable Sparse Expert Models, April

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T15:32:41.168219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:3be879871b7a3194ada6b6ec1c7d910afa4643bc718cc18b7b192d5658357cb0

Observation a8545f5e-dea0-4647-b36b-3b57af46b590 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Sparse Layers are Critical to Scaling Looped Language Models ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:35:29.151888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:90d591c798366efc4d1700716ae21e9599c05fedf46bfcf388a8574d2e890784

Observation ace89dca-a025-4f73-a417-b3b1e17513b4 · outbound

This paper cites Mixtral of Experts.

Sparse Layers are Critical to Scaling Looped Language Models Mixtral of Experts

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.139849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:fbe3efc1904d9a5fba88f5f7f253c7ee2e46eeca3d9bc0dc0590d9966cd2228a

Observation 8e1563b3-a639-4f80-b3cb-96f1a24bc552 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Sparse Layers are Critical to Scaling Looped Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:35:29.134236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:936f57ace6a4bde50b1d7899be07003c8c9f0fb7b40ec821ececa08d107d80bd

Observation 8d7e4e47-4e11-4c1f-9fc7-d0c71b3b12b6 · outbound

This paper cites µ-parametrization for mixture of experts.

Sparse Layers are Critical to Scaling Looped Language Models µ-parametrization for mixture of experts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.137134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:464b484163d8fa4bfe3b35c2443e210e7f13ec0bab46e36ca6dec786fe2511b7

Observation 78d1846f-9335-4d9d-9343-6f08a72c761e · outbound

This paper cites Training Compute-Optimal Large Language Models.

Sparse Layers are Critical to Scaling Looped Language Models Training Compute-Optimal Large Language Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:35:29.154592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:e678f6095583af52c6decf05296ddf2ef4ec23075ad61570c4113ae4e9330d32

Observation 0c737458-8940-4103-a9d9-b1a1f75a123a · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Sparse Layers are Critical to Scaling Looped Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.164556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:7895f73e05de5d12c0f24c28cc8c388e633e32b299ba47650b9c2225c6ff1649

Observation 0a558ddc-639f-478e-a3e9-510344677a4e · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Sparse Layers are Critical to Scaling Looped Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.142171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:e12bf3cc626d6b0e384fc681f94a0d2cd71cf60cf0754d6bbb9c8ba2946ef045

Observation 912ed991-2177-4a6b-84d1-51791829ffd9 · outbound

This paper cites Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler.

Sparse Layers are Critical to Scaling Looped Language Models Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.147307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:626b39356fdee87e5b32c25f6689f4f5e9fe7af34510ac546767a90c6cad9d1c

Observation 19f4615f-ead4-43ca-8814-52ef30da456b · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Sparse Layers are Critical to Scaling Looped Language Models OLMES: A Standard for Language Model Evaluations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.159520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:219397a5324c4be9141558d4f74f37bb2a6946a784efbba8b322b013e8f37e79

Observation 1678ac11-6a11-4607-b776-7df0c13b453f · outbound

This paper cites an unresolved cited work.

Sparse Layers are Critical to Scaling Looped Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-07-06T15:32:41.170030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:23a3203e29bfc66418bf98c2df91bb2a65288443598d37576794be63118e6d4d

Observation 8ea638cf-0782-4ebb-b23c-db06c9603c45 · outbound

This paper cites interpreting GPT: the logit lens — LessWrong.

Sparse Layers are Critical to Scaling Looped Language Models interpreting GPT: the logit lens — LessWrong

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T15:32:41.179388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:27def277a52995df0216cf79152ac497cb7e99476832d15e49dce37ebe79c00f

Observation 66c4dcfe-c260-41f7-bad6-0cee2d519e0a · outbound

This paper cites an unresolved cited work.

Sparse Layers are Critical to Scaling Looped Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-07-06T15:32:41.177501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:eb37fdc2e6392558e31577ac297a9f8c8fda8c7cb3a7e2ebca6b00be7f8a3320

Observation 012e7808-59e0-460c-b772-3acd216a5213 · outbound

This paper cites Approximating Two-Layer Feedforward Networks for Efficient Transformers.

Sparse Layers are Critical to Scaling Looped Language Models Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.121225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:f6b75ef0cf30aae9f1eb7a7ea16346d93b46e2bf89452df838b413d7323d61f2

Observation 82b6ef2c-655b-4686-ac88-54e268505ad6 · outbound

This paper cites SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention.

Sparse Layers are Critical to Scaling Looped Language Models SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.123697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:c9bac2cdb12eb6357fa9fee55f87125e0b3d3d2e72c31480cf7bd5be33fb9d1f

Observation 16c6fe74-1d7f-4843-a6e4-2b631db4cdb7 · outbound

This paper cites Layer Normalization.

Sparse Layers are Critical to Scaling Looped Language Models Layer Normalization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.126150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:e79b1bbea385c3448e32e6bd4c4504f3d13d4d26bf61349796939c3e12a9db6c

Observation 19915ffc-48c3-483d-833a-1635d2273075 · outbound

This paper cites L ayer S kip: Enabling Early Exit Inference and Self-Speculative Decoding.

Sparse Layers are Critical to Scaling Looped Language Models L ayer S kip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 31

Resolution
metadata mismatch
doi, observed 2026-07-01T07:35:28.530348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:eecf21ff461e8076721f61bc9fb244ffa59c3d6193af66dedce3165cc63814fd

Observation 1de4db74-4a21-4dd7-b6b2-3b4b7cb1b546 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Sparse Layers are Critical to Scaling Looped Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:29.113057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:a7b44fc1c62c0291d7166b6c045aa60a802db3b4b6633b44ecb7629d720381bd

Observation 59aaef27-6760-4cdf-a2ec-0da6d6763c27 · outbound

This paper cites Mixture-of-recursions: Learning dynamic recur- sive depths for adaptive token-level computation.arXiv preprint arXiv:2507.10524.

Sparse Layers are Critical to Scaling Looped Language Models Mixture-of-recursions: Learning dynamic recur- sive depths for adaptive token-level computation.arXiv preprint arXiv:2507.10524

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.115807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:55d3c8cb9f97333a8d09f46a2f355f0c6fc88774745936b7c4e687b299f5f238

Observation f9b073a9-3594-4bd7-a615-fd79224e0d35 · outbound

This paper cites Don’t be lazy: Completep enables compute-efficient deep transformers.

Sparse Layers are Critical to Scaling Looped Language Models Don’t be lazy: Completep enables compute-efficient deep transformers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.118338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:4bc0ba68d2aa89a4c93e905af0860ac4b6ba0ca70f0d7fe46a3be1d42ce81636

Pith citing papers

Observation 956782ee-58a1-44a1-bbe2-b9b7db90e466 · inbound

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models cites this paper.

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models Sparse Layers are Critical to Scaling Looped Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:08:44.008245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T04:27:48.590008Z digest=sha256:25d149590699aaee38b892e34d26f164ca296c3696ab1af596c5338c7a8c3906

Observation 3f0ffe69-439e-470a-86ee-8a5bbddbc55f · inbound

Loop the Loopies! cites this paper.

Loop the Loopies! Sparse Layers are Critical to Scaling Looped Language Models

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-01T21:34:28.320684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:34:28.320684Z digest=sha256:243e76e7a19f617518711fef03c573705a4d4a2340a027ed7666a34cf303803e

Observation 5d5002f9-3e51-4bb0-8e9d-c5940efec03f · inbound

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching cites this paper.

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching Sparse Layers are Critical to Scaling Looped Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:26:53.829565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:26:53.829565Z digest=sha256:01366167d9057fbe13454117272efb8dc690b82601aec5c86eb453ca99583a12