Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:14:20.461704Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.22389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:14:20.461704Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0975920-8ccb-4fc5-bfeb-2264c07accdc · outbound
Towards Distributed Neural Architectures GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2482fa62-7024-41dd-992c-ebe0d7109451 · outbound
Towards Distributed Neural Architectures Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280f67af-b0b0-4421-a7a3-b66c0271f621 · outbound
Towards Distributed Neural Architectures Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca5aea9-79ea-4d95-a551-6c355c8775e5 · outbound
Towards Distributed Neural Architectures This may be because we are not considering a setting with high sparsity
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16ee2b9c-a491-44bb-b343-14df6bd663b4 · outbound
Towards Distributed Neural Architectures DeepSeek-V3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37c74e1-eadf-4001-8422-2a432370ff72 · outbound
Towards Distributed Neural Architectures DARTS: Differentiable Architecture Search
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45d82e4b-be64-4207-a54c-5c5bbec99369 · outbound
Towards Distributed Neural Architectures Fineweb-edu: the finest collection of educational content, 2024.https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 691fbcff-885e-4b63-9d50-80ac488b6cdc · outbound
Towards Distributed Neural Architectures Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1811a32-701b-46bc-9a0f-716da38d2905 · outbound
Towards Distributed Neural Architectures Deep Information Propagation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c29bac-d4fe-4df1-bc35-eb61e82108aa · outbound
Towards Distributed Neural Architectures Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4618a6-8bc6-4de1-82ac-58aa2f2a9a07 · outbound
Towards Distributed Neural Architectures Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66a3de26-abb7-40ef-a9d7-45045dcee1bd · outbound
Towards Distributed Neural Architectures torchtune: Pytorch’s finetuning library, April 2024.https//github.com/ pytorch/torchtune
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c6aa4a5-2d58-45df-9390-c1ea56231cb4 · outbound
Towards Distributed Neural Architectures Neural Architecture Search with Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126d581e-3eb2-482f-8a24-b3638f6ea7a0 · outbound
Towards Distributed Neural Architectures Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba1c9cb3-de61-4bc8-baaf-a7e2e568a5b4 · outbound
Towards Distributed Neural Architectures B Module Usage and Load Balancing We plot the module usage distribution for all DNA models used in the main text in Fig
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6395a7bd-f507-47d1-b607-48576b07aec8 · outbound
Towards Distributed Neural Architectures The random noise is per-pixel zero-mean, and has a linearly decaying variance, starting at 1 and ending at 0 by the end of the optimization procedure
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbb959d7-42ec-4192-bbe4-8f53e54d3e56 · outbound
Towards Distributed Neural Architectures 3, we find that the patches following the same path in a randomly initialized model share much greater visual similarities
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08fb4c44-9031-43fa-a905-51093eb1bc92 · outbound
Towards Distributed Neural Architectures 9 is a zoomed-in version of those two figures
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88848d09-fc48-4bfe-9fbd-977fbd2abc37 · outbound
Towards Distributed Neural Architectures Scaling Laws for Neural Language Models
Reference 1991
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5161ef2f-64d8-456b-9d59-5bfba7bd16c5 · outbound
Towards Distributed Neural Architectures RACE: Large-scale ReAding comprehension dataset from examinations
Reference 2012
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9f46559-0637-4498-a104-7320a4abcc83 · outbound
Towards Distributed Neural Architectures LLM Pretraining with Continuous Concepts
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99607442-268c-4eac-9c9c-65012a38d471 · outbound
Towards Distributed Neural Architectures Distilling the Knowledge in a Neural Network
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d627122-77c7-40cc-9e44-a97146762004 · outbound
Towards Distributed Neural Architectures PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f7d93b-bee0-48bd-bfb2-1abc51917bee · outbound
Towards Distributed Neural Architectures Do language models use their depth efficiently? arXiv preprint arXiv:2505.13898,
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b911b4-b597-4b18-8395-38a785d1ba97 · outbound
Towards Distributed Neural Architectures Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8483eb-53bc-4772-9926-b0885c75eb1e · outbound
Towards Distributed Neural Architectures FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9122c42c-cc6e-464e-84c7-8063b72564ae · outbound
Towards Distributed Neural Architectures LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538acbe2-ac9c-413a-928f-44ade83f4014 · outbound
Towards Distributed Neural Architectures The Llama 3 Herd of Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1d56b4-b56f-412e-8ea8-f24089f2cff0 · outbound
Towards Distributed Neural Architectures Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a084e5-e6ef-4663-85d2-3565f95c0be8 · outbound
Towards Distributed Neural Architectures The Unreasonable Ineffectiveness of the Deeper Layers
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98349d3c-4cf6-4b1e-bfa9-44ccd5c843da · outbound
Towards Distributed Neural Architectures What do Vision Transformers Learn? A Visual Exploration
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.