Pith. sign in

Paper Citation Record · LEDGER

Towards Distributed Neural Architectures

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.22389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22389 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:14:20.461704Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0975920-8ccb-4fc5-bfeb-2264c07accdc · outbound

This paper cites GPT-4 Technical Report.

Towards Distributed Neural Architectures GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.370267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.370267Z digest=sha256:b9ee985ea12f4c729b2c7dfd676a183282dbb26595b366dd33d69043f40e2034

Observation 2482fa62-7024-41dd-992c-ebe0d7109451 · outbound

This paper cites Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei.

Towards Distributed Neural Architectures Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.796382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.796382Z digest=sha256:3946e2a872df0b7e788d46905f592fe6f9e6b331adb280a532d9aef48d32f90e

Observation 280f67af-b0b0-4421-a7a3-b66c0271f621 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Towards Distributed Neural Architectures Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.944144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.944144Z digest=sha256:c7762bb3e1511debf9ca72b8863d2394a0a250ca6d2eeb3410440f1c9f97c78a

Observation 1ca5aea9-79ea-4d95-a551-6c355c8775e5 · outbound

This paper cites This may be because we are not considering a setting with high sparsity.

Towards Distributed Neural Architectures This may be because we are not considering a setting with high sparsity

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.567025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:20.148222Z digest=sha256:ec81d10c467d7692c6d20b0812f8b7ec9ffc94cd0d1778960f2492a43040caae

Observation 16ee2b9c-a491-44bb-b343-14df6bd663b4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Distributed Neural Architectures DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.357714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.357714Z digest=sha256:f0e716cd89ea0a15875246d782381e8adc71d74c07c74c6098549e8ef254c011

Observation b37c74e1-eadf-4001-8422-2a432370ff72 · outbound

This paper cites DARTS: Differentiable Architecture Search.

Towards Distributed Neural Architectures DARTS: Differentiable Architecture Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.416905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.416905Z digest=sha256:136c6fe44b0c8c5c956346017c9d0a43321c505ff54255bb4f218a25807c91f4

Observation 45d82e4b-be64-4207-a54c-5c5bbec99369 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu.

Towards Distributed Neural Architectures Fineweb-edu: the finest collection of educational content, 2024.https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:22.111671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:19.468676Z digest=sha256:67907007da5c56db73e9e067afc811747889f710b49d6d36806db0f17047e38b

Observation 691fbcff-885e-4b63-9d50-80ac488b6cdc · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Towards Distributed Neural Architectures Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.596966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.596966Z digest=sha256:0e2ed6c75d2463ed8ec699393fc9e97e70c512752180d94ab8a0fcdfbd397f4b

Observation d1811a32-701b-46bc-9a0f-716da38d2905 · outbound

This paper cites Deep Information Propagation.

Towards Distributed Neural Architectures Deep Information Propagation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.643965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.643965Z digest=sha256:3cfec6150ef07c18d2f6571ae8bb092e6f7a012540c7382dec7b2057f30c2bc1

Observation b3c29bac-d4fe-4df1-bc35-eb61e82108aa · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Towards Distributed Neural Architectures Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.690467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.690467Z digest=sha256:0f8286316cd4163a50a36fd5822a17a51e63899e38cffa768ac2efac7bece525

Observation 2a4618a6-8bc6-4de1-82ac-58aa2f2a9a07 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,.

Towards Distributed Neural Architectures Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:22.010935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:19.742140Z digest=sha256:4c98907b4c2d3ad3a8772b854ac4a366b85be5147662332f9e6a3b02941483cf

Observation 66a3de26-abb7-40ef-a9d7-45045dcee1bd · outbound

This paper cites torchtune: Pytorch’s finetuning library, April 2024.https//github.com/ pytorch/torchtune.

Towards Distributed Neural Architectures torchtune: Pytorch’s finetuning library, April 2024.https//github.com/ pytorch/torchtune

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.905199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:19.869553Z digest=sha256:eca2ba6f4b50fc56776f3d0d464e3d3d55b135d6cb020669d90fafd8e5563129

Observation 9c6aa4a5-2d58-45df-9390-c1ea56231cb4 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Towards Distributed Neural Architectures Neural Architecture Search with Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.952473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.952473Z digest=sha256:d5a148afed797b577b73eb2fb3a093d517f7640ea5f1d35912bc6780dea286c4

Observation 126d581e-3eb2-482f-8a24-b3638f6ea7a0 · outbound

This paper cites an unresolved cited work.

Towards Distributed Neural Architectures Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:14:21.820871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:20.040098Z digest=sha256:04e6daefc0da8226030c397656b7902f407ab4c8d88f7b69f1b3e8bd44a34728

Observation ba1c9cb3-de61-4bc8-baaf-a7e2e568a5b4 · outbound

This paper cites B Module Usage and Load Balancing We plot the module usage distribution for all DNA models used in the main text in Fig.

Towards Distributed Neural Architectures B Module Usage and Load Balancing We plot the module usage distribution for all DNA models used in the main text in Fig

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.678283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:20.045284Z digest=sha256:08a74f26fd57abe08a490ff7c4168151c32a6fa3091827659574a6ed2470f775

Observation 6395a7bd-f507-47d1-b607-48576b07aec8 · outbound

This paper cites The random noise is per-pixel zero-mean, and has a linearly decaying variance, starting at 1 and ending at 0 by the end of the optimization procedure.

Towards Distributed Neural Architectures The random noise is per-pixel zero-mean, and has a linearly decaying variance, starting at 1 and ending at 0 by the end of the optimization procedure

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.382742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:20.256809Z digest=sha256:e819a2e652c7d0ac74316449e3252314194211ab19fb644424fda6e51569be7a

Observation bbb959d7-42ec-4192-bbe4-8f53e54d3e56 · outbound

This paper cites 3, we find that the patches following the same path in a randomly initialized model share much greater visual similarities.

Towards Distributed Neural Architectures 3, we find that the patches following the same path in a randomly initialized model share much greater visual similarities

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.216489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:20.374982Z digest=sha256:66f8e05eb2931b14bd8d5499e9618fdc8a927e1a619bc865ac6aeea57dce5bd3

Observation 08fb4c44-9031-43fa-a905-51093eb1bc92 · outbound

This paper cites 9 is a zoomed-in version of those two figures.

Towards Distributed Neural Architectures 9 is a zoomed-in version of those two figures

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.112995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:20.461704Z digest=sha256:2d45183d2958119c9dcd92480f873ae074f2a0d010c7329d72bdeb7d28127746

Observation 88848d09-fc48-4bfe-9fbd-977fbd2abc37 · outbound

This paper cites Scaling Laws for Neural Language Models.

Towards Distributed Neural Architectures Scaling Laws for Neural Language Models

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.232811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.232811Z digest=sha256:fd23cac9f9ac540ad75fa2725c81b09c6565c096b5ddf3e898bd6c7c338e2098

Observation 5161ef2f-64d8-456b-9d59-5bfba7bd16c5 · outbound

This paper cites RACE: Large-scale ReAding comprehension dataset from examinations.

Towards Distributed Neural Architectures RACE: Large-scale ReAding comprehension dataset from examinations

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:22.229680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:14:19.291521Z digest=sha256:7d04692d71310428a3a7731e8a76398ff75f786c72778ecae9e45dd874d7ee82

Observation e9f46559-0637-4498-a104-7320a4abcc83 · outbound

This paper cites LLM Pretraining with Continuous Concepts.

Towards Distributed Neural Architectures LLM Pretraining with Continuous Concepts

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.808871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.808871Z digest=sha256:d85cfdd1f77e8ff56876fc0a4b33dcdc3dcb712df9c5be5d8a011768575ba675

Observation 99607442-268c-4eac-9c9c-65012a38d471 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Towards Distributed Neural Architectures Distilling the Knowledge in a Neural Network

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.192043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.192043Z digest=sha256:535b86a7c514b573999b37c3d3c8a9ca72384da9bcda4e41cbaa5ea105f75dd8

Observation 1d627122-77c7-40cc-9e44-a97146762004 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Towards Distributed Neural Architectures PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:14:19.516841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.516841Z digest=sha256:646bc31b9b55f25140b1eb4d20d6fc0dd29edcc7b666731a0dbc16e4b4152cf4

Observation b2f7d93b-bee0-48bd-bfb2-1abc51917bee · outbound

This paper cites Do language models use their depth efficiently? arXiv preprint arXiv:2505.13898,.

Towards Distributed Neural Architectures Do language models use their depth efficiently? arXiv preprint arXiv:2505.13898,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.721573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.721573Z digest=sha256:b27284b396bcf0b262514fd7db33ef8f452439184377c77fd1a653260443349c

Observation e1b911b4-b597-4b18-8395-38a785d1ba97 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Towards Distributed Neural Architectures Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.635699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.635699Z digest=sha256:2b338979ea508f38bfa05ee90844f682a700b8aa975175798b7d952cc018b947

Observation ef8483eb-53bc-4772-9926-b0885c75eb1e · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Towards Distributed Neural Architectures FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.547802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.547802Z digest=sha256:c9719c5646cd0005fb311f0b969c5665bbcc0275ab2cda15f67f2a5c25babc02

Observation 9122c42c-cc6e-464e-84c7-8063b72564ae · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

Towards Distributed Neural Architectures LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.881708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.881708Z digest=sha256:7275ec15ec988b34ff86c97c7d408f35f404c8ba017f19d77fb684f669902845

Observation 538acbe2-ac9c-413a-928f-44ade83f4014 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Distributed Neural Architectures The Llama 3 Herd of Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.055624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.055624Z digest=sha256:d95729381a3e039867f298bba0178c2f07a85cdd2d19ff7060523d1f2ebdc859

Observation 1e1d56b4-b56f-412e-8ea8-f24089f2cff0 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Towards Distributed Neural Architectures Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.461421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.461421Z digest=sha256:2550ef36d1038277b07158a83a524b88f323093b02d5d34bafbeb27b626f25cf

Observation a0a084e5-e6ef-4663-85d2-3565f95c0be8 · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Towards Distributed Neural Architectures The Unreasonable Ineffectiveness of the Deeper Layers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.114883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.114883Z digest=sha256:18250ab64bb5b07aae3d5101af17181169c1b83ed02b077271b04be9ed7f087a

Observation 98349d3c-4cf6-4b1e-bfa9-44ccd5c843da · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Towards Distributed Neural Architectures What do Vision Transformers Learn? A Visual Exploration

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.008165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.008165Z digest=sha256:d1ea144ece8f9018abc201c4ae13e3b3d6df90b09654f6fe49a3f4aeb2de6eff

Pith citing papers

No inbound Pith citation observations are available.