Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:13.305354Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2505.20802.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:13.305354Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aa05ce8e-8208-4cdb-91a9-9f0d85dfee56 · outbound
Leaner Transformers: More Heads, Less Depth A deep conditioning treatment of neural networks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01750a21-676d-4050-b7cd-1a2ce70de4ec · outbound
Leaner Transformers: More Heads, Less Depth Xcit: Cross-covariance image transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 622a335c-181c-42a2-be4b-66791f259c5f · outbound
Leaner Transformers: More Heads, Less Depth On the op- timization of deep networks: Implicit acceleration by over- parameterization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d4441ca-1002-4665-9104-6606be01699b · outbound
Leaner Transformers: More Heads, Less Depth End-to- end object detection with transformers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832261df-1c02-4e3d-9e0e-6dc5476c78b7 · outbound
Leaner Transformers: More Heads, Less Depth Rethinking attention with performers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a1419588-3169-40f8-ad3a-796f240d693b · outbound
Leaner Transformers: More Heads, Less Depth BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980b9ea1-4735-4fe4-986c-45f30d8d05f2 · outbound
Leaner Transformers: More Heads, Less Depth Davit: Dual attention vision transform- ers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93c57150-92b7-43bf-bdaf-4e2adab4804a · outbound
Leaner Transformers: More Heads, Less Depth An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef95e9bf-f52e-4ea4-ad23-0d8c634941a0 · outbound
Leaner Transformers: More Heads, Less Depth TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f8fbce-d0ed-46bb-ae0b-cc288c2dc65a · outbound
Leaner Transformers: More Heads, Less Depth Drive like a human: Rethinking au- tonomous driving with large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4076683d-2bb3-46c1-844a-813c4778f542 · outbound
Leaner Transformers: More Heads, Less Depth The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94949c45-019d-4f6a-8097-fccfe740ac80 · outbound
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8122e36f-47df-4a41-b51d-88980b3c09d1 · outbound
Leaner Transformers: More Heads, Less Depth Cramming: Training a language model on a single gpu in one day
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d44dec49-3d12-4736-8a91-3f0a003c6f62 · outbound
Leaner Transformers: More Heads, Less Depth Transformer in transformer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a817ca96-d9b8-4b60-b04e-993a1772915f · outbound
Leaner Transformers: More Heads, Less Depth Neu- ral tangent kernel: Convergence and generalization in neural networks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4433458-c758-451a-b672-a7cd124caa67 · outbound
Leaner Transformers: More Heads, Less Depth On the size of convolutional neural networks and generalization per- formance
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28f6b77f-0022-4929-9d0e-ccb0d27f4a67 · outbound
Leaner Transformers: More Heads, Less Depth Re- former: The efficient transformer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c440c53-4fd6-4377-88bc-30bded8e3cc9 · outbound
Leaner Transformers: More Heads, Less Depth The Depth-to-Width Interplay in Self-Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4656912-f2d7-4c95-9724-11c96f5988fe · outbound
Leaner Transformers: More Heads, Less Depth Limits to depth efficiencies of self-attention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bcb3d35-6155-4a67-adfa-a7f17636c20f · outbound
Leaner Transformers: More Heads, Less Depth On Tighter Generalization Bound for Deep Neural Networks: CNNs, ResNets, and Beyond
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 424b5107-c515-41ca-ab89-a6e7104cf9e4 · outbound
Leaner Transformers: More Heads, Less Depth Loss land- scapes and optimization in over-parameterized non-linear systems and neural networks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4b3b603c-76cd-40f3-89ff-27477b4cdb5b · outbound
Leaner Transformers: More Heads, Less Depth Swin transformer: Hierarchical vision transformer using shifted windows
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0b4e89-b93e-4d38-b190-339e79bcc449 · outbound
Leaner Transformers: More Heads, Less Depth The expressive power of neural networks: A view from the width
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 721e7453-42f3-4eeb-9e84-7b01e24dea3b · outbound
Leaner Transformers: More Heads, Less Depth Transfusion: Multi-modal fusion network for semantic segmentation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bf7a85e-954d-44b9-934e-6add32e88502 · outbound
Leaner Transformers: More Heads, Less Depth Numerical optimiza- tion
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afafbf8c-2619-4c9a-be87-d9271319f90d · outbound
Leaner Transformers: More Heads, Less Depth The impact of depth and width on transformer language model generalization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc66014d-698a-487f-b293-a4faf56edf7e · outbound
Leaner Transformers: More Heads, Less Depth Exponential expressivity in deep neural networks through transient chaos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation adb98fdc-43f3-4d81-bbe6-48334022aafa · outbound
Leaner Transformers: More Heads, Less Depth Tiny-stories-gpt
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e46a6c6-03c5-4f7f-a075-8794bc32805f · outbound
Leaner Transformers: More Heads, Less Depth Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43510b91-c664-4914-85d7-8c226607df28 · outbound
Leaner Transformers: More Heads, Less Depth Representational strengths and limitations of transformers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1985930e-d536-4375-a2b6-5680575faf4e · outbound
Leaner Transformers: More Heads, Less Depth Real analysis: measure theory, integration, and Hilbert spaces
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 764911ee-d831-4bb5-84c0-49aa58cfe182 · outbound
Leaner Transformers: More Heads, Less Depth How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d27f5e-ddd2-4935-8e67-82d97cfd151a · outbound
Leaner Transformers: More Heads, Less Depth Long Range Arena: A Benchmark for Efficient Transformers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36dc04fe-54e3-4747-b9aa-194264362edc · outbound
Leaner Transformers: More Heads, Less Depth Training data-efficient image transformers & distillation through at- tention
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5849d952-ba2b-4f5b-be49-8e315ee16638 · outbound
Leaner Transformers: More Heads, Less Depth Width is less important than depth in relu neural networks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4244fd36-db2f-401f-9ad1-5a334b7590c3 · outbound
Leaner Transformers: More Heads, Less Depth Attention is all you need
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80c06c48-fc80-4a8e-8931-33c6daf085a9 · outbound
Leaner Transformers: More Heads, Less Depth High-dimensional probability: An intro- duction with applications in data science
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27206a34-a8a9-450f-aa85-cd59e28c352d · outbound
Leaner Transformers: More Heads, Less Depth GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b162034-64cc-4bd9-bcea-25ad16772ef1 · outbound
Leaner Transformers: More Heads, Less Depth Linformer: Self-attention with linear complex- ity
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d50ddb4e-77f8-4be6-9e21-843ccd74a516 · outbound
Leaner Transformers: More Heads, Less Depth Github repository, 2021
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80d8fcc6-bb68-4c83-986a-3ff2466558af · outbound
Leaner Transformers: More Heads, Less Depth Nystr¨omformer: A nystr¨om-based algorithm for approximat- ing self-attention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 622798dd-8cab-497c-95d4-262d39936b60 · outbound
Leaner Transformers: More Heads, Less Depth V olo: Vision outlooker for visual recog- nition
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 73f1affc-3c99-4491-a8f3-2d1395b2a604 · outbound
Leaner Transformers: More Heads, Less Depth cosformer: rethinking softmax in attention
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2d421d34-2722-4ebe-b59a-3ee49888125a · outbound
Leaner Transformers: More Heads, Less Depth Understanding generalization and optimization performance of deep cnns
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation edd4c23a-d92b-4435-82b4-51a4a181152e · outbound
Leaner Transformers: More Heads, Less Depth A robustly optimized bert pre-training approach with post-training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.