Pith. sign in

Paper Citation Record · LEDGER

Leaner Transformers: More Heads, Less Depth

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2505.20802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20802 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:13.305354Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa05ce8e-8208-4cdb-91a9-9f0d85dfee56 · outbound

This paper cites A deep conditioning treatment of neural networks.

Leaner Transformers: More Heads, Less Depth A deep conditioning treatment of neural networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:19.425180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:09.034328Z digest=sha256:1d312335d50dae9d0730cbccedd4cb3d224e49e3bba911770818ebaa62f65c99

Observation 01750a21-676d-4050-b7cd-1a2ce70de4ec · outbound

This paper cites Xcit: Cross-covariance image transformers.

Leaner Transformers: More Heads, Less Depth Xcit: Cross-covariance image transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:19.156179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:09.162406Z digest=sha256:5648491798243a03a4cbd8e5ef23a1344441ac2a34421abbb35d64f39d69fa1e

Observation 622a335c-181c-42a2-be4b-66791f259c5f · outbound

This paper cites On the op- timization of deep networks: Implicit acceleration by over- parameterization.

Leaner Transformers: More Heads, Less Depth On the op- timization of deep networks: Implicit acceleration by over- parameterization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.867260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:09.316954Z digest=sha256:7752c2b4a0f7c92ae8687cc87cdc883ba9dd353d88f1ae6ec5d7ff7d8f31cb8c

Observation 3d4441ca-1002-4665-9104-6606be01699b · outbound

This paper cites End-to- end object detection with transformers.

Leaner Transformers: More Heads, Less Depth End-to- end object detection with transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:09.459325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:09.459325Z digest=sha256:561d0332b05295a27924b68eb916733621ec6bccad238dda7f71b6838609bcb9

Observation 832261df-1c02-4e3d-9e0e-6dc5476c78b7 · outbound

This paper cites Rethinking attention with performers.

Leaner Transformers: More Heads, Less Depth Rethinking attention with performers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.721286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:09.561350Z digest=sha256:8fc55e25a2bdfa5347497a3dddaa2701ad43c96831418f96a6c5e1d4e55f912e

Observation a1419588-3169-40f8-ad3a-796f240d693b · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Leaner Transformers: More Heads, Less Depth BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:09.658901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:09.658901Z digest=sha256:38e97959d056a04512860da12b88e3cd10b004fecf1defdd92b68e9f1b752658

Observation 980b9ea1-4735-4fe4-986c-45f30d8d05f2 · outbound

This paper cites Davit: Dual attention vision transform- ers.

Leaner Transformers: More Heads, Less Depth Davit: Dual attention vision transform- ers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.551570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:09.807743Z digest=sha256:efc642f628d2d50bce7bd25cbfb24d67990d113205e0ec775822f348b11caec0

Observation 93c57150-92b7-43bf-bdaf-4e2adab4804a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Leaner Transformers: More Heads, Less Depth An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:09.946194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:09.946194Z digest=sha256:1b603a3d5aaaf97866c53b5a62b48c8ab968737032b7a1902cd420228592033d

Observation ef95e9bf-f52e-4ea4-ad23-0d8c634941a0 · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Leaner Transformers: More Heads, Less Depth TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.115483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.115483Z digest=sha256:03efe9cdea4aaa06e105d166c3d446c436ca01949b2051a74e099f2969dd6a52

Observation c2f8fbce-d0ed-46bb-ae0b-cc288c2dc65a · outbound

This paper cites Drive like a human: Rethinking au- tonomous driving with large language models.

Leaner Transformers: More Heads, Less Depth Drive like a human: Rethinking au- tonomous driving with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.332722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:10.232092Z digest=sha256:03d7af3711960777ca2f9b3511bf706ad698da8a62d80089b6dc4888d15d7ed0

Observation 4076683d-2bb3-46c1-844a-813c4778f542 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Leaner Transformers: More Heads, Less Depth The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.351485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.351485Z digest=sha256:1b97c3481f5bfae8afc65e382234ae5308fc985f81407bcc3db634900b137f5e

Observation 94949c45-019d-4f6a-8097-fccfe740ac80 · outbound

This paper cites Cramming.

Leaner Transformers: More Heads, Less Depth Cramming

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.181276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:10.468447Z digest=sha256:2ba4f5418220b88a94028898f4cbdae8338975f302ec1806ef10fb6376c377a7

Observation 8122e36f-47df-4a41-b51d-88980b3c09d1 · outbound

This paper cites Cramming: Training a language model on a single gpu in one day.

Leaner Transformers: More Heads, Less Depth Cramming: Training a language model on a single gpu in one day

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:18.003646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:10.536029Z digest=sha256:f9b8085c32141f0d0af553abec78b32895877a6d3c12f077ae6f3fcba40aab30

Observation d44dec49-3d12-4736-8a91-3f0a003c6f62 · outbound

This paper cites Transformer in transformer.

Leaner Transformers: More Heads, Less Depth Transformer in transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.628588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.628588Z digest=sha256:0783b0e00b105b9698b7ab86d0fd65f6cf9409cd0c35361e455d185ff33ba7d6

Observation a817ca96-d9b8-4b60-b04e-993a1772915f · outbound

This paper cites Neu- ral tangent kernel: Convergence and generalization in neural networks.

Leaner Transformers: More Heads, Less Depth Neu- ral tangent kernel: Convergence and generalization in neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.755706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:10.710054Z digest=sha256:920dace61a17896d9907a0bc693c9015641a32e82bdf7a053736a7c80fa95ede

Observation d4433458-c758-451a-b672-a7cd124caa67 · outbound

This paper cites On the size of convolutional neural networks and generalization per- formance.

Leaner Transformers: More Heads, Less Depth On the size of convolutional neural networks and generalization per- formance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.540790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:10.798388Z digest=sha256:917b23e96ec978e227f45dd0ad6f40744c8bd24e5b7c2c0ce381a731a4f89ee8

Observation 28f6b77f-0022-4929-9d0e-ccb0d27f4a67 · outbound

This paper cites Re- former: The efficient transformer.

Leaner Transformers: More Heads, Less Depth Re- former: The efficient transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.374388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:10.880408Z digest=sha256:91f138339941dfc96178efde514ed77869a97f1a43c8fca04ca4d7d3574bbc13

Observation 8c440c53-4fd6-4377-88bc-30bded8e3cc9 · outbound

This paper cites The Depth-to-Width Interplay in Self-Attention.

Leaner Transformers: More Heads, Less Depth The Depth-to-Width Interplay in Self-Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:10.980819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:10.980819Z digest=sha256:3de15e0a0df777e2bcd8e75cd91753fb8e9b9a5fbbd72a59a080da09e10d4a98

Observation a4656912-f2d7-4c95-9724-11c96f5988fe · outbound

This paper cites Limits to depth efficiencies of self-attention.

Leaner Transformers: More Heads, Less Depth Limits to depth efficiencies of self-attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:17.164882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.067802Z digest=sha256:b699d7c8f4ad9916a879d8cc4c17f33116dc917ebbf0eefe32ff097c55a674bb

Observation 8bcb3d35-6155-4a67-adfa-a7f17636c20f · outbound

This paper cites On Tighter Generalization Bound for Deep Neural Networks: CNNs, ResNets, and Beyond.

Leaner Transformers: More Heads, Less Depth On Tighter Generalization Bound for Deep Neural Networks: CNNs, ResNets, and Beyond

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:11.154002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:11.154002Z digest=sha256:df29f8644c9da9641a3c4c75cc0a956cbbc8a86756a5992a63f24dd1ac83d26d

Observation 424b5107-c515-41ca-ab89-a6e7104cf9e4 · outbound

This paper cites Loss land- scapes and optimization in over-parameterized non-linear systems and neural networks.

Leaner Transformers: More Heads, Less Depth Loss land- scapes and optimization in over-parameterized non-linear systems and neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.915596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.267568Z digest=sha256:b854bcc861925f33921d7405b1b479829a3f9a600bb26bf7a1db3f6c30541579

Observation 4b3b603c-76cd-40f3-89ff-27477b4cdb5b · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Leaner Transformers: More Heads, Less Depth Swin transformer: Hierarchical vision transformer using shifted windows

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:11.348724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:11.348724Z digest=sha256:1abc8faec93d30dcca04682b980efb5a0c380d0903556bbf3244a627ab74d2ff

Observation ed0b4e89-b93e-4d38-b190-339e79bcc449 · outbound

This paper cites The expressive power of neural networks: A view from the width.

Leaner Transformers: More Heads, Less Depth The expressive power of neural networks: A view from the width

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.716353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.438467Z digest=sha256:e7b3b911ee857fb18a460915ccefa61d50b41c102684324c29cd5e68803a3206

Observation 721e7453-42f3-4eeb-9e84-7b01e24dea3b · outbound

This paper cites Transfusion: Multi-modal fusion network for semantic segmentation.

Leaner Transformers: More Heads, Less Depth Transfusion: Multi-modal fusion network for semantic segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.571609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.544733Z digest=sha256:3669ba239a582aa2843382b15b089db3d25344dbeae30cec25f1abb4ff57045f

Observation 8bf7a85e-954d-44b9-934e-6add32e88502 · outbound

This paper cites Numerical optimiza- tion.

Leaner Transformers: More Heads, Less Depth Numerical optimiza- tion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.426135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.625386Z digest=sha256:57b8d9147cf162406ee052465b05211239fc73a0c9f874023086f5f8bbfd531c

Observation afafbf8c-2619-4c9a-be87-d9271319f90d · outbound

This paper cites The impact of depth and width on transformer language model generalization.

Leaner Transformers: More Heads, Less Depth The impact of depth and width on transformer language model generalization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.267841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.758211Z digest=sha256:d473d7dfbbbe4df58df2dfb0644ac944fd34621ad09a452003014d6c95acfb68

Observation fc66014d-698a-487f-b293-a4faf56edf7e · outbound

This paper cites Exponential expressivity in deep neural networks through transient chaos.

Leaner Transformers: More Heads, Less Depth Exponential expressivity in deep neural networks through transient chaos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:16.130245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.859104Z digest=sha256:aa23c69c87f96396638998b1af65d8e37ed7d226dd4c5dae620443380f6be619

Observation adb98fdc-43f3-4d81-bbe6-48334022aafa · outbound

This paper cites Tiny-stories-gpt.

Leaner Transformers: More Heads, Less Depth Tiny-stories-gpt

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.929889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:11.950438Z digest=sha256:c14d93d3ad7c59c0415e3941ee286f7365b33b65aacb053392d0d767dd6796a7

Observation 4e46a6c6-03c5-4f7f-a075-8794bc32805f · outbound

This paper cites Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data.

Leaner Transformers: More Heads, Less Depth Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.040288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.040288Z digest=sha256:a7659b0c23180db49a61111d95e39c16d54822cbd4d022452a83e9e5b4632f03

Observation 43510b91-c664-4914-85d7-8c226607df28 · outbound

This paper cites Representational strengths and limitations of transformers.

Leaner Transformers: More Heads, Less Depth Representational strengths and limitations of transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.767405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.116907Z digest=sha256:9b1b71d378a7797361bdbbcebbf9fc902650596166cb3435b6823a205a7faa08

Observation 1985930e-d536-4375-a2b6-5680575faf4e · outbound

This paper cites Real analysis: measure theory, integration, and Hilbert spaces.

Leaner Transformers: More Heads, Less Depth Real analysis: measure theory, integration, and Hilbert spaces

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.626931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.186595Z digest=sha256:55d380fbf37bee1fc5bb63f944c4678db388d1da1336bb5abc123cb28b84110a

Observation 764911ee-d831-4bb5-84c0-49aa58cfe182 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

Leaner Transformers: More Heads, Less Depth How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.281412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.281412Z digest=sha256:36e4f02c0b80d569534e20ed4f673d439be9f824d2c1a0e84523eee32f932555

Observation 31d27f5e-ddd2-4935-8e67-82d97cfd151a · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Leaner Transformers: More Heads, Less Depth Long Range Arena: A Benchmark for Efficient Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.361815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.361815Z digest=sha256:7afe4024ef60d5390cbfe8a7eeadcdb5f53e9ffbc398ca54d48e4956f36e72ca

Observation 36dc04fe-54e3-4747-b9aa-194264362edc · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Leaner Transformers: More Heads, Less Depth Training data-efficient image transformers & distillation through at- tention

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.394585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.442508Z digest=sha256:6026c9345309d63caf7e2f5e317cd16429a41d1181d9b98a509926acf6a01013

Observation 5849d952-ba2b-4f5b-be49-8e315ee16638 · outbound

This paper cites Width is less important than depth in relu neural networks.

Leaner Transformers: More Heads, Less Depth Width is less important than depth in relu neural networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:15.177550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.535662Z digest=sha256:ed0e4c8c847a5e251cc6ed0f1d05e5577ee740d7786f7035c001161a3480538d

Observation 4244fd36-db2f-401f-9ad1-5a334b7590c3 · outbound

This paper cites Attention is all you need.

Leaner Transformers: More Heads, Less Depth Attention is all you need

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.854796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.619327Z digest=sha256:eff12cb0db11fafa55a81c54e5f4fdbea02384c01b107b77a71c1fde252768b2

Observation 80c06c48-fc80-4a8e-8931-33c6daf085a9 · outbound

This paper cites High-dimensional probability: An intro- duction with applications in data science.

Leaner Transformers: More Heads, Less Depth High-dimensional probability: An intro- duction with applications in data science

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.720110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.713775Z digest=sha256:b6103baf511b884b16c4fe1f966ca2277a066c244e00c3f56adf7400ae841e2d

Observation 27206a34-a8a9-450f-aa85-cd59e28c352d · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Leaner Transformers: More Heads, Less Depth GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.781563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.781563Z digest=sha256:7c85f53d52c9cc977959622d0f45a842a0aa23ba821ca2faccf6e3f162020f1c

Observation 1b162034-64cc-4bd9-bcea-25ad16772ef1 · outbound

This paper cites Linformer: Self-attention with linear complex- ity.

Leaner Transformers: More Heads, Less Depth Linformer: Self-attention with linear complex- ity

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.514716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.836097Z digest=sha256:1aefbef7b5708309b38edc64fd9c5922604ed132383f20cb60822d4df8f61826

Observation d50ddb4e-77f8-4be6-9e21-843ccd74a516 · outbound

This paper cites Github repository, 2021.

Leaner Transformers: More Heads, Less Depth Github repository, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.325590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.907534Z digest=sha256:11cb22678bf43be67729e0b72ae606510d232fa334d238b1daa0baa4418fa950

Observation 80d8fcc6-bb68-4c83-986a-3ff2466558af · outbound

This paper cites Nystr¨omformer: A nystr¨om-based algorithm for approximat- ing self-attention.

Leaner Transformers: More Heads, Less Depth Nystr¨omformer: A nystr¨om-based algorithm for approximat- ing self-attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.156866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:12.982547Z digest=sha256:d581bc54c4bdb7eb9b82834a85bac10f857e761663dc241dcfd30e35fa45c032

Observation 622798dd-8cab-497c-95d4-262d39936b60 · outbound

This paper cites V olo: Vision outlooker for visual recog- nition.

Leaner Transformers: More Heads, Less Depth V olo: Vision outlooker for visual recog- nition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:14.014039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:13.044588Z digest=sha256:e89bf0273375a041a606127f86b10f1b14065e2c5c445305062532bf164bc0b0

Observation 73f1affc-3c99-4491-a8f3-2d1395b2a604 · outbound

This paper cites cosformer: rethinking softmax in attention.

Leaner Transformers: More Heads, Less Depth cosformer: rethinking softmax in attention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:13.848094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:13.118002Z digest=sha256:a7bdc541b849f31821c3ac35eab926191ba514c7bde3c691c647fc9056fed0fc

Observation 2d421d34-2722-4ebe-b59a-3ee49888125a · outbound

This paper cites Understanding generalization and optimization performance of deep cnns.

Leaner Transformers: More Heads, Less Depth Understanding generalization and optimization performance of deep cnns

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:13.683260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:13.197390Z digest=sha256:66ae106c1c62f0f2840297cd951365e39ff07850a80e7127dcc8e17311e1c910

Observation edd4c23a-d92b-4435-82b4-51a4a181152e · outbound

This paper cites A robustly optimized bert pre-training approach with post-training.

Leaner Transformers: More Heads, Less Depth A robustly optimized bert pre-training approach with post-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:13.523878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:53:13.305354Z digest=sha256:ddcc9dca7ce9346fb1f58197529b63c858d80efb85642cde1aaf656974390d97

Pith citing papers

No inbound Pith citation observations are available.