Pith. sign in

Paper Citation Record · LEDGER

Reducing Transformer Depth on Demand with Structured Dropout

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:1909.11556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.11556 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:17.834796Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

273
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1ad9f5c-7c94-4fca-b410-26e8483eec12 · inbound

PyTorch Distributed: Experiences on Accelerating Data Parallel Training cites this paper.

PyTorch Distributed: Experiences on Accelerating Data Parallel Training Reducing Transformer Depth on Demand with Structured Dropout

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:15:26.852762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T19:15:26.764929Z digest=sha256:a2298392a0c0855f6d43eaaec0085d3d0b9195fa19403e5563e06a64c276b4a3

Observation 686627a5-f0e5-4e24-a436-285b6556240f · inbound

Eliciting Latent Predictions from Transformers with the Tuned Lens cites this paper.

Eliciting Latent Predictions from Transformers with the Tuned Lens Reducing Transformer Depth on Demand with Structured Dropout

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:54:37.594176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T16:54:37.382049Z digest=sha256:85a5a3cbd06ae5a61f6066899b41a8fdc732ac1c98c3db6d65cf0ceb05ef4c2d

Observation 2a778857-1a0e-42dd-9422-7548c2e75db6 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Reducing Transformer Depth on Demand with Structured Dropout

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.141062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:8059bdee46e1f0da6571e42d1972d38cb68128fb19b999ff5022cbf8d82d67ec

Observation 2ac5c55d-a864-4dac-b941-0025254eda76 · inbound

AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer cites this paper.

AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer Reducing Transformer Depth on Demand with Structured Dropout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:17.834796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:17.834796Z digest=sha256:8178e4c1c562e380ba56de1755b683b9b098029f68d8a1d4efe089076de7569b

Observation 211c76de-123e-4481-a77a-df405572e788 · inbound

Position: The Future of Bayesian Prediction Is Prior-Fitted cites this paper.

Position: The Future of Bayesian Prediction Is Prior-Fitted Reducing Transformer Depth on Demand with Structured Dropout

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:36.354788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:36.354788Z digest=sha256:67bae0cfe0c852b2e9ebf13abf4b47807ce85926149d47233b6ee242d04e70c6

Observation 3223e1a1-e09b-4363-b4b1-f661be8cbb40 · inbound

Learning to Skip the Middle Layers of Transformers cites this paper.

Learning to Skip the Middle Layers of Transformers Reducing Transformer Depth on Demand with Structured Dropout

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:53.679247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:40:53.679247Z digest=sha256:6cd6c8a64a7eae4c0f6ec13492cd2743eb9bbc30802af2ea267656438e41abf6

Observation 282d4541-9bd4-46a8-b938-768d71b1ce3b · inbound

Towards Universal & Efficient Model Compression via Exponential Torque Pruning cites this paper.

Towards Universal & Efficient Model Compression via Exponential Torque Pruning Reducing Transformer Depth on Demand with Structured Dropout

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:42.899383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:42.899383Z digest=sha256:cee9983ee069b82585b64056b7b6f22e1db6b7942d9ee0750bc04576512d498c

Observation 54e80b2d-9a4e-45cd-b894-d3ffe7f781a1 · inbound

Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs cites this paper.

Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs Reducing Transformer Depth on Demand with Structured Dropout

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:19.194766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:19.194766Z digest=sha256:9a585d8d662f7c4e348bb1c877e78b64d0bd07dbd20b70aebd45d3b2dd67209b

Observation ec931f3a-5d9b-4e1a-b351-2bc27cd39f80 · inbound

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling cites this paper.

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling Reducing Transformer Depth on Demand with Structured Dropout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.447217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.447217Z digest=sha256:a6dfbbe4d9132110e7c71cb9088732bedcd56343c1052e6c00b51b5e5f2eb63e

Observation 3776073e-b025-4559-a80c-f2d61610c801 · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study Reducing Transformer Depth on Demand with Structured Dropout

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:44.726762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:44.726762Z digest=sha256:5ce1b9afc967f4a5809a4e29325b766937d56d6ae81b9af4870514b2fd4c262a

Observation b220cf03-cacc-40ac-99fe-602821af9ace · inbound

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code cites this paper.

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code Reducing Transformer Depth on Demand with Structured Dropout

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:01:55.897855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T00:01:42.228190Z digest=sha256:fa9033a795005d535d0c59685a90f89c1eb90a9635f903d3766400b9ef9654f2

Observation e4d35bd6-cf40-4be1-9924-1ff35b2f3be8 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models Reducing Transformer Depth on Demand with Structured Dropout

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:36:56.313975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T00:34:38.247099Z digest=sha256:e87c73369eadf03ef4f396e1b06788746be9b795b4345a71e253b8ca0f5f7b8e

Observation 8b34d409-72f4-470e-8e2e-b23b693ea384 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models Reducing Transformer Depth on Demand with Structured Dropout

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T00:02:58.946950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:02:58.946950Z digest=sha256:82a8ab92da1172903769cddcdc5c94f59071fe329c3d934606a622b6d955ab85

Observation 1beebdf1-25db-475a-9bc5-eb1303dad0e5 · inbound

Harnessing Input-Adaptive Inference for Efficient VLN cites this paper.

Harnessing Input-Adaptive Inference for Efficient VLN Reducing Transformer Depth on Demand with Structured Dropout

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:22.237618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:22.237618Z digest=sha256:58d1e8d66811fb5248ed7462d457ca9a89d4fc58685cd15617192d72b50839ac

Observation 483f8e97-ee2b-4ef9-8ae3-c0b17994d552 · inbound

VISP: Volatility Informed Stochastic Projection for Adaptive Regularization cites this paper.

VISP: Volatility Informed Stochastic Projection for Adaptive Regularization Reducing Transformer Depth on Demand with Structured Dropout

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:37.008912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:37.008912Z digest=sha256:f29e2e0df5f2027b146898d73737f9edce68a49af03f396b1413d64571acee4f

Observation 3f4346a0-e7d6-4651-ba8f-7c985a2c2f7b · inbound

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures cites this paper.

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:38:07.556104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T18:35:23.814733Z digest=sha256:939c7a48a95f2cd0042299e7b68207f647121088727392d9a7172f39eaf4e40e

Observation d1ac1b49-3e8e-4b59-a63e-c429923851db · inbound

Depth Adaptive Efficient Visual Autoregressive Modeling cites this paper.

Depth Adaptive Efficient Visual Autoregressive Modeling Reducing Transformer Depth on Demand with Structured Dropout

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.625427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:22:55.035749Z digest=sha256:65bc968b254b595192654ceef00dab270956f987fe4b7f5473ab7ce01be57b7c

Observation c01b4b4a-0906-42c7-af38-7334de1511bf · inbound

Language models recognize dropout and Gaussian noise applied to their activations cites this paper.

Language models recognize dropout and Gaussian noise applied to their activations Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:02.186788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:38:12.882914Z digest=sha256:2c791066662e0768e20cd26245e29485eae12231cd7694c719d27e6847d83e58

Observation 7351aec5-c8c2-4627-b739-ae141d653acf · inbound

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing cites this paper.

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing Reducing Transformer Depth on Demand with Structured Dropout

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:58:12.041938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T19:56:48.015363Z digest=sha256:44efec316907885312309db161a9fd74a15894f705ad8b2a50bae581edb88e63

Observation 67ff6be4-e9a7-4148-b7d1-cbc5aefda9fb · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency Reducing Transformer Depth on Demand with Structured Dropout

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:14.586000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:9f810f7bb65c23564f4a8b9f17b1fe9b41b7aa29b34208a9775371fbb380dcca

Observation bf2da7a7-a8d6-447c-a3d8-8ceccf243a33 · inbound

SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations cites this paper.

SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations Reducing Transformer Depth on Demand with Structured Dropout

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:29.267468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:08:12.351139Z digest=sha256:c7e0f75ce32b916bb8905cf404e2a34efd7a52d808e1ec03cd36932001b73b0c

Observation c13c2308-d5ae-45b4-981d-07a7b3f47ed0 · inbound

Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility cites this paper.

Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:09.011862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T10:33:52.457620Z digest=sha256:4863451fa47ec80def1bdb179384686cc4da874ccd017e2b5a9e0a52660eb716

Observation e9c90031-af5b-441b-9322-bd19b77d0272 · inbound

Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs cites this paper.

Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs Reducing Transformer Depth on Demand with Structured Dropout

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:58:04.030648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:50:10.564922Z digest=sha256:563ebb86a8d2b1d83035d54281a9e3fecf43ac6bd32087ea573189610c8d133b

Observation 972ad1df-c680-4ad2-807f-6c3adf3cb646 · inbound

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility cites this paper.

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility Reducing Transformer Depth on Demand with Structured Dropout

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:39:47.975021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:35:09.705532Z digest=sha256:402cea7936ec22824f72da531a8219c76edfe61c9ad5da568c7794a4e99b1055

Observation a9ae38b3-1ea4-455c-add4-905c249bdbf4 · inbound

Skip a Layer or Loop It? Learning Program-of-Layers in LLMs cites this paper.

Skip a Layer or Loop It? Learning Program-of-Layers in LLMs Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:57.370073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:56:34.435152Z digest=sha256:ee4b046c0fc8d42dc9010e6e3e28120eeeebf5fe562c8f3d7b523a2d13caaebe

Observation 7cada08f-0e51-4b6a-b246-8e3b26a7c03d · inbound

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation cites this paper.

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation Reducing Transformer Depth on Demand with Structured Dropout

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:41:03.364124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:32:09.784716Z digest=sha256:08777539dce735d34cf4628c44431b21e37a6108b7cd4d262080345389984f35

Observation 557c46b0-6da5-4ea7-a4ed-dee2dfdb2bb1 · inbound

Tapered Language Models cites this paper.

Tapered Language Models Reducing Transformer Depth on Demand with Structured Dropout

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.954130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:11:20.341634Z digest=sha256:feec1cb352f56b46153a55463a35574bcf574e696ce0ad59d63094cab526642a

Observation 23ae4396-631b-4c90-a1cd-57c22b8282ba · inbound

Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping cites this paper.

Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping Reducing Transformer Depth on Demand with Structured Dropout

Reference 34

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T07:34:21.633951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:29:23.786653Z digest=sha256:43d00dd43e6786760068595570820f09e5bc1fc93b9c8c531c6da159113713de

Observation ee248a96-9780-4b07-b3fe-04d7b3b56d60 · inbound

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study cites this paper.

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study Reducing Transformer Depth on Demand with Structured Dropout

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:20:38.521564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:20:38.521564Z digest=sha256:cde953c1f543c5d6905eb57aa8de0a88ad37ca7caf3b4c8173de416eb57776bb