Pith. sign in

Paper Citation Record · LEDGER

Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.01733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01733 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:23:43.299853Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68593e84-379b-40e8-b2fd-bbb5509dca58 · inbound

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference cites this paper.

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:18:42.398669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:17:11.261301Z digest=sha256:5a412da5705ed88addb9b39c752771ff3180966510f8b30660cfc85e5bc970e2

Observation d02a5606-e915-4bf8-892a-ba68de491c37 · inbound

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching cites this paper.

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:30:44.403100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:28:08.873654Z digest=sha256:1f70a02849eefdfb6f8583ab74bcd1b9cdc5acbd844af7238edca41c438b67ac

Observation 4625368b-bcb5-4845-8a89-9cebe998ecd0 · inbound

AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers cites this paper.

AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:56:50.255035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:55:15.572867Z digest=sha256:defe604c2b0926894c6ff9b2d28d6c43e2a2689b4eb271a42343d35bbdaac757

Observation e1433d7d-a275-4bc1-968f-53b6163b80d6 · inbound

LESA: Learnable Stage-Aware Predictors for Diffusion Model Acceleration cites this paper.

LESA: Learnable Stage-Aware Predictors for Diffusion Model Acceleration Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:23:43.299853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:23:43.299853Z digest=sha256:b95d161cc80873ee3bde8f7a845b220073ec7e7467751cd389fcdbf27f092186

Observation e4195a67-9b26-4346-8864-d8dc672cffc6 · inbound

S2O: Early Stopping for Sparse Attention via Online Permutation cites this paper.

S2O: Early Stopping for Sparse Attention via Online Permutation Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:32.944227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T19:32:52.948154Z digest=sha256:282a008bb76218a3cba844361d2bf38258b448c6e9f0a0ddb5f3f08692823033

Observation 4923d8da-09eb-4f6b-a655-f3c9fadb43cc · inbound

Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models cites this paper.

Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:47:33.002955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:43:45.415485Z digest=sha256:19c8bbb576d8e8a086b2bc4d6319ee51e96f68831098c96f931e4704f8a73181

Observation e3c1db61-0a81-487f-bf8c-4f61f74457fd · inbound

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism cites this paper.

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:20.640324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T10:23:19.236803Z digest=sha256:15400f67567d9ec3970512ba414b068134acf2ff16306386c3c6fb2639ab3525

Observation 17d2dd36-afec-4d65-9b70-5e6a87ecde36 · inbound

DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse cites this paper.

DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:27.256030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:27.256030Z digest=sha256:7195ad521b1e97173ddbd4fdd818c87f2ea36f1080b72a05da805601fc24c742