Pith. sign in

Paper Citation Record · LEDGER

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

As of 22 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 10 inbound Pith citation observations for arXiv:2505.14687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14687 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:12.833852Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:30:07.769382Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:45:46.109771Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1aee1501-f069-4741-a986-6b2df5d8ef2a · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers All are worth words: A vit backbone for diffusion models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.008274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.008274Z digest=sha256:997a9cf4862b29131e0187241e6d6791fd07c675ff821c46a90adf07081de589

Observation 79c19ce1-eb41-4073-bc28-c88e517ff778 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.118350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.118350Z digest=sha256:171a3096d3197b827ff9e7332b1a267d8b2f3b3f4714f8d879d74bdeb1d78ade

Observation 570f195f-2620-41dc-914c-3d3c2f7f226a · outbound

This paper cites Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:18.248784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:08.197422Z digest=sha256:2beac5686e76e4c5a083508d3e80013b99d66dd49c3a696ba824491fd86749f6

Observation 8e1e9712-4957-4e21-bf0d-88a4e3ac9617 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.290073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.290073Z digest=sha256:2bfceff59409f13fb90878281453ae4202edb4828649564085df46da4e8c467c

Observation f3978d47-c442-4331-bc50-de86e6ed8d5d · outbound

This paper cites COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.388520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.388520Z digest=sha256:272d4f1f4bae8463b2ebca40fb880b2c159d99260c87320e284c8148a772234c

Observation b65944f0-5ef0-4d1f-90ca-b981e5a77b81 · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.466620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.466620Z digest=sha256:61f9004c8b220c14449c87211556f9f81f801081cd30ff2ab577fea99dfb4251

Observation 55148d36-2d75-4aad-8792-03b2310999da · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.532943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.532943Z digest=sha256:3ce8c1360976ca3e74b7b2f0e0ab9ad78203059f89ab915458972387ef089b63

Observation 4ed792c8-8bfc-49e3-a976-f12d51407e7b · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Scaling rectified flow transformers for high-resolution image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.640437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.640437Z digest=sha256:475c88e73a421692cedcd82a656b8ff9be9fae3136341534d14fc4165f58811d

Observation c084b80d-dd89-4d79-a5c9-8a0a46f8ebde · outbound

This paper cites Structural pruning for diffusion models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Structural pruning for diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:18.103909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:08.725477Z digest=sha256:9a617513226cb528c17833b505c4f90727636c52e7a3a464b2eda99be442d2f9

Observation aa39bbf9-2236-4de0-b5dd-b96f1eddd47d · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.981770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:08.815181Z digest=sha256:0d1ad80585f9c5c0f391365602d4b1a829b9f1f86392c12f4bc33751b1dd0dbc

Observation 4b10cff7-cce4-4bf4-a41e-21bfe2969bf6 · outbound

This paper cites Neighborhood attention transformer.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Neighborhood attention transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.816114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:08.892465Z digest=sha256:fadc4b8accf84179200920febe19ad19a0a0222c7b8c51f88b62dbf887ebbe7c

Observation 0eea724e-96ed-45c3-81e6-87c7c20ea960 · outbound

This paper cites A Simple Video Segmenter by Tracking Objects Along Axial Trajectories.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers A Simple Video Segmenter by Tracking Objects Along Axial Trajectories

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:33:13.584587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:08.965378Z digest=sha256:10254e44b3a45e45f3361aa67cd086bec89e3ac8d868ca0286a8155ba5508b66

Observation 03515be4-1de5-4509-b57b-0eb2ae6b54e8 · outbound

This paper cites Flowtok: Flowing seamlessly across text and image tokens.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Flowtok: Flowing seamlessly across text and image tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.065536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.065536Z digest=sha256:370182cdaf91f5c88da22b1cf5e991373bf161f4c35e6ab4c415a685e8f9549b

Observation 03e5c20f-ce90-4244-ba60-26c8237e3e2f · outbound

This paper cites Masked autoencoders are scalable vision learners.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Masked autoencoders are scalable vision learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.139544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.139544Z digest=sha256:0cf2a91341c9c5fea89ed91b2b4a85939ae92bfa7c10a07369d09eead338b95c

Observation 788e1d04-561d-4298-ba60-ebb3ee82799e · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.231873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.231873Z digest=sha256:532fea0c9fcaf3b1023483ad01be68673756029e2cf10e97499c8d6f3b2b6b58

Observation 37d84659-5348-4166-a3cf-c3936698a787 · outbound

This paper cites Denoising diffusion probabilistic models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.321274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.321274Z digest=sha256:cd0d00eeea422f440c5013178a39fed3212463c9980aa362a2027aa25e648556

Observation 1f4363f6-381f-42b0-a7ea-d36456e50415 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Vbench: Comprehensive benchmark suite for video generative models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.415489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.415489Z digest=sha256:a753b6b5cc45bf34cb108e7c671c6b58a5fc7d25fa271b22c59a80f3f959d566

Observation d74470f6-32c6-427b-9b13-b484f376cc39 · outbound

This paper cites BK-SDM: A Lightweight, Fast, and Cheap Version of Stable Diffusion.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers BK-SDM: A Lightweight, Fast, and Cheap Version of Stable Diffusion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.528176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.528176Z digest=sha256:05c801c5c341729d557196888426a564bfcfd51a214a1d0accaa753165319e0a

Observation 2a26b896-c5e6-445e-8391-19d4a2ddce8e · outbound

This paper cites Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.611097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.611097Z digest=sha256:70685db3dc622450964a3422ae67b6f4e19ca9517de597aa674be9354076b5ce

Observation 1787af12-90c5-4b72-9f49-63638d80ea82 · outbound

This paper cites Auto-encoding variational bayes.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Auto-encoding variational bayes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.450202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:09.705790Z digest=sha256:414a5455545fea04875d4378ef4599739a6c604a78eec31f0edaee3abc24955c

Observation d75af1e5-a66f-4b6b-97b7-02db52c580e4 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:09.802048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:09.802048Z digest=sha256:64078972db486e212637b79d61928600cc1ef7f9fb066d23e6af8ac4eab11c42

Observation b2bb5ecf-4d31-43cb-8817-516329cc1819 · outbound

This paper cites Flux: Official inference repository for flux.1 models, 2024.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Flux: Official inference repository for flux.1 models, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.223842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:09.890736Z digest=sha256:184c62fcabc6e43ca8079d865aee6f370ff6e46f5e8e45386d3c29d4369a0522

Observation 9cc94751-2b32-4437-b9e8-535fa9d70074 · outbound

This paper cites Set transformer: A framework for attention-based permutation-invariant neural networks.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Set transformer: A framework for attention-based permutation-invariant neural networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.040318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:09.990647Z digest=sha256:042041d65f639030bbd6a017d4c9d08e903ee786782469e882653bced4d279ba

Observation f2caf802-0014-48e2-b3de-918c4742d692 · outbound

This paper cites Koala: Empirical lessons toward memory-efficient and fast diffusion models for text-to-image synthesis.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Koala: Empirical lessons toward memory-efficient and fast diffusion models for text-to-image synthesis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.801008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:10.054589Z digest=sha256:1a8778f10ddd0b1db306b4110c87429ff9c98b50067bda225e58e7dd66d1e336

Observation d5ca0487-ccd4-49f5-ae78-14d7505d2ada · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.117315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.117315Z digest=sha256:9b3a029b78d43df9dfc30393958536e8733c8d8c3dc0db48488b82614a050601

Observation 2a90d978-d06c-48d5-b03f-3b411c00ea1b · outbound

This paper cites Q-diffusion: Quantizing diffusion models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Q-diffusion: Quantizing diffusion models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.609258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:10.175585Z digest=sha256:30d59b7dc9e6e5efd15e3570df5302a3a1f62ead74a12f9215d8b0d8a2523165

Observation b92128ce-cfe6-4f37-9d28-195864e60914 · outbound

This paper cites Microsoft coco: Common objects in context.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Microsoft coco: Common objects in context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.217019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.217019Z digest=sha256:2486a396d193cc4e3d6f85be2066912f737c1d3789300cdb536489946bee0b28

Observation 972035e3-7f4a-4f27-92ca-ed1145ceb719 · outbound

This paper cites Flow matching for generative modeling.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Flow matching for generative modeling

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.438238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:10.273734Z digest=sha256:eebe35f57eceb11623ccafee0c12826fb536525ee667fc9e145e622341314ae2

Observation 0f3c9bf9-39e8-4cd3-8cf7-717e30be88b4 · outbound

This paper cites Pseudo numerical methods for diffusion models on manifolds.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Pseudo numerical methods for diffusion models on manifolds

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.268542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:10.339537Z digest=sha256:f7cc148148d6f977f227a692e124a86fcc89768b667a85e3da6a44e82f0142b2

Observation 0cd2cb5c-2199-4a34-a237-179c20409316 · outbound

This paper cites Alleviating distortion in image generation via multi-resolution diffusion models and time-dependent layer normalization.NeurIPS, 2024.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Alleviating distortion in image generation via multi-resolution diffusion models and time-dependent layer normalization.NeurIPS, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.093310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:10.406077Z digest=sha256:ac8bc0204a4fef26a81ba44d05c62447268b1ef92e4a89f08716e0dcc5e5f436

Observation 19652a10-b693-4a25-877e-120a4bd9af9c · outbound

This paper cites Revision: High-quality, low-cost video generation with explicit 3d physics modeling for complex motion and interaction.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Revision: High-quality, low-cost video generation with explicit 3d physics modeling for complex motion and interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.485556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.485556Z digest=sha256:b00d18afb3a74f19a2fdc5c58d7614ef78bc01a1c4110b85fca9c67142e76016

Observation 35b85fd9-8839-4faa-a332-90f6e4ea280d · outbound

This paper cites CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.540296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.540296Z digest=sha256:4a645dd2532903dc9d412eee39b8e81a61051e1e26b070fbdbb668681eae06bc

Observation f4577a30-29a3-43a2-9d6c-eab7419c672d · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.610639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.610639Z digest=sha256:bf15e6a3c5cf989dd9e9a0015a88ccc8a0104b433b8dc9ae8b90764977de14d7

Observation 4d94bc75-79ca-4c30-a430-b229abadc9d4 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.696982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.696982Z digest=sha256:abc9b5322e315e0a822884fe4190bc428841df5a85df7be3009d00f714db0e36

Observation 1dbe9e44-1332-4068-bf74-324f3776c0db · outbound

This paper cites Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.882654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:10.753989Z digest=sha256:48dfdcd7f3129a4ba557e4efacc0eeafc5fdcfee34d39e0796109f55ce1b3187

Observation f3338114-3ca9-4c04-8fe0-8ad611cde28c · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Pytorch: An imperative style, high-performance deep learning library

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.831193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.831193Z digest=sha256:1799d7784ebf25b8c1beba6d0226acae2d215488f8d422f863cd6ea1ed9ece00

Observation 86184171-4eff-455c-bce7-036aa8a98720 · outbound

This paper cites Scalable diffusion models with transformers.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Scalable diffusion models with transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.906479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.906479Z digest=sha256:a143ad6b426529c1da85d02a1dccee0f1232551df6d7c82ecc6c21889e5b965a

Observation 043dafce-3774-4c63-ac1a-3847466d5983 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Learning transferable visual models from natural language supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.970631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.970631Z digest=sha256:2c14a50426cb4e8ef79df4563532075984e60f46c5a7c55005eea547f1107b9e

Observation a275f21e-61a1-43de-8f5a-b6f5a8de9e99 · outbound

This paper cites Stand-alone self-attention in vision models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Stand-alone self-attention in vision models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.688154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:11.011275Z digest=sha256:29baeb1da172051ba832e01f74aebece6375d701b01a7375188a1649b3ccb6d4

Observation 2bbd127e-dbbb-433b-813f-e49370cc231c · outbound

This paper cites Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.114818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.114818Z digest=sha256:26a39722c2ca8746854c76073f4dd8ca9429a16d534886125f9069b281fd1d5f

Observation 9f72c516-3ad7-4c04-8ece-4ef59bbc362f · outbound

This paper cites Flowar: Scale-wise autoregressive image generation meets flow matching.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Flowar: Scale-wise autoregressive image generation meets flow matching

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.477595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:11.195760Z digest=sha256:7011f9805e23ce5da4b6dd29ee6b243f83e31ffa1c440136c05921a8507f99d3

Observation 94c821a4-de10-4790-81ad-2d7123d4bd17 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers High-resolution image synthesis with latent diffusion models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.277092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.277092Z digest=sha256:afe5dc2198fc7b5fb29d94ec2120219c8ea3868fb6d1bf9e42efba13320c639b

Observation 6398b433-bfa5-4f43-8e83-f8eeeefeec59 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Progressive distillation for fast sampling of diffusion models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.272853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:11.359332Z digest=sha256:9380488ba14436b9566ceb89e2252afd424c5b1101a977a44c52c0198350361b

Observation 321737fb-993e-48a6-a12b-32fbfd077602 · outbound

This paper cites Post-training quantization on diffusion models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Post-training quantization on diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.032194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:11.455816Z digest=sha256:4c1bcb78748e4edcd13f530cd29496f70f7f1d6dbe48bf045c66027a6e91de36

Observation c24fecb9-67b7-4ab2-a03d-d5e391791211 · outbound

This paper cites Deeply Supervised Flow-Based Generative Models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Deeply Supervised Flow-Based Generative Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.513331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.513331Z digest=sha256:6019cdede9a8980c56b9882b175bd06550f1c9ab6b71d708d559c936f6990c7c

Observation 19ff55c3-075f-45ca-bbf3-7053f4ce474c · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Score-Based Generative Modeling through Stochastic Differential Equations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.607502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.607502Z digest=sha256:64bda4fb917fd5c2f1aa8825931df0ec574bd8e4ca9d5438e39c6c5dcd63aa43

Observation 5ab29220-e2f7-49e4-a5c8-71683cfec651 · outbound

This paper cites Attention is all you need.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Attention is all you need

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.705712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.705712Z digest=sha256:2e73970cb37fc09eec4e43073eb2732125611946d84299ab9d2fde3c70d0470f

Observation 0a5d231c-f14c-4c32-ad4f-523e11f1d95d · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Wan: Open and Advanced Large-Scale Video Generative Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.772789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.772789Z digest=sha256:314cd24d41013c40daed30f2e1d3ba0d0ac0674d1f0df269691b21529a3db3a2

Observation 335f927a-0a5c-4a67-a57f-1cefeef14486 · outbound

This paper cites Axial- deeplab: Stand-alone axial-attention for panoptic segmentation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Axial- deeplab: Stand-alone axial-attention for panoptic segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:14.822785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:11.883552Z digest=sha256:d5999826128207d57e71a076b62bd054506c6b5344c0a3f5cff6da13588f2575

Observation 38e89115-e881-4974-8c2e-638c6bcb5b66 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Linformer: Self-Attention with Linear Complexity

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.991320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.991320Z digest=sha256:35c304a85858e6df52d0457bf9516ae05262714ec084d1110a4514b6383ec26b

Observation 26888ff7-48f0-4f2b-a890-68aa6548ffe5 · outbound

This paper cites PPT: Token Pruning and Pooling for Efficient Vision Transformers.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers PPT: Token Pruning and Pooling for Efficient Vision Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.123223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.123223Z digest=sha256:cb2e397b7e3edf2060364dbb25600d61d95075619ee1d8a6d6d85b7a5b64f93a

Observation 7d171f7b-9679-4dd8-98b7-0f4657292cc7 · outbound

This paper cites Revealing the dark secrets of masked image modeling.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Revealing the dark secrets of masked image modeling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:14.625332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:12.199902Z digest=sha256:534cae8dea2bc1f59ea298c9a6b8c8083812f40d2879e0057c965d47a92c3491

Observation 21e1b526-6c24-4eb8-9732-6fa63313b9fa · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:14.439072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:12.267943Z digest=sha256:72a5c3863819c2758e1f7e7d76ef2c16d855f76b9be5bf5f4491f913d44f1b52

Observation 20031a9d-1774-46c2-9f24-63e9b5f8b4bd · outbound

This paper cites 1.58-bit FLUX.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers 1.58-bit FLUX

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.353042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.353042Z digest=sha256:3826fc2e084ac1f13e5e4be770fa80a754761e0fc1720b83cecc393f4d772932

Observation 5ca922d8-374f-4f87-84ea-6bbce513e84b · outbound

This paper cites Randomized Autoregressive Visual Generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Randomized Autoregressive Visual Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.464495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.464495Z digest=sha256:564dec1721c0425aada80ec9e441ca462fdcc676f15c7bc4bfa685c4204fdb6c

Observation 7620ea40-4ef0-47f6-b852-c29001f5d12e · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers An image is worth 32 tokens for reconstruction and generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:14.221647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:12.577214Z digest=sha256:1df237a940b597db4a2d227c72eb99b7be9a720c7743bb8b4cb2a750de572fdb

Observation 57186a5d-b8a1-4890-9d3e-67454e39df8b · outbound

This paper cites Ditfastattn: Attention compression for diffusion transformer models.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Ditfastattn: Attention compression for diffusion transformer models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:14.001466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:12.642624Z digest=sha256:c6a8276d8d26c452b9f0f1677c8eb9f6e27c3b67d8c08782db6c1d428de74859

Observation 09a73369-39d8-4ca8-9def-216df835b5ed · outbound

This paper cites Big bird: Transformers for longer sequences.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Big bird: Transformers for longer sequences

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:13.837906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:33:12.738393Z digest=sha256:1b9bf51221d0362bc0c331434f6f6ef34513525be05eae266c8dc75a646e1295

Observation 5f1ba8ba-8080-4e44-b3d1-2a3d53862c32 · outbound

This paper cites Fast Video Generation with Sliding Tile Attention.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Fast Video Generation with Sliding Tile Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.833852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.833852Z digest=sha256:444e58758f1dc6dd2bca82f1bca6367f69c9e31d61344f227b3fb9139233d0b9

Pith citing papers

Observation 4f4f36e8-d914-4ce3-ad57-e1bc149db04c · inbound

Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens cites this paper.

Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:59.571181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:59.571181Z digest=sha256:b5d4ecd732b13ff24330b419ae7c7db03d1735ab0886634151b01fee613b86ee

Observation e35993a9-d991-427b-84a6-22aa4778c62d · inbound

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking cites this paper.

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:10:05.851830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T15:10:00.146694Z digest=sha256:2557ff5c4634c8117e38a46334414796b90820d55f69009f759812206e3feee4

Observation 43ca65b5-c884-476c-89eb-9a67ba031992 · inbound

A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens cites this paper.

A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:50.174151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:19:17.796377Z digest=sha256:54c4e7df7381950e54b31871ce39527825136a3b1ee167b52134621314da9536

Observation b86c1ea3-1ce8-4a08-a0f2-0ad7cacd5cdf · inbound

Frequency-Aware Flow Matching for High-Quality Image Generation cites this paper.

Frequency-Aware Flow Matching for High-Quality Image Generation Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:00:04.105673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T10:58:05.541289Z digest=sha256:9ac51069ff4a49e6f2cf6dd35eca4375c79eb58f8b8f2632f8ec65173c5169d0

Observation a1252ef0-c61f-4411-b75a-250093626758 · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:25.769904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:5d13c1584b173be1144b1d316077232569abaa778be73ad8c573e6338bc1f90e

Observation 46388b3a-5e1a-407a-8fa7-2421d098f3b5 · inbound

Attention Sinks in Diffusion Transformers: A Causal Analysis cites this paper.

Attention Sinks in Diffusion Transformers: A Causal Analysis Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.044397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:18:26.883899Z digest=sha256:49625248d1855a7d6815126483d1880e5d8d5def67abe63034b77613a7de6788

Observation a5fff18d-56e4-433c-a0a9-601a2ebc0192 · inbound

Attention Sinks in Diffusion Transformers: A Causal Analysis cites this paper.

Attention Sinks in Diffusion Transformers: A Causal Analysis Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.283147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T05:59:16.233848Z digest=sha256:6d8df37e95507523acf502434c3f9052043ee175d0961b78fc1f55eaa2b61144

Observation 5a6911fb-01fb-4563-a1ba-bcbc644f8972 · inbound

Attention Sinks in Diffusion Transformers: A Causal Analysis cites this paper.

Attention Sinks in Diffusion Transformers: A Causal Analysis Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:46.111412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T22:47:11.360530Z digest=sha256:f7835d3576d1026bd0f30482a0bfadca38edce9a4a21f24b044d15202dfd9c32

Observation 03ed245a-a299-4c3a-9bee-234a7289ec20 · inbound

ACID: Adaptive Caching for vIDeo generation cites this paper.

ACID: Adaptive Caching for vIDeo generation Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:40:16.892454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:40:16.892454Z digest=sha256:94fda63c85d5d3b11fe7681cfdf0e3bdcc7e10346415995b64fc7eefb815c1d6

Observation 8e5acc8b-c681-4a6e-85f2-429ec8ba351e · inbound

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention cites this paper.

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:07.769382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:07.769382Z digest=sha256:1ddd3694591dfc6e53664cc53aa783643d9d3ff3bfe2f5b378f04fad66779712