Pith. sign in

Paper Citation Record · LEDGER

Playing with Transformer at 30+ FPS via Next-Frame Diffusion

As of 8 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 7 inbound Pith citation observations for arXiv:2506.01380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01380 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:46.876113Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:43:37.510074Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:54:57.961666Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 244b70f9-ebf8-4ce1-8757-fca7a6da178c · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:41.669315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:41.669315Z digest=sha256:5ea5c03042086f3d57cf71a0a604239ba2beec28d6338235b975f03af83e18b1

Observation fb49a71e-b600-4a27-bc94-5c19ddb3f2af · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion for world modeling: Visual details matter in atari

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:41.724851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:41.724851Z digest=sha256:6f1ed436a7488747e3852ee73234dc51eb0d98180bb6b1b3f35326eb2f1324cd

Observation 01b31326-a147-4769-abcc-850662967b95 · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffusion language models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Block diffusion: Interpolating between autoregressive and diffusion language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:50.211033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:41.780871Z digest=sha256:10630d146272b536e9fa3d535c68475ce86ac883a2a3efbd6798c731486996c7

Observation 5052c2e3-eb6e-4d5b-ad36-2245576cd7ff · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.982602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:41.831486Z digest=sha256:455c9158a3aee4998923e479d86c464df340b8951598dcdb6f9de599e81e181e

Observation ed5dc038-8008-495b-86fb-10591f58257b · outbound

This paper cites Language models are few-shot learners.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.811589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:41.879455Z digest=sha256:20a4dc3e9b47579b52ee57e0f6973bc092660b07db2b1886ce674db580d00098

Observation 0ec95660-fa27-44ff-895f-bd061699fefa · outbound

This paper cites Genie: Generative interactive environments.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Genie: Generative interactive environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:41.967656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:41.967656Z digest=sha256:af5d8327aaaa10099a3932b063204bf4280c918e9602879fc4be89ff66087a27

Observation d699f9dc-f578-41f3-9d70-722e4d6eaaa6 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.051150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.051150Z digest=sha256:0871bf45f007812fcfdb1de765a6a039a1782fcc73cbc2b9dd125d8196513675

Observation 4419b65c-9ab0-42a4-971f-c21882a4a9c5 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.114414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.114414Z digest=sha256:fc73a926d5a078ccc577a64ffe85cee8f67fc2fd016f3fabe9116dce7e6ccafd

Observation 4f067ec0-c780-4392-ad75-32d7e4a8449d · outbound

This paper cites Sana-sprint: One-step diffusion with continuous-time consistency distillation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Sana-sprint: One-step diffusion with continuous-time consistency distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.153720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.153720Z digest=sha256:2e13832ee2a3277e8f68f2ea945a16076f463af943cf2755343cc88b9434d972

Observation 794491b1-9a61-44bd-8399-e6d5650962a6 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.207829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.207829Z digest=sha256:acbd72282b3656fd67e1e29a2d4ea8e8166730d38b8e5906d35053ee86a5b258

Observation 55a430e0-325c-4524-8c13-2f8e5b2eb418 · outbound

This paper cites CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.283209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.283209Z digest=sha256:dee16f4d6d027338a26dff63922a7a3c105da395e3d1ea0b1dcd9b12dd7aa721

Observation 05958ecd-d84a-4b4a-8956-9ace4dfdf4f1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion policy: Visuomotor policy learning via action diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.338890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.338890Z digest=sha256:6fb05eee69eb7cf167724dc26a34226ec8536766d152fa8521fcb8fd4fe52a9c

Observation 614d844f-25fa-4288-b6c3-c257825fa05b · outbound

This paper cites Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.399750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.399750Z digest=sha256:3d9f02b1f7eb39201554f6a2d4637de6b85878e210d6c7d8e93d69e6064b69ac

Observation 5a8c4109-0bdf-4b38-b94a-da12a55010be · outbound

This paper cites Accelerated Diffusion Models via Speculative Sampling.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Accelerated Diffusion Models via Speculative Sampling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.452683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.452683Z digest=sha256:7dd909caed5aba95c752f90571202951e83ec8339eaad65ef467304fbf76b59c

Observation 98710a6b-25bd-40cd-8280-d7ecc15005d5 · outbound

This paper cites Oasis: A universe in a transformer.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Oasis: A universe in a transformer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.594996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:42.512389Z digest=sha256:baf56ae3bbade1d8de9257f17cebb59a5ee1d2a6cfb1bfb144af74e0865cdd6f

Observation 555ffc70-71d1-4176-8a2d-8315aef0adf2 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion models beat gans on image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.552838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.552838Z digest=sha256:36f2251a4796a9ea244269b835f1ba62d91b2ef37aeb08cbd1d0c92aead961bd

Observation 7e9145c7-2721-4514-a17c-e447792d10fb · outbound

This paper cites Learning universal policies via text-guided video generation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Learning universal policies via text-guided video generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.628427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.628427Z digest=sha256:d48ed7d1e7054ae7974d1605e92637826df95c3b858706bc4ef77271f2f97bf8

Observation 896a152c-dc74-46c7-9765-1834df733d95 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.677994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.677994Z digest=sha256:1f266cdcc363f36ccf4c7d8d248a972ff1050a6ec1123011aa55e05535d7d5c0

Observation 9b5c2a66-effd-4e70-91f4-e9e926a1d3d7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Taming transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.727555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.727555Z digest=sha256:8dd525f566fe275e89a170d6a494a755122b2e8f0a31340a0a8aab2cf678d2b6

Observation 25872947-c2bd-4564-9b4b-80de0aa5fc1c · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Vista: A generalizable driving world model with high fidelity and versatile controllability

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.346082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:42.792368Z digest=sha256:67a4a98b265b1152aefe154eea204cc1f4cc5de56205af2624c144a5521ec985

Observation 4ba8dbc8-c918-4eb4-b0d6-ad4fea3cc9e3 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.840907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.840907Z digest=sha256:fb72074298f3228913664e32751042fc5e77208dbb7558da489acf3f199653b6

Observation 92d671bb-a3a7-4acd-bea1-766cfceef618 · outbound

This paper cites MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.913195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.913195Z digest=sha256:dbea2fd19e6a530e534263a7a4a7ca5119e677b85dc44e2b9c9b0aedd6d76257

Observation 84a16ecb-f50c-4525-9fb5-88aa4bc0fa82 · outbound

This paper cites World Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion World Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.962948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.962948Z digest=sha256:14a3eb0fd1aabf20422b3b7d1a7e3ac298e44c51b94c9b47b94f21d8b024d8b5

Observation a041ef9d-72a0-4fe4-b8e4-6ee3aa4b97e1 · outbound

This paper cites Mastering Diverse Domains through World Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Mastering Diverse Domains through World Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.016848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.016848Z digest=sha256:88d290b2de226e2c10cce817392e0fd74551a18360bae5300074d37514dccc06

Observation 27b1f4b2-18df-4db5-91f1-8708ce935c74 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Imagen Video: High Definition Video Generation with Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.063242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.063242Z digest=sha256:a150d85f57dd4070b08b1f0966e98f98dfd105af3b30ed479c997a6770092e87

Observation 9b507f31-7b1a-4daf-ada6-45498c6f5a7b · outbound

This paper cites Denoising diffusion probabilistic models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Denoising diffusion probabilistic models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.232940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:43.103930Z digest=sha256:fc79cef1cd564ffa4b64e5f1f76ae2e0059105a6e497fa09764a2fb0d70ba864

Observation 8022c61b-1d09-4c0b-abe9-1a3830976dc7 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion GAIA-1: A Generative World Model for Autonomous Driving

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.159204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.159204Z digest=sha256:32ccd2b9d69d6ec0d96cea43436ae040aed706e543d2e22d103430f75f3f288f

Observation 7eb626a3-f55d-4c06-b5b6-288a638f1d38 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Elucidating the design space of diffusion-based generative models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.244093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.244093Z digest=sha256:b4b34fc7544a96c873bbaca664b6ec7f90dac3a9b2533050c4595df7ad9e54d3

Observation 9b8a44f7-e077-45ec-9dc4-9a958dc95356 · outbound

This paper cites Learning to simulate dynamic environments with gamegan.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Learning to simulate dynamic environments with gamegan

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:49.072910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:43.295706Z digest=sha256:2406e853154c3821c25d57f563b47943055a4f4e4d1f7ff946310552176a769a

Observation f94f38e9-fc56-4ca3-9d86-bed4c5f7feca · outbound

This paper cites Adam: A method for stochastic optimization.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Adam: A method for stochastic optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.344193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.344193Z digest=sha256:6a0c7d177d006cc3d23b1cadb1ce2d410934ae2cb9940fc43e7ba6f91bce9b01

Observation d457942c-d70c-44ed-a610-bb856991f381 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Videopoet: A large language model for zero-shot video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.865142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:43.410746Z digest=sha256:9b9eecc8a42669cb742cf747c423e47ee4f6a0d15d27cfa3badd4e0c53fc6d82

Observation 0f3acd39-c673-4d9f-bd75-efd65314df84 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.465981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.465981Z digest=sha256:69f864592128c1a8d8b335c8be9463b485aff1977a0ef4b8740863bea71ea3ca

Observation 831cc28a-985a-49d5-aeb2-ad88c629ebdd · outbound

This paper cites Fast inference from transformers via speculative decoding.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Fast inference from transformers via speculative decoding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.547324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.547324Z digest=sha256:b41330a43d2e182d561aa104889fe36581b72707ba358094a8dba18f7cbe6667

Observation 9a681093-c219-414d-bba4-996fb72b1165 · outbound

This paper cites Autoregressive image generation without vector quantization.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Autoregressive image generation without vector quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.588354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.588354Z digest=sha256:e29b28b8de423251c2c4cb1b2916b35a5954d1fe9e911fa31391d348cfce8ef3

Observation 5a430c13-2701-412e-906a-22546058089c · outbound

This paper cites Flow matching for generative modeling.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Flow matching for generative modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.645431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.645431Z digest=sha256:6015ef9e977308fa4a4f435c1950fc67187f4b8c6ce1ca700bdc6dce85391240

Observation cb3697ee-4b98-462b-a585-1e7d43bee3cc · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.695257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.695257Z digest=sha256:436efdeb7edcc417049c55478a31cbf3ed144f37cd83116b5ce77995ee7568ed

Observation a127a3b0-bf61-42ad-bc05-c61bcf86d895 · outbound

This paper cites Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.757089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.757089Z digest=sha256:cea0850f7517bfcede8a188fd0f0ecbad1e0fcb6effb66c4fe77fcdfa9a5c994

Observation 064070d2-898e-499e-9c37-14aa8aba7101 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.826343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.826343Z digest=sha256:2666fdc8f69c85afb39e38d360b1a28847e75b0b1c3a8052c8b2077e562a4ab8

Observation 6c9bea53-8c1e-4823-9143-fe3bebe56106 · outbound

This paper cites an unresolved cited work.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.894267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.894267Z digest=sha256:65c941c542a4f009ac82f954ffa1f7c356a2f5fcff13f1d61fcf85e444824b73

Observation 50e4430f-ac5a-49a2-a0ec-fcb87473229b · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Pytorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:43.972665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:43.972665Z digest=sha256:2293a55391e1c6fd1fc8366d9615a680681573aec79ad24749cbe64f52746a49

Observation 997a0c0e-7b37-40d6-a318-66fe5816d0a3 · outbound

This paper cites Scalable diffusion models with transformers.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.051513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.051513Z digest=sha256:92f47e128f9069b664e1a289eced5bebcf75dad3040c482c831df6237c414532

Observation 36275562-bb18-48f9-bcf3-fcbefeaf1c6f · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Movie Gen: A Cast of Media Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.157194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.157194Z digest=sha256:54e0f6e0490c47bea7e61d201641783ff8002bbfeeaee794ae47c992a17198cd

Observation d3a2258f-eda7-4680-9303-375ca7af16d0 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion U-net: Convolutional networks for biomedical image segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.282569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.282569Z digest=sha256:0a99020aeb7bb4d8feadd972bc6257999141ab015f5a6a09d190f5b300dbdbe7

Observation 758a92ba-c5cf-4ec5-bd22-2dad472cd609 · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.373058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.373058Z digest=sha256:6a26acd9858ffb8b94457022329a3d321a546ed820948a25b9e7908daf9d4c6f

Observation 7ea15702-e171-45c5-b4bd-5e1ffcb05278 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Progressive Distillation for Fast Sampling of Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.482801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.482801Z digest=sha256:c9f404790b3f526b29d8beb4162dbc9df848a427da92d89c0a16d9804c94e4d7

Observation 3afca255-798f-4e8b-93f6-76d4ff56c7fd · outbound

This paper cites an unresolved cited work.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:49:48.625980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:44.610609Z digest=sha256:cb843e231ee314062dfee549fe6cb54fb543d683e8f662cfb5ab9d145cfe4f03

Observation 9014dacf-a4df-4dd1-abe8-536867e34a55 · outbound

This paper cites Mastering atari, go, chess and shogi by planning with a learned model.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Mastering atari, go, chess and shogi by planning with a learned model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.715413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.715413Z digest=sha256:19d9830c948a56d62088324be5acd0d307767d39cf8f4407721e4aa39ecfd71d

Observation 4139b3eb-4354-4a14-897b-c21ada20b60c · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.809000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.809000Z digest=sha256:2222d1fd1c0e9a0b806517bde3ab1912beb10def0a40bb920bf62a3a5f5e3085

Observation 669706dd-e7f9-4c9b-a152-c74416627bd0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Denoising Diffusion Implicit Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.911941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.911941Z digest=sha256:c93bd44762b31acb2ff56d819f4370e2269a0f2f71f7872b8440978b029a4744

Observation 4b69d8d2-a957-404b-bce3-5b9631ea15e5 · outbound

This paper cites Consistency models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Consistency models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.039508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.039508Z digest=sha256:37b26505776888371746eb7a06f9ba07aac966778c55ddda793a0087d8c5f15d

Observation ef69d201-f105-4960-bf9c-554ac4c7e766 · outbound

This paper cites VidTok: A Versatile and Open-Source Video Tokenizer.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion VidTok: A Versatile and Open-Source Video Tokenizer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.159313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.159313Z digest=sha256:c9c513ff34f7feebf9726edd4bbd6e666611dfc31a438b75c12dfffd7d5adc84

Observation 82f4af78-8928-40ad-8d5b-fe5b1a6d6a93 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.258191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.258191Z digest=sha256:d598d041c4a705cf6092f008965339c621d329d96a7a790d9c8504dda3242c0f

Observation 8df5ef5d-437e-44bd-80e9-3c2672772843 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.338761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.338761Z digest=sha256:05515a84d2176286c506a74228709f77fb396089c8a4c9b108a82e03e977727f

Observation a15bba91-5386-4aaa-bd82-07dc7b2dfcff · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Diffusion Models Are Real-Time Game Engines

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.459923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.459923Z digest=sha256:b20234fc16ba207cc82b37fede592bd079bd43a8048918604ab08ef49bfdb35b

Observation a71b5ed8-f426-4118-bc98-ee61850d411f · outbound

This paper cites SparseDM: Toward Sparse Efficient Diffusion Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion SparseDM: Toward Sparse Efficient Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.587972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.587972Z digest=sha256:09c7dd1bbd8ca6e873cf3197a9d61ca4e42bac770c6b329dd4702ebd6e7fbd91

Observation de718956-e0a8-4e10-ac15-e6ff202e9df5 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Image quality assessment: from error visibility to structural similarity

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.688193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.688193Z digest=sha256:af8a2ae344f97153fe45922c4ad2e894b8622169cf544c8fcebc9b26c3fa3845

Observation 944b68fc-c7cf-41f4-a0b2-7e15d97f381f · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion ivideogpt: Interactive videogpts are scalable world models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.793677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.793677Z digest=sha256:354579cc2495dc00537b9445cb8b687a11f35ea2862da8d5ee2acaa0e56e68c6

Observation 5867bbd9-0c42-4daa-9eea-7c3cdbca4bb9 · outbound

This paper cites Structured 3d latents for scalable and versatile 3d generation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Structured 3d latents for scalable and versatile 3d generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.368033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:45.891921Z digest=sha256:a7d7827ae9a7b9091e2cf28d0b4c8fa0a72dbb1bdb4c1530aeb0daa493bbd9ab

Observation 006165f0-d22d-45b4-9a01-8dac9f634583 · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.038383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.038383Z digest=sha256:573f0570e5432c5ada45e59f2d675c39e366383576b62e0beffd023df7066445

Observation 0e869d7d-1207-4229-9ba0-c76a48058a43 · outbound

This paper cites Learning interactive real-world simulators.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Learning interactive real-world simulators

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.094491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.094491Z digest=sha256:c63cc915de06b78f3d579915631951418e9f78fa5b9cfad0a99bd60dc5f12a83

Observation 76a209fa-6616-4f33-a3b4-a610861537af · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.205356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.205356Z digest=sha256:6c178fb9a88226a714869bde2861aa6064e2418afa6f8d713a5579610554fa7f

Observation 3679763c-168a-4206-98dc-f77eea973d1c · outbound

This paper cites One-step diffusion with distribution matching distillation.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion One-step diffusion with distribution matching distillation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.325346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.325346Z digest=sha256:c8dd3bd87a0f1dfbad93d991f7b7532e160d18562ca5c6ab6e5fc0507d63404a

Observation 595a32b9-49df-4289-9282-8ed39999edc2 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion From slow bidirectional to fast autoregressive video diffusion models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.430341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.430341Z digest=sha256:04acf2c0024783d73aeb3e54918fae88d353136c0b455210b4e546c81cdcebb2

Observation fdf71b7a-fb16-4023-b770-e1a77b564d81 · outbound

This paper cites Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.509630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.509630Z digest=sha256:2681707680346b07d6ed28549a6510cc7e4440480cab47a5c76190b93a7c97cd

Observation 3cf5599f-d2c9-4bae-95ac-a89b1cab702f · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.182571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:46.632889Z digest=sha256:1a1c24886a7315e0cf59cb7645aaa65c199b33a4698c8a02b3b812ea5bbde35c

Observation 9653b660-e1e2-4245-9f08-35784f3d887a · outbound

This paper cites Drivedreamer-2: Llm-enhanced world models for diverse driving video gen- eration.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Drivedreamer-2: Llm-enhanced world models for diverse driving video gen- eration

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:48.009638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:46.727368Z digest=sha256:fe55821241d01f8a6f42dd337afac9ba28b6eee964b9cd9e56640f25b1091cef

Observation d75d06a7-22a2-4eef-add7-1fb68169bf7b · outbound

This paper cites Genad: Gen- erative end-to-end autonomous driving.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Genad: Gen- erative end-to-end autonomous driving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:46.822591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:46.822591Z digest=sha256:c428ceff360cf755f479a0b349488fb7c43e166e1edf269dd37ecd99f8d769ae

Observation 2bba489d-707c-40e4-b46a-73d633d75228 · outbound

This paper cites "" model : Distilled NFD + model vid : Input video tensor act : Action sequence tensor.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion "" model : Distilled NFD + model vid : Input video tensor act : Action sequence tensor

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:47.872390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:46.876113Z digest=sha256:0361235b9b3a050ff7f0936fe4527d7ee4aa9981e1381b2eb05310b4d61e8ce6

Pith citing papers

Observation 85ed9549-5a85-4e0b-9220-f9468c6cb2a5 · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:06.579621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:703ebf00b44313c8f27c0f88f49ca0b42d211de67d0c7ce97ed0ddc0667a7118

Observation 6fe79ce5-4e09-44a6-9f1a-f37ed8288403 · inbound

Matrix-game 2.0: An open-source real-time and streaming interactive world model cites this paper.

Matrix-game 2.0: An open-source real-time and streaming interactive world model Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:36:53.261124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:36:31.044743Z digest=sha256:559edd6660c8571a7e73ce2981a6d64f430c2ab58a2fcf13d5fb1910056d582e

Observation 2b2e644a-d463-4e90-9339-514e08ab82df · inbound

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation cites this paper.

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:43:37.510074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:43:37.510074Z digest=sha256:a7ea13fff273833075094b2b8dc62576823bc24d9d5194c076ed4a0d77701ead

Observation 73a3b47d-225f-4037-bc31-628ccd0cdf46 · inbound

Envisioning the Future, One Step at a Time cites this paper.

Envisioning the Future, One Step at a Time Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:05.951388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:31:41.904925Z digest=sha256:ac3238e4b8a37aaa0e023056f6571ed43a414746679cbaef61dd915d6ad73268

Observation 63b7fccf-b51d-41fe-a17e-679b03228f9c · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:26.144191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:f3c4bf88192c6e3409561083521fe23a8c1d9ed11fe3d99e30be629693a38697

Observation 7f4661b4-163d-45c9-a1d4-182b8e7013ae · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.422283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:657a477ca9b61d3212b9d9f9dc5a590f8af39e680d901c90d55fe7d11bf87a44

Observation 304a5c1b-a27a-4192-b7c2-c58efde9de4b · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playing with Transformer at 30+ FPS via Next-Frame Diffusion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.963353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:aefd0e0802b8ab89eb15677a7379c0f70dbf32408df09ca03bb82c7505dece5e