Pith. sign in

Paper Citation Record · LEDGER

Muse: Text-To-Image Generation via Masked Generative Transformers

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2301.00704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.00704 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:24.622157Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

119
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9232872e-81f2-4cf9-80d4-dadf982e014b · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.586456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:d6cdc449aba1e81dd62415ecf559d86327ed123d9866a4ea9ec0fdf5932083bf

Observation 0f796b13-7c6f-4551-8c62-4e22c30ac520 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:57.973110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:3ee60b22133c325417a4104acc23d4a30ad7ac7b0019a7789525f2b2fd60eb30

Observation a694b0ae-382a-41c6-b127-2f87bba00510 · inbound

Finite Scalar Quantization: VQ-VAE Made Simple cites this paper.

Finite Scalar Quantization: VQ-VAE Made Simple Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:34:13.478920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T18:34:13.439993Z digest=sha256:d8cc74d6056676016e80e44de21806a0532c4ebae3dbff16c869ad554fd8b62b

Observation c77dcd03-48dd-492f-9de5-fdf451d90738 · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:15:18.633063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:ff539390b0584d01ad00709aaff961c83469faa4302ee394cd593e7b20728ac7

Observation a07e5ce6-1751-4c76-9b09-d8f18e2bd385 · inbound

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation cites this paper.

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:40:44.012282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:40:43.956642Z digest=sha256:d9477a8c20c2b55aa96722bcbaa3e310ade0f54df02661dc8234b71b22b4cbda

Observation c40f8024-858c-4265-a131-657112c79235 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.648073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:3928020eb72ad89499dc1726eae57e97eb897e5e2cd3040773566cf271c2a368

Observation 5b3eb7f1-a16f-4980-8df3-66b9a26b42c7 · inbound

Controllable Image Generation with Composed Parallel Token Prediction cites this paper.

Controllable Image Generation with Composed Parallel Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:53:40.560109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T00:53:36.830943Z digest=sha256:24e69986af999ca721fb53cbfd01288ea13b7ac151cabb2ca85934c95fbf481d

Observation 53a8d595-7b89-4e86-9456-b2e176e8dd26 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:09:16.842300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:dfdd6bf56d1632891b72612f9e855ee0f4b457853dde5e7f087aeb9a8095bbec

Observation 131f433b-66db-4afd-a722-86e083c996f1 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:03:33.528653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:016d03e6c995095d6faa211c0962b8abf1ad91ddb3d926e7c4cf90a28cef0d60

Observation 80dc1c22-9878-45d2-9c2a-0b0e873d3836 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.400212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:c74b3f68dd52096d42e39628e92dc74a0700b805c1070e6be7b4fb0eb363e269

Observation 66e0d9ea-dd01-44ae-baca-dd2d23af5fab · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.770391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:818ec5f37246462f6ce67528534bd00f7e9754df5d9f65fdfd204313f42fa926

Observation e4c62dd1-97f8-43f1-b85b-6c6e62cec95a · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:42:54.605300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:92786e33ad3cb1a1b4bd0fc099b8e6086249754c1eea349cb4fc906101d34ff9

Observation 2a1143d1-7a46-454e-9f27-222bada2c9f2 · inbound

MSDformer: Multi-scale Discrete Transformer For Time Series Generation cites this paper.

MSDformer: Multi-scale Discrete Transformer For Time Series Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:44:52.815486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:44:08.891339Z digest=sha256:68095e31aeb234a27c55e0f5548e10eabf774999f684a7485b341222800896cc

Observation 25db7335-c11d-44ce-a2bb-695de5c292dd · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.789654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d03e33dc8658e6530908ebe0ad3bc07e365804cfa4f87dddf13f04e3cac9d8b7

Observation e77c68c5-45da-461b-aeba-416f1234010b · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.486525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.486525Z digest=sha256:4fa5f01915e56c35baa8c92601d0a4677fdc73a449f2c4d66af0287b9d10c551

Observation 9e99710c-9d77-427f-8763-c4e96e5bdcd7 · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:32.902329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:32.902329Z digest=sha256:f0d7f3d9c629d74ff8ebae4572bb40d78c91ef56bc524f49362aa3c6e8fb3d4b

Observation f51b3af5-c512-4de2-977c-a437fe25ccf4 · inbound

Semantics-Aware Human Motion Generation from Audio Instructions cites this paper.

Semantics-Aware Human Motion Generation from Audio Instructions Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:35.069080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:35.069080Z digest=sha256:71a09ebe21104d673aee958c4b5195c005de3218a9f2ab4ca3d39b13cf8c1d33

Observation 20a3d936-6f55-46db-9e63-6e5bf523bf11 · inbound

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling cites this paper.

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:23.180725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:23.180725Z digest=sha256:226f89373d4339740f56a97f3b74fd8e0494959cd742235626bf3c82f1a0a551

Observation 36887104-8b08-42f6-9869-9f24112971e3 · inbound

Native-Resolution Image Synthesis cites this paper.

Native-Resolution Image Synthesis Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:21.740312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:21.740312Z digest=sha256:151f519ebd9f3fc878e7fbffe181685f4b212ac3f9658c5cc2be67d94348ac0f

Observation 4b2b60c7-2598-43c8-ad79-1de3049e13ba · inbound

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation cites this paper.

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:47:52.273933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:47:52.273933Z digest=sha256:f8b804ad8801d1a88b0b8563be48fc463c1dad38ef53fc4f6a0890f59f08fb0c

Observation c3712f10-5504-470b-a3ff-dff63c008640 · inbound

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis cites this paper.

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:47.443309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:47.443309Z digest=sha256:24db930de5dc6f0d2febe809322ea7f65d4530253bf8c9e1594ac18251515389

Observation c9c346ec-e6fd-49ab-adf8-2a271f52af36 · inbound

MapBERT: Bitwise Masked Modeling for Real-Time Semantic Mapping Generation cites this paper.

MapBERT: Bitwise Masked Modeling for Real-Time Semantic Mapping Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:23.832417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:23.832417Z digest=sha256:86a7a2f47219d4fd303cd6199bf7e6a7e870e48cd4cd443262946c2fdbdfb90b

Observation ff38e35f-58a8-4028-8e43-25e4f3c590c4 · inbound

Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model cites this paper.

Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:24.622157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:24.622157Z digest=sha256:9f6597b83ac15bcf52392c9e7948383009bdb479ffcec4c8fa397499b6b99442

Observation 93898f86-93a1-4007-bbe9-fd4ddcd0c70e · inbound

MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention cites this paper.

MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:36.281876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:36.281876Z digest=sha256:9723501fbfff264f4c41671d56468b8448b6ce583a4df2968dd36df9b6bdb40e

Observation 7f962a37-9a2e-400b-9a7d-5a0dbce28b9d · inbound

Learning golf swing signatures from a single wrist-worn inertial sensor cites this paper.

Learning golf swing signatures from a single wrist-worn inertial sensor Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:33.442154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:33.442154Z digest=sha256:3e36d6fdac89a316f18542580464646de3bb5d4c077ef27d87ca568bbcd946d5

Observation e8f6cd43-d7dc-455d-a05e-11b2a34633df · inbound

MVGBench: Comprehensive Benchmark for Multi-view Generation Models cites this paper.

MVGBench: Comprehensive Benchmark for Multi-view Generation Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:09.670357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:09.670357Z digest=sha256:a2895d3ea0c2d726cc8c6f562227c5eeab459827075d45ef461340490bfa79b2

Observation 53d434c6-451e-4728-bb7a-2559e37af9e4 · inbound

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? cites this paper.

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:05.800608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:05.800608Z digest=sha256:5ee06f5498ccbf122da6063b5b148585fc829a377ce997fcef0d05dfe1e45369

Observation dc6d1286-0035-4aa1-9ff1-5e25eab8a452 · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:53.858179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:53.858179Z digest=sha256:81b1b380a2162870d0084bc6e1fe25671920934ee3d67fb294d5781b62e2822e

Observation 0757d32b-5882-4aac-a7ab-e83195aef7f6 · inbound

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer cites this paper.

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:23.621281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:23.621281Z digest=sha256:3d56d45f052dce898eb38d0fec3ad3f89998927c709a944440eb7a42b118fc66

Observation ed205e6c-4d82-4db4-bc82-fb3e72b207a6 · inbound

Room Impulse Response Generation Conditioned on Acoustic Parameters cites this paper.

Room Impulse Response Generation Conditioned on Acoustic Parameters Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:28.623018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:58:28.623018Z digest=sha256:36c00fe1b8fc9f1d166bfb9081e125fcfb492d08f6821a2aed0c85c1db74a878

Observation 07764351-64d3-4d07-9710-39ec82cf5df2 · inbound

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing cites this paper.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:29.960911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:29.960911Z digest=sha256:b3949d5e9bd6d3bad9ee8e7c191abde72e1932e82fce708ddea0a8c815d403b0

Observation 3469a9a0-887f-4059-ad72-6d7dc63615e5 · inbound

Learning neuro-symbolic convergent term rewriting systems cites this paper.

Learning neuro-symbolic convergent term rewriting systems Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:28:45.675804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:28:45.675804Z digest=sha256:efc2664a5eba3a6d91ceb59abefa2dff6d26d1a23bd933cb9637da14f3e7765a

Observation 7235f01d-ba09-49de-a38a-6496680f10d2 · inbound

Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model cites this paper.

Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:32.077010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:32.077010Z digest=sha256:6b0dd684321b8e370de376a9ac5d061c7413db459fd3141f693f09b8894a8c17

Observation c0233a12-2300-4dcf-8ad9-faf649f90181 · inbound

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy cites this paper.

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T18:37:17.815731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:37:17.815731Z digest=sha256:2ee4c9b21c2ec03ae3a669540337ab7b1979286755e893d5bcf59a401c27baa1

Observation 03f319a0-c123-4471-9679-1616fb66726a · inbound

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking cites this paper.

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:43:49.262777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:43:49.262777Z digest=sha256:7254efb723681402588429dceec26a3022a6d6fdb749f76b6c67aba416e51647

Observation a3eeabf2-40d5-4655-b1f5-f0bdfb7aee7c · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:34.022747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:34.022747Z digest=sha256:c0a0ea7b7ad60f40fe851e7b50f94a06656bfa1e6935593ac46eb708b1c55615

Observation 5b9524c7-760c-40a4-b39b-46c5de90dd0f · inbound

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction cites this paper.

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T11:08:02.386459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:08:02.386459Z digest=sha256:19baf4d0e97624185832d071d3c1131daf0334bb5fb063c0a75841b524cae840

Observation 71063a29-1ae8-45a6-87fa-202e586ea3d0 · inbound

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation cites this paper.

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:25.700535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:25.700535Z digest=sha256:fd9b36bed04c82207a8070ce6d254a20a0614a827740fc59c671f66cd0428cd2

Observation ab85a340-a5c4-4404-a48d-c2a01f7d38e5 · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.725029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.725029Z digest=sha256:45b774df2971f8ce4711f03f9541f2bda7c59d12a77870a327d4a75e8bff757d

Observation e823a537-1fdc-4c00-bb45-09ecb3b07f47 · inbound

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation cites this paper.

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T08:16:22.354776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:16:22.354776Z digest=sha256:e49973efc4e9b6a5a61b495ac05ddc5b420d2f15505f34acc795ff3fdac7f202

Observation e8362087-437c-4cbf-a03b-83b1d4e73e9b · inbound

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards cites this paper.

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:28.773088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:40:18.940889Z digest=sha256:3fea57e48a8951a65cf0a4be73afbbcc772bede8bbd20fa583e87c6432a94c80

Observation a83873ff-e267-4133-8856-ff170a96f15b · inbound

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation cites this paper.

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:06.998166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T12:15:57.019720Z digest=sha256:7ded6fb04dd131603ce05fde7e578a12f45f782bba4143944e787d2f86336f4d

Observation ef4784c5-9a87-4b1d-9891-5fb6a6b4e38f · inbound

Controllable Image Generation with Composed Parallel Token Prediction cites this paper.

Controllable Image Generation with Composed Parallel Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.709020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:33:03.066564Z digest=sha256:848ddd0d50095b7fd2c8a785c2c6e91978f38c7d423271831ea0321d4552fdec

Observation 14212172-9d6b-4810-9261-eae613d889da · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.840743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:32:54.235578Z digest=sha256:d89ba4beb64fd25b711b9fb2d90ce5e32a256eb786641e278a48bcd4ca2e39a8

Observation 7d97494e-1a26-412c-a643-e19525e82873 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T11:41:02.733153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-05T11:32:38.636335Z digest=sha256:162d37606ad6e057d90506aaac5d6ec1cf1f2517fa671170cfd53613153daf2b

Observation 83ed988c-ef11-45c2-a602-6aad93d637dc · inbound

Coupling Models for One-Step Discrete Generation cites this paper.

Coupling Models for One-Step Discrete Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:00:55.161047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:58:10.909499Z digest=sha256:bbf5389f3d6624f9eccab4060c9b21e20ebdd5bcb71d8e9d75fd9a912e98dc51

Observation 9d6ce223-0082-4686-ae56-d033d04f4f12 · inbound

VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation cites this paper.

VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:33:53.079392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:33:47.573164Z digest=sha256:da512415016561cd115a04ec0511cfc83f184e9a10b654660259003571763200

Observation 609ca211-2bb8-4d53-b578-5bbb2630e701 · inbound

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space cites this paper.

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.247806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:47:33.474443Z digest=sha256:d03bbb6d563a5ba5cdd4cd2da15351cf174b6c24c94158bce8ba9a32fd6a7f44

Observation 04ea7e06-9606-4226-9f00-13106dcbd614 · inbound

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space cites this paper.

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:00.171208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T19:13:35.794787Z digest=sha256:377f5ffa7b7b7a53d0fa52cd04e6dc61852eafca8b43b27ce92f16e849c63ecb

Observation 281bbec6-d233-4b26-9316-e46384d8f824 · inbound

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space cites this paper.

Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:01.557172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:01.557172Z digest=sha256:bd0d4d6036c14208a2046e924c8eec0390d13afd9e56571df02262e337dec9ab

Observation bc82891e-2ef2-4cef-886d-9807b393be94 · inbound

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South cites this paper.

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:08:07.000756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T07:06:57.555070Z digest=sha256:c4875d333b1c2a9c03aade69f5bfd89d84716d6765fc33f093c36a1c31595403

Observation 41295a1b-dadb-4525-8144-73c64613c3ba · inbound

SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting cites this paper.

SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:04:00.978197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:02:13.380306Z digest=sha256:a97b9487da8394e42d00f7ce45809d982e2e6afd3197222e7911992c74ebd46b

Observation 98eb01a1-65d2-4d36-b016-935ee963d365 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:27.646872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:ecb6776e114f7aeb5e40066c1da303bc84b82fbf9ee32f6e972014581bddbe0c

Observation b7772da1-d62b-4175-b2b5-9909ce03c38b · inbound

What Type of Inference is Active Inference? cites this paper.

What Type of Inference is Active Inference? Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-28T06:51:44.836825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T06:45:53.233544Z digest=sha256:b8e2cd392c21b40f998620c7432e7d865b9e3474adcea2b0b5cc86dce5632ee0

Observation 05d80a24-c1eb-495f-8d8c-cf224c716547 · inbound

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations cites this paper.

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.441534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:29:11.526106Z digest=sha256:9d316193b84c9b688301d1c089b29353e9c92931b61acf952e169b5bd58d2406

Observation 5a832f26-c88c-40e4-81b6-d31f4a969b93 · inbound

Expected Free Energy-based Planning as Variational Inference cites this paper.

Expected Free Energy-based Planning as Variational Inference Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:40:57.096595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T13:39:28.078342Z digest=sha256:fcb3cd9e985266f4bda4035a33e7edf395f4433b26b73ab89ff1e3faec935240

Observation 4e6b0e0c-9da9-4d51-bef2-c9129619da42 · inbound

Co-occurring associated retained concepts in Diffusion Unlearning cites this paper.

Co-occurring associated retained concepts in Diffusion Unlearning Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:09:57.475269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:50:58.689718Z digest=sha256:94b8206baea7aab66c97f266c5af68d91c28413840e72d90fb7f0d2af5946ece

Observation 7d4417cc-c09e-40ff-9e0e-29344603c31c · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.259206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:a3485302b00071ad3826a167eb9d74e5f0b6e693fb25a709ac1c5cee88f46741

Observation bda92eb2-a1ff-4c83-8505-db6d71b2ede1 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:fc0dd580ac9c4d295d40f1646a7d5032078ce29f0d8eaefc1eea5c5cdc259076