Pith. sign in

Paper Citation Record · LEDGER

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2506.07999.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07999 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:02.330264Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 544ef75f-082c-41c0-8402-9ca8daa351ab · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.141730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.141730Z digest=sha256:b32911fcacbc7088b4b4fdd8a6912b5010d59e176c0646b9e55b31c9fecddab5

Observation 494bde0b-a381-4826-bc7e-74cc21da2b3b · outbound

This paper cites Semantic-conditional diffusion networks for image captioning*.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Semantic-conditional diffusion networks for image captioning*

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:03.035305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.148374Z digest=sha256:e05babf341e3b421cce63d30fa6d4cd44e41341dc96962eb42d1a34e4c579b0d

Observation 7210748d-ffab-4ce4-84b6-868e679ba753 · outbound

This paper cites Metaxas, and S.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Metaxas, and S

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:03.019678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.154536Z digest=sha256:660a4ac097a6233cfa6baabc2b4fc8377b4fb6189a4ac098e4d1e73c2dc28d18

Observation 4fa29346-b78e-407d-ba57-8d494f21b8e0 · outbound

This paper cites Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.159306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.159306Z digest=sha256:dc5eb3bcded287135010adaca62b9dd5af9f8749e6c2d74c7b74cb4680b054b0

Observation 1acfa949-3647-486a-8f7a-3dcc41e3ac57 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Taming transformers for high-resolution image synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:03.003472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.164485Z digest=sha256:df8ac2d28c63c8da4adcc2c766078dbf34375bb27898939e9c9ef36b27564542

Observation a9af8c98-d79e-47ad-a980-60bd0d5c3382 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.169516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.169516Z digest=sha256:2b48a11d12787da0ad02f3be559cd5916a4c50b6678a2b9c5bc00de6fbf35a1c

Observation e296f37f-7262-4e36-9ab5-699e0e348ad8 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Cosmos World Foundation Model Platform for Physical AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.174857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.174857Z digest=sha256:e95984d0e475cb31576c5fb689c5f0cebfbf65cf9798d1811ca180e4623b6794

Observation 861fa5a2-c893-4650-8695-d3812e7ed8e9 · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.987674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.179899Z digest=sha256:b46be038bb1b30480e6dcde41dbb2d45295e77827ca81c95562626fb38338d55

Observation 40f1b01b-18b6-4358-a0fd-bcefc82a9227 · outbound

This paper cites Peebles and Saining Xie.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Peebles and Saining Xie

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.972644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.184466Z digest=sha256:2d595c4d5025dffda9220a3fdafcb1ba32c43732f8893082f7be116ad3626baf

Observation 9b1e849c-805a-457e-bf94-fd03f743576b · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.189060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.189060Z digest=sha256:bc169b72f0f6b481d7a66efa97735ecc5ad2d65635906a428305a1309344f762

Observation 68056b33-8d3e-40df-834f-f1d5b816e72a · outbound

This paper cites Announcing the flux pro finetuning api.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Announcing the flux pro finetuning api

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.956433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.195206Z digest=sha256:0251770bb52880c75c7f7f0467ffa1c3581b3494607bb849da4230e26dabf726

Observation 5f1e3e6b-773d-4c51-8945-3fe4e201cf36 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Autoregressive Image Generation without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.200117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.200117Z digest=sha256:259a74c9dda4e42494c62326bf80c2bdaf965cc97e3342a1b746ca8f3839b654

Observation c225a258-a4c0-40ad-ad33-589605391c9d · outbound

This paper cites Acdit: Interpolating autoregressive conditional modeling and diffusion transformer.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Acdit: Interpolating autoregressive conditional modeling and diffusion transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.205671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.205671Z digest=sha256:2f4eb0f38ea03e07814eb67404c9bff175c2ee64a4b519d67a11e860460c921f

Observation 6fbc3b17-ad2e-44d7-8d58-a644b733a2a1 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.210366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.210366Z digest=sha256:86bf16d18482e586d6910fd3abeac5b7ed6497b2db7bd844e883a4c0bae3929f

Observation efe7341d-83ae-420a-bc88-319e07ac5896 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.215293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.215293Z digest=sha256:c4867bf7106d278bf02d9aec9fa12e6e0ffc34674500d238d92d0a03df5be0ef

Observation a6927dc3-ef34-4fbc-beff-b841f147b555 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Denoising Diffusion Probabilistic Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.219967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.219967Z digest=sha256:473fc784238debfc32964685f73bc73b7d1a4e9b8d01bd5e9797755336dfbbac

Observation 592032d5-4885-454e-9b6a-b31df77da8b3 · outbound

This paper cites Neural discrete representation learning.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Neural discrete representation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.938173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.224761Z digest=sha256:e7387c28b7ba964536e4decae477bbf79873a42a0b0f72da40addb730a0df682

Observation 52069c2e-0ba9-437b-8cac-ff9b27c1b071 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Finite Scalar Quantization: VQ-VAE Made Simple

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.229503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.229503Z digest=sha256:0c3b57f8d0527f710744b7099a725fca407f1c824b38129d9fb125ad5473d95c

Observation fbaad28b-6672-4741-97a4-0bfbedee047a · outbound

This paper cites Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.921578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.235441Z digest=sha256:aaeaaad2c14756e2296fb8071c651961d49ddc663acdda67a18746df998ec0ab

Observation 1d12df17-efb9-41c9-8cea-fa22c4d26da0 · outbound

This paper cites GIVT: Generative Infinite-Vocabulary Transformers.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation GIVT: Generative Infinite-Vocabulary Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.240085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.240085Z digest=sha256:4072a2518f670e46b18b19fc84189a4a8a8a267c04d85890c77e828024ef8224

Observation 58e5e446-5684-49ed-a9fb-a57566674917 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.245217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.245217Z digest=sha256:db8e7e79573c1016b51936429b9cf49fb8dc715a381690ef981c7188fc817fe9

Observation bd83759b-6a84-4ff7-9ff3-95f419ceca7e · outbound

This paper cites The Llama 3 Herd of Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.250162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.250162Z digest=sha256:ac837a730f59a085cef23db2bafa6db37c4841a7db5f1b424bccfb372f97efba

Observation f8e9cff5-7b83-49e2-8c18-1caf4883be3c · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Improved Denoising Diffusion Probabilistic Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.255039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.255039Z digest=sha256:49f8343f4ab59bfbcaf9f129062f546dcac5e2ce27e0cde21e147eb4411c0ec3

Observation 4fba9174-df27-40ef-a29b-901eeec878d1 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.259935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.259935Z digest=sha256:de506a59365d35dd1c9f30f6275baef3902c2596c5962c169af7cf06e985ee97

Observation b1c20c57-3a1e-4fa7-ad98-4618190b6905 · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.265059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.265059Z digest=sha256:fa2a7fe702fbeeed04ada3d92f6a9d5b0d4395221f35fa63ecb0619e760d6be1

Observation 4db2e591-13d1-4e93-93a7-dde297427fd4 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Progressive Distillation for Fast Sampling of Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.270330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.270330Z digest=sha256:49c0bcd14f913d33329a8c1163f6c7239af4ec668e79314df8cdf61a7c64fae3

Observation 62d869dd-57ac-48be-966f-c3e3d4421037 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation A style-based generator architecture for generative adversarial networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.905182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.275413Z digest=sha256:0a8f78e3c6a64c17f14c2820f13dc2dbf9368633f3eaf59a9afa0bf8cde663b5

Observation dec8887e-146a-497a-9fcb-531037e4a453 · outbound

This paper cites Berg, and Li Fei-Fei.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Berg, and Li Fei-Fei

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.280007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.280007Z digest=sha256:fac34299d53a09cb2fae68b833543d9b5611d2845618b36a5693aab7df84df82

Observation e2b7d0af-685f-409a-9696-75ee733aa485 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.284927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.284927Z digest=sha256:bf18f0c294fdb0e3d514c46c52afda592f6f00ae76dcf26459bdec6d4f9d32d2

Observation b359263d-7854-4482-bff0-19015f64197b · outbound

This paper cites Denoising Diffusion Implicit Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Denoising Diffusion Implicit Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.290423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.290423Z digest=sha256:128c0cf9c4802faa99081f395fbc450b15357b0f84bb2b3c38f8daaf5c85d7a8

Observation 87f949b3-56af-4d57-b1b8-c37f3b2ffbad · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.295862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.295862Z digest=sha256:e00d6c17fcffccd2789397b8ed6a4a624cdc64531f3746eb6c6dfa38fac6bf2f

Observation 69d98c81-9564-48fe-8a07-aeefee81b03d · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.300549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.300549Z digest=sha256:f0ed1a36e1c717351924490e272b32958b5c8cc798e66df8947d908808d79317

Observation c019fe1d-5e97-499b-8135-b37cb455203b · outbound

This paper cites Multimodal Latent Language Modeling with Next-Token Diffusion.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Multimodal Latent Language Modeling with Next-Token Diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.305545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.305545Z digest=sha256:898d398d32b18557552bd368c98be935513f70735f2b372a797857028c8f6c1f

Observation 446505ae-d8ab-4536-8a30-77cdf3a0a24c · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.310361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.310361Z digest=sha256:f01b289b02cfa212e2c28fb2bd1ffd9680b16d04c6484e4b81dedcb092bb78b1

Observation 736137f4-926f-4033-9fb4-1590dd309969 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.887324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.315361Z digest=sha256:198f6a4223699adffd25feec8724655b477ea6d1d3ce59741ccbcd6f2fd6c056

Observation d9adef13-b6d4-4578-a126-2b3c99bbc6c7 · outbound

This paper cites Addendum to gpt-4o system card: Native image generation, 2025.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Addendum to gpt-4o system card: Native image generation, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.867367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.320633Z digest=sha256:c509a74852bfbb7c33b607728a365554a79f06cdade713abde592fb036d18e72

Observation 5a82c5ce-1d22-45c3-b8f0-cae910510371 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.325132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.325132Z digest=sha256:e837ad1a03e82ce68889a3531decbc333af9f9cf13786cfdd91874f289f69327

Observation 7afd44aa-38ce-4c14-996a-5b623d8e2a32 · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.330264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.330264Z digest=sha256:fed20df7e58e6ec7bc0b8289c16c807ea5f14460d12a59056e2dd28bcf946a36

Pith citing papers

No inbound Pith citation observations are available.